scale AI coding agent

Claude Code

Anthropic's terminal-native AI agent for deep, agentic work on real codebases

●●●●● Non-coder rating · Updated September 2026
Visit Claude Code →
$20/mo (Claude Pro)
subscription
Best for

Developers who want a powerful terminal-native AI agent for complex codebases

Not for

Non-technical founders; this is a developer tool, full stop

Claude Code visual overview

Claude Code in context: product setup, workflows, and operations

Claude Code is Anthropic’s agentic coding tool, and it represents a fundamentally different philosophy from most things in this space. While other tools give you a chat interface or a visual canvas, Claude Code lives in your terminal. It reads your codebase, understands it, and executes multi-step tasks: writing code, running tests, fixing failures, making commits. It operates more like a pair programmer you can leave running than a chatbot you query.

New in September 2026: hitting your 5-hour limit no longer cuts you off mid-edit

Anthropic shipped a small but genuinely useful fix on September 25: when you hit your 5-hour usage limit in the middle of a task, Claude Code now tries to find a clean stopping point instead of cutting off mid-edit. It draws a small, fixed allowance from your weekly limit to wrap up whatever it can — finish the file it’s touching, leave the repo in a working state — rather than leaving you with a half-written change and no way to fix it until the limit resets.

The allowance isn’t unlimited and it isn’t the same for every plan: Pro subscribers get it once a week, while Max and Team Premium get it at every 5-hour limit they hit. Anthropic confirmed the split directly on its ClaudeDevs account, and it was independently corroborated by several people testing it the same day.

Practical read: this doesn’t change how much you can use Claude Code, it changes how badly a cutoff can hurt. Getting stopped mid-refactor with an inconsistent codebase was the single most annoying part of the 5-hour window; this doesn’t remove the window, it just softens the landing. If you’re on Pro, don’t expect it to bail you out more than once a week — plan your longest tasks accordingly.

New in September 2026: Cloud sessions go GA, with $100–$250 in one-time credit

Claude Code cloud sessions came out of research preview on September 23 and are now available to all Pro, Max, Team, and Enterprise seats. A cloud session runs your task on Anthropic’s own infrastructure instead of your laptop — you kick off a job, close the lid, and check back later for the result, similar in spirit to Cursor’s cloud agents or Devin’s background runs.

Anthropic is handing out one-time promotional balances to mark the GA launch: $100 for Pro subscribers, $250 for Max, on top of your normal plan allowance. Three dates matter here if you want the credit: your subscription had to be active on September 23, you have to claim the balance by October 7, and whatever you don’t spend expires November 4. After that, cloud session usage draws from your regular weekly limits like everything else — this is a trial, not a permanent bonus.

Alongside GA, Claude Code also picked up Projects, which lets you split work into parallel threads instead of one linear conversation, plus a local-thread option for keeping some work on your own machine while cloud threads run elsewhere.

Practical read: if you already pay for Pro or Max, claiming the credit before October 7 costs nothing and is worth doing even if you don’t have an immediate use for cloud sessions — it’s free runway to try offloading a long-running task (a big refactor, a test suite, a migration) without burning your weekly cap. Just don’t let it auto-expire unclaimed. Sources: Claude Code cloud sessions coverage.

New in September 2026: Opus 5.5 is the new default, and it costs 20% less

Anthropic released Claude Opus 5.5 on September 22 and made it the default Opus model in Claude Code 2.1.280, shipped the same day. If you use Claude Code, the model underneath you changed on Tuesday without you doing anything.

The price moved for the first time in five releases. Opus 4.5 through Opus 5 all sat at $5 per million input tokens and $25 per million output; Opus 5.5 is $4 and $20 — a 20% cut. Cache reads dropped 60%, from $0.50 to $0.20 per million, which matters more than the headline for this tool specifically: in a long agentic session most of your input is cache reads. Anthropic’s claim is that typical workloads land around 40% cheaper overall because the model also finishes tasks in fewer tokens. That’s a vendor figure, but the per-token half of it is verifiable and real.

This only reaches your invoice if you’re on API billing. The $20/mo Pro seat is still $20/mo, with the same weekly limits described below. What Pro users may notice is hitting those limits slightly later, since the same work now consumes fewer tokens.

One caution on effort levels. Simon Willison ran Opus 5.5 at “max” thinking and it failed to return a response at all, twice — it reasoned past the 128,000-token output ceiling before producing anything, at $2.56 and roughly 20 minutes per attempt. One test on one prompt, but it’s a reminder that the top effort setting is a cost decision, not a free quality upgrade. Full context in three frontier models in 48 hours.

September 14, 2026: Weekly limits dropped 17%, at an unchanged price

Anthropic’s weekly usage limits on Claude Code changed on September 14, and the framing was confusing enough that it is worth stating plainly.

A temporary 50% boost had been running on weekly limits since May. It expired September 13. A permanent 25% increase over the original baseline took effect September 14. Anthropic’s announcement led with the 25% figure, which is accurate against the pre-promotion floor and misleading against what customers actually had the day before. On a baseline of 100 units: 150 before, 125 after — a 17% cut relative to your September 13 allowance. Anthropic deleted the original post and reposted with that number spelled out, which is a faster and more honest correction than this site has seen from most vendors. BleepingComputer documented the sequence.

The seat price did not move. $20/mo Pro is still $20/mo Pro. Limits are not denominated in prompts — Anthropic says consumption varies by conversation length, model choice, tool usage and effort level — so the practical impact is entirely a function of how close you were already running to the cap. If you were hitting weekly limits, you now hit them roughly a day and a half earlier. If you weren’t close, you will not notice. Full context in two AI coding plans shrank this month.

September 17, 2026: Update to 2.1.179 or newer — plugin pinning was bypassable

Air Security’s Plugin4Shell disclosure, reported by The Register, found that Claude Code checked out a plugin’s pinned commit hash without verifying that the commit was what actually landed. An attacker who controls a plugin’s repository could create a branch named after the pinned hash, make it the default branch, and have git resolve the pin to attacker-controlled code — while the install still reported success at the pinned commit.

Because Claude Code auto-updates installed plugins in the background by default, this was zero-click: no install step, no prompt, no way to notice.

Anthropic patched it in 2.1.179, confirmed June 17, after coordinated disclosure in June. Run claude --version; anything older is exposed. Note that the attack only works on marketplaces hosted where a branch can be named like a 40-character hash — GitHub rejects those names, but Bitbucket and self-hosted git servers do not, and Anthropic’s own documentation lists both as valid marketplace backends. Full breakdown in the fix for last month’s agent supply chain problem doesn’t work.

September 7, 2026: A blocked request took over a domain, 9 times out of 10

Tenet Security’s GhostJacking research, presented at DEF CON 34 and reported by SecurityWeek and Infosecurity on August 10, is the sharpest result yet on this page — and it is worse than it first sounds because nothing in the chain malfunctioned.

An attacker sends a request that Cloudflare’s recommended managed rules will block, with instructions hidden in the User-Agent header. Cloudflare blocks it and logs it verbatim, which is correct behaviour. An engineer later asks Claude Code to review blocked traffic. Claude Code reads the log, treats the planted text as a finding to remediate, rewrites the DNS records to point at an attacker-controlled domain, and reports the issue resolved.

Tenet measured a 90% success rate against Claude Code. Every request had already been blocked by the firewall. The domain was taken over anyway. Parallel chains got Claude Code to execute a command and exfiltrate cloud credentials via a poisoned Datadog alert, and to run attacker code that Sentry’s own AI agent had first laundered into a “proposed fix.”

Tenet also found and disclosed a Claude Desktop sandbox escape that allowed exfiltration to a remote server. Anthropic patched it before the talk, without issuing a CVE — worth knowing if you track this vendor’s disclosure record.

The practical rule, and it is the one that actually generalises: do not give a single agent both read access to outside data and write access to something that matters. If Claude Code can read your Cloudflare logs and also change DNS, or read Sentry and also deploy, you built the vulnerability. Split the credentials or put a human on the write side. Full pattern in every file your agent reads is executable.

Rating held at 2. Nothing here changes the recommendation for non-technical founders, who should not be running this tool at production credentials in the first place. But this is the third September security finding on this page, and the second where Claude Code’s own competence is the delivery mechanism.

September 6, 2026: Claude Code installed unowned packages inside real corporate networks

Research from Pandex, covered by Ars Technica and picked up by Bruce Schneier on September 4, found 227 install commands across 120 companies’ llms.txt documentation files pointing at package names and domains nobody owned. The researchers registered some of them, hosted harmless packages that phoned home, and logged which agents executed them. Claude Code was one of three named, alongside Codex and Hermes. Anthropic did not respond to a request for comment.

The trigger is not an attack. It is a normal prompt — “using [vendor]‘s docs, build and run a project with their SDK” — and the agent doing exactly what it should: reading the vendor’s official machine-readable docs and following the install instructions in them. The instructions were written in good faith. The package name they pointed at was simply never claimed.

Claude models were the least likely to execute the payload of anything tested — Opus 4.8 at medium effort ran it around 30%, against 90%+ for GPT-5 Luna and Sol. That is a real relative advantage and it is not a defence. Thirty percent of your dependency installs is not a rate you can live with.

Practical guidance: carve npm install and pip install out of any auto-approve rules you have, and read the package name before you accept it. Full detail and the check to run in your agent trusts the vendor’s docs.

September 5, 2026: one GitSpawn path is fixed, one isn’t

Manifold Security’s GitSpawn research, published September 1, contains two Claude Code findings, and they resolved very differently.

The first is the core.fsmonitor path: opening a folder with claude runs git status as an internal subprocess before the workspace-trust prompt, and the repository’s own .git/config gets to name a program git will execute. Reported June 26, fixed by 2.1.196 (CVE-2026-55607). If you are on any current version you are covered here.

The second is not fixed. claude ultrareview reaches the same class of bug through a different git config key that the review path does not strip, and Manifold has deliberately withheld which key while it stays live. Reported July 15 on 2.1.210, closed as a duplicate of an internal ticket, and re-confirmed reproducible on 2.1.252 on September 1. Seven weeks.

Practical guidance until it ships: do not run claude ultrareview on a repository you did not clone yourself. git clone does not carry a .git/config, so cloned repos are safe; zips, shared drives, synced folders and dev-container images are not. See the repo someone sent you can run code before you type anything for the check to run first.

Also in v2.1.261 (September 4). Two additions are worth flagging. /skill-doctor now reports which loaded skills go unused and what they cost you in context — a real answer to the slow context bloat anyone with a dozen installed skills has been living with. And auto mode now treats a link that packs content into a public diagram renderer’s URL as an upload to that site, so it is no longer auto-approved unless you asked for it. That is a quiet fix for a genuine exfiltration path: encoding your source into a URL and calling it a rendering request.

Housekeeping in the same release: bashOutputMaxChars and taskOutputMaxChars let you raise inline command output to 128K characters before it spills to a file, and the /model picker now shows a model’s real name instead of a raw Bedrock or Vertex ID.

New in September 2026: Fable 5.1 becomes the default, and a batch of permission fixes worth reading

Two things in the first three days of September, both verified against Anthropic’s own changelog rather than a release tracker.

Claude Fable 5.1 shipped September 1 (v2.1.257) as the default Fable model — 1M context, $10/$50 per million tokens with $0.25/Mtok cache reads. If you’re on a Claude apps gateway, note that fable and best deliberately still resolve to Fable 5, because gateways not yet configured for 5.1 reject it; pick Fable 5.1 explicitly in /model to use it. The same release added a Containment Escape rule to auto mode, so cloud metadata-credential fetches, egress evasion and cross-tenant reach are no longer auto-approved unless your environment marks them expected. That is a sensible narrowing of what auto mode will do without asking.

The September 3 release (v2.1.260) is mostly bug fixes, and two of them matter more than a bug fix normally would. A permission rule whose path contained parentheses — Edit(./my-project (old)/**), say — was being dropped as invalid or ignored by the Bash sandbox, “which left ‘read-only’ folders writable.” A second fix covers a single file rule with an uncompilable pattern making every file edit fail. And a third stops Bash permission checks auto-approving zsh commands that hide a command substitution inside a REPORTTIME, REPORTMEMORY or DIRSTACKSIZE assignment.

The practical read: if you have been relying on deny rules to keep an agent out of a directory, and any of those paths contain brackets or parentheses, the rule may not have been doing anything. Update, then re-check your rules rather than assuming they held. This is the second month running where the notable Claude Code security work is in the gap between what a permission rule looks like it says and what it actually matched.

Also in 2.1.260: a /diff panel that opens beside the conversation in fullscreen and shows uncommitted changes as Claude edits, and /cost now names the likely cause of a prompt-cache miss. Both are small, both are the kind of thing you notice every day.

New in August 2026: A mode that gives Claude Code less, and a credential leak it just closed

Two entries from this week’s releases belong together, because one is a fix for a problem and the other is the beginning of a structural answer to it.

The fix, in 2.1.247 on August 26: /ultrareview and locally seeded cloud sessions had been uploading uncommitted edits to prod.env-style and *.tfvars files, along with editor swap, temp and backup copies of credential files — the changelog names examples like key.pem.tmp and id_rsa.swo. Those now stay on your machine. If you’ve run /ultrareview on a working tree with production environment files or Terraform variables in it, rotate what was in them.

The structural answer, in 2.1.248 on August 27: a new --restricted flag (or CLAUDE_CODE_RESTRICTED=1) that “removes the built-in tools that run commands or code and WebFetch (unless named in --tools), keeps file tools inside the working directory, refuses bypassPermissions, and ignores user, project and local settings files.” That last clause is the interesting one — a restricted session can’t be quietly re-widened by a settings file someone dropped in the project. Note that as of today the flag is in the changelog but not yet on the permission-modes docs page, so the changelog is the authoritative description.

This is a notable turn. Two weeks ago the headline change here was auto mode becoming the default, which gives Claude more latitude. --restricted goes the other way, and it’s the right tool for a specific job: pointing an agent at code, dependencies or issues you didn’t write. Use auto mode for your own work; use --restricted when the input is untrusted.

Same release also fixed a prompt-cache miss that was firing roughly once an hour in long sessions after an OAuth token refresh, which quietly cost money and dropped extended-thinking context, and stopped Claude Desktop and Cowork sessions disappearing after 30 days. Source: Claude Code changelog. Founder-facing context on what agents read: your .env file is the first thing your coding agent reads.

One caution on vocabulary, added August 28: “restricted mode” does not mean the same thing across tools. Claude Code’s --restricted gives you a working agent with fewer powers. Cursor also has a restricted mode, but it’s VS Code’s workspace trust feature, it’s off by default, and Cursor’s docs say it breaks AI features and recommend a plain text editor for untrusted repos instead. Don’t assume a control transfers because the label matches. Comparison of what each tool actually gates: what is your coding agent actually allowed to touch?.

New in August 2026: You can finally see which loop is burning your tokens

Version 2.1.243, shipped August 25, adds a Loops breakdown to /usage — per-loop run count, total tokens, tokens per run, and last run time. If you’ve set up recurring /loop tasks and watched your usage climb without knowing which one was responsible, that’s now a single command.

This is a small feature with an outsized effect on the thing this site keeps flagging: agent costs are unpredictable because nobody can see them at the right granularity until the bill arrives. A per-loop token figure is the granularity that lets you kill the chatty one.

Two other settings landed in the same release worth knowing about. modelPricing lets an organisation plug in its contracted per-model rates and discount multiplier so /cost, the status line, and telemetry report what you actually pay rather than list price — if you’ve negotiated an enterprise rate, your cost figures have been wrong until now. And promptCacheTtl / subagentPromptCacheTtl let API-key and cloud-provider users hold a one-hour prompt cache on the main conversation while subagents stay at five minutes, which is a direct lever on the cache-read costs that dominate long sessions.

Version 2.1.246, same day, adds an Auto mode tab to /permissions for viewing and editing the auto-mode classifier rules. Given that auto mode became the default on August 14 (below), being able to inspect and edit what the classifier will and won’t wave through is the missing half of that change. Source: Claude Code changelog.

New in August 2026: Claude Code gets a front door that isn’t a terminal

On August 20, Slack launched Slack Code, and Claude Code is one of four founding partner agents alongside Devin, GitHub Copilot, and Vercel’s agent. Tag Claude from a Slack conversation and it opens a project-specific code channel where the whole team can watch the work: code diffs, live previews, a running plan in dedicated tabs, and an approval step before anything ships. The channel archives itself when the task completes and leaves an audit log. It’s available on every Slack plan including free workspaces, though you still need your own Claude subscription.

This doesn’t change what Claude Code is, but it does change who can reach it. The rating stays at 2 because the underlying tool is still a developer product and still assumes someone can judge a diff. What Slack Code removes is the requirement that the requester be that person. A support lead who spots a bug in a Slack thread can now tag Claude directly instead of hoping the report survives a trip through a backlog.

Worth noting the lineage: Anthropic shipped a research preview of Claude Code in Slack threads back in December 2025, so this is the productised version of a pattern that has been quietly working for eight months. If you’ve written Claude Code off as terminal-only, this is the release that makes it worth a second look. Source: VentureBeat, Unite.AI, Salesforce.

New in August 2026: Auto mode becomes the default, and a classifier holds the permission button

As of August 14, new Claude Code sessions on Pro, Max, and Team start in auto mode by default. Instead of stopping to ask you before each file write, shell command, or network call, every tool call is routed through a safety classifier that blocks anything irreversible, destructive, or aimed at resources outside your environment. When it blocks something, Claude either finds a safer route or asks you. Three consecutive blocks, or twenty in a session, and it drops back to manual approvals. If you’d previously set a different default, you get a one-time prompt asking whether to switch. It stays opt-in on Enterprise, the API, Bedrock, Google Cloud’s Agent Platform, and Microsoft Foundry for now, with the same default expected there within a month.

The justification is more interesting than the feature. Anthropic says developers approve 97% of permission prompts — which is a polite way of saying the prompts stopped working. In a controlled study with 1,053 paid testers, humans blocked 13.6% of dangerous commands; the classifier blocked 89%. Independent testing by Trajectory Labs found none of 720 prompt-injection attacks succeeded against Claude in auto mode, versus a 5.83% success rate against GPT-5.6 Sol in Codex’s auto-review mode and 19.03% with Codex in full access. Anthropic has also stopped charging for the classifier’s extra tokens, so this doesn’t come out of your usage.

The honest caveat: auto mode reduces risk, it doesn’t remove it. You’re swapping a human rubber-stamp for an AI reviewer that Apollo Research still got past 7% of the time on deliberately adversarial attacks. Anthropic’s own advice is to keep reviewing before high-risk changes — production systems, anything with real customer data. If you want the old behaviour, it’s still a setting. For the founder-facing read on what this means for every agent you run, see Your coding agent now approves its own actions. Source: Anthropic, Help Net Security, The Register.

New in June 2026: Claude Fable 5 lands as the new top-class model

On June 9, Anthropic shipped Claude Fable 5 and made it available in Claude Code. It’s a genuine step up on exactly the work Claude Code does: it scores 80.3% on SWE-bench Pro (versus 69.2% for Opus 4.8) on the hard, whole-codebase variant, plus roughly 95% on SWE-bench Verified, with a 1M-token default context window. In practice that means fewer silent mistakes on long multi-file refactors and better reasoning on the gnarly bugs where earlier models stalled.

Two things to know before you lean on it. First, the economics: Fable 5 is included free in Pro, Max, Team, and seat-based Enterprise plans only through June 22. On June 23 it moves to metered usage credits at API rates ($10 per million input, $50 per million output tokens, roughly double Opus 4.8). Second, this doesn’t change who Claude Code is for. It’s still a terminal tool for developers, not non-coders. But if Claude Code is already in your stack, restart your sessions to pick up Fable 5, and use the free window to test it on a real, hard task before deciding whether the post-June-23 cost is worth it. We break down the implications for builders in our Claude Fable 5 guide. Source: Anthropic, TechCrunch, Simon Willison.

New in June 2026: The pricing reset we predicted arrives on June 15

The plan restructure we’ve been flagging since April is now confirmed and dated. On June 15, Anthropic splits Claude usage into two billing pools. Interactive usage, meaning you typing into Claude.ai, the desktop and mobile apps, Cowork, or Claude Code live in the terminal, stays on your normal subscription limits, unchanged. Programmatic usage (the Agent SDK, the claude -p command, Claude Code in GitHub Actions, and third-party agents built on Claude) moves onto a separate monthly credit pool (roughly $20 on Pro, $100 on Max 5x, $200 on Max 20x), metered at full API list prices with no rollover.

What this means in practice: if you drive Claude Code by hand in the terminal, your day-to-day is unaffected. If you’ve wired Claude Code into CI, scheduled jobs, or autonomous pipelines, that spend is now a visible, capped line item rather than a hidden subsidy on a $20 plan. The change closes an arbitrage where heavy Agent SDK users were drawing several hundred dollars of API-equivalent compute for $20, so it’s a correction, not a broad price hike on interactive users. If Claude Code is central to an automated workflow, audit what runs without a human before June 15 and budget for the credit cap. Full breakdown in our guide to the June 15 billing change. Source: Codersera, DevToolPicks, Zed.

The non-coder rating here is 2. Not because Claude Code is bad, since it’s arguably the most capable agent in this category, but because it was built explicitly for developers. If you’re reading this and you’re not writing code yourself, you can stop here.

New in May 2026: Claude Opus 4.8 ships as the default model

On May 28, Anthropic released Claude Opus 4.8 and made it the default model in Claude Code. It’s a point release on the same $5/$25-per-million rate card as Opus 4.7, so there’s no price change, just a better model under the hood. Two things matter for the work Claude Code actually does. First, the context window moves to 1M tokens by default, so the agent can hold a much larger slice of your codebase before it starts losing the thread on big multi-file tasks. Second, Anthropic reports the model is roughly four times less likely than Opus 4.7 to let a flaw in code it just wrote pass unflagged, and it scores 69.2% on SWE-bench Pro (up from 64.3%), which is incremental on paper, but it shows up as fewer silent mistakes on long agentic runs.

Alongside the model, Anthropic shipped Dynamic Workflows for Claude Code (research preview), which lets the agent adjust its plan mid-task rather than committing to a fixed sequence up front, plus effort control so you can dial reasoning up for hard problems or down for speed and rate-limit savings. The same model also rolled out to Cursor and the Claude API the same week, so the upgrade isn’t a Claude Code exclusive.

Nothing here changes the verdict. This is still a terminal tool for developers. But if you’d already adopted Claude Code, the upgrade is free and worth restarting your sessions for. Source: Anthropic, The New Stack, Simon Willison.

New in May 2026: Self-hosted sandboxes and MCP tunnels at Code with Claude London

On May 19, at the Code with Claude London event, Anthropic shipped two new infrastructure features for Claude Managed Agents that matter for any team running agents in a regulated or security-conscious environment.

Self-hosted sandboxes (now in public beta) let tool execution happen on infrastructure you control: your own servers or a managed provider you choose (Cloudflare, Daytona, Modal, and Vercel are all supported). The agent loop that handles orchestration and context management stays on Anthropic’s infrastructure; what moves to your environment is the actual execution of code and tool calls. For teams where data governance requires that sensitive files, packages, or services never leave their own perimeter, this removes the biggest blocker to deploying Claude Managed Agents on real work.

MCP tunnels (research preview) let Managed Agents and the Messages API connect to private MCP servers without exposing them to the public internet. If your company has internal tooling (a proprietary data store, an internal Jira instance, a private API) you can now give Claude agents access to those systems over an encrypted tunnel rather than having to open a public endpoint.

Neither feature is aimed at non-technical founders. Both require infrastructure setup and are firmly in developer/DevOps territory. But for any technical co-founder whose development team is evaluating Claude Code for production-grade agentic work, these are the two features that unlock the enterprise deployment conversation. Source: The Decoder, InfoQ.

New in May 2026: Usage limits doubled as the SpaceX compute deal closes

On May 6–7, Anthropic announced it has signed a compute agreement with SpaceX to access the Colossus 1 data center in Memphis: more than 300 megawatts of capacity, over 220,000 NVIDIA GPUs, with more potentially following in orbit. That compute is flowing directly into Claude Code’s rate limits.

What changed for users: the five-hour rate limits for Claude Code are now doubled across Pro, Max, Team, and seat-based Enterprise plans. Anthropic has also removed the peak-hour limit reductions for Pro and Max, which means you’re no longer throttled during business hours. API users are seeing even bigger jumps: Tier 1 API users saw a 1500% increase in maximum input tokens per minute and a 900% increase in output tokens per minute for Opus models. These changes are live now and apply automatically, with no plan changes required.

For technical founders who’ve been hitting limits mid-session on long agentic runs, this is material. The throttle that cut your Routine short or killed your context mid-refactor is largely gone. The caveat: the underlying economics still don’t add up at $20/mo for heavy usage, and the pricing change we flagged in April is still likely coming. Use the current limits while they hold.

New in May 2026: Code Review and Remote Agents

Also announced at the Code with Claude 2026 developer event (May 6, San Francisco): Claude Code now includes a native built-in Code Review tool that can analyze a PR or diff and return opinionated, structured feedback: not just syntax notes, but logic errors, test coverage gaps, and security concerns surfaced in the same terminal-native interface. This is Anthropic’s direct response to BugBot in Cursor.

Remote Agents extend Claude Code’s reach beyond the laptop: you can now trigger and monitor Claude Code sessions from your phone, with the agent running on Anthropic’s cloud infrastructure (same as Routines from April). If you kicked off a long refactor and stepped away, you no longer need to be at your desk to check progress or intervene. Source: Code with Claude 2026 event coverage.

New in late April 2026: Anthropic owns the monthlong quality decline

On April 23, Anthropic published an engineering postmortem acknowledging that a series of engineering missteps, not user error and not phantom regressions, were behind the widely-reported drop in Claude Code quality between early March and mid-April. Fortune and VentureBeat both ran the story prominently, and the user response on the Anthropic Discord and on X was sharp: a notable number of paid users said they’d cancelled, and a senior AMD AI exec called the tool “unusable for complex engineering tasks” during the worst of the period.

Three changes were responsible. On March 4, Anthropic cut Claude Code’s default reasoning effort from high to medium to reduce latency. Anthropic now says that tradeoff was wrong and reverted it on April 7. On March 26, a caching change meant to clear stale thinking from idle sessions instead cleared it every turn, which is why Claude Code felt forgetful and repetitive for weeks; that bug was patched on April 10. On April 16, a system-prompt instruction was added to cap responses at 25 words between tool calls. It measurably hurt coding output and was reverted on April 20 (in v2.1.116). The API was unaffected throughout; the regressions only hit Claude Code’s product-side defaults.

Two things matter for non-technical founders considering Claude Code today. First, the quality issues are over. If you tried Claude Code in March or early April and bounced, the underlying tool you’re returning to in late April is materially better. Second, Anthropic’s initial communication implied users were largely to blame before the company walked that back. That’s a yellow flag on trust, not a red one, but it’s worth weighing alongside the technical strengths. Vibe coding tools depend on the underlying AI lab being honest about regressions in real time. Anthropic eventually got there; “eventually” is the operative word.

New in April 2026: Opus 4.7, /ultrareview, and task budgets

On April 16, Anthropic shipped Claude Opus 4.7 and rolled it out as the default model in Claude Code. Three changes matter for day-to-day work. First, there’s a new /ultrareview slash command that scans a file for bugs, security issues, and logical gaps in a single pass, useful as the final pre-commit sweep before you push a PR. Second, Opus 4.7 introduces an xhigh (“extra high”) reasoning effort level that sits between high and max, giving you a middle lever when high is underthinking a problem but max is overkill on latency and cost. Third, task budgets graduated from beta: you can now cap token spend on any autonomous run, which matters a lot once Routines are handling production work unattended.

The model itself shows up in long-horizon agentic work. Anthropic’s stated Opus 4.7 improvement is in “systems engineering and complex code reasoning,” the kind of task where Claude Code has to hold multiple files, a test suite, and a desired end state in its head at once. Early reports on the Anthropic Discord and on X describe fewer turns to complete multi-file refactors and a lower false-start rate on hard bugs. If you haven’t bumped a stuck session to xhigh yet, try it. It’s the first reasoning lever in a while that feels genuinely different from “more of the same.”

New in April 2026: Routines and a redesigned desktop app

On April 14, Anthropic shipped two updates that change how Claude Code fits into a real engineering workflow. The first is a full redesign of the Mac and Windows desktop apps: integrated terminal, faster diff viewer, in-app file editor, expanded preview area, and proper multi-session support so you can run several Claude Code instances in parallel without constant app-switching. This is the first time the desktop experience has felt like a primary surface rather than a thin wrapper over the CLI.

The bigger news is Routines, in research preview for Pro, Max, Team, and Enterprise subscribers. A Routine is a saved Claude Code configuration (a prompt, one or more repositories, and a set of connectors) that runs automatically in the cloud instead of on your machine. There are three flavors: Scheduled Routines (cron-like jobs for things like nightly docs-drift scans or backlog triage), API Routines (HTTP endpoints with auth tokens you can hit from Datadog, PagerDuty, or a CI pipeline), and event-driven Routines (trigger on a GitHub webhook, for example). Because Routines run on Anthropic’s web infrastructure, your laptop doesn’t need to be open.

For technical founders and ops-minded engineers, this is the most interesting update Claude Code has shipped in months. It moves the tool from “agent you run in a terminal” to “agent that runs your on-call response, your nightly cleanup, and your triage queue.” The obvious caveat: the more autonomous the agent, the more carefully you need to scope what it can touch. Start with read-only Routines (reports, scans, summaries) before you let one open PRs unattended.

What makes it different

Most AI coding assistants operate in one of two modes: chat-based code generation (you ask, it answers) or IDE-integrated suggestions (Copilot-style autocomplete). Claude Code operates in a third mode: genuine agentic execution. Give it a task (“refactor this authentication module to use JWT” or “find and fix all the broken tests in the payments service”) and it will work through the problem methodically, using tools to read files, run commands, check output, and iterate.

The context window handling is exceptional. Claude Code is built on Claude’s large context window and uses it to hold an accurate model of your entire codebase, not just the file you have open. This makes it markedly better than most alternatives at tasks that require understanding relationships across files and modules.

Terminal-native matters

The decision to ship this as a terminal tool rather than an IDE plugin or web interface was deliberate. It means Claude Code integrates cleanly with any development environment and workflow. It works with your existing version control, your test runners, your build tools. There’s no proprietary layer that interposes between the AI and your actual code.

Pricing reality

The $20/mo Claude Pro subscription gets you access, but heavy usage will hit rate limits. Teams doing serious agentic work will likely need the API-based billing path, which is pay-as-you-go and can add up depending on codebase complexity and task length. Budget accordingly if you’re planning to use this heavily for large codebases.

Pricing uncertainty worth noting (April 21-22)

Between April 21 and April 22, Anthropic quietly removed Claude Code from the Pro plan feature list on its public pricing page for a subset of new signups. The Register and Simon Willison both documented the change while it was live. Anthropic’s head of growth called it “a small test of 2% of new prosumer signups,” and the Pro pricing page was reverted within a day. Existing Pro subscribers were not affected and the official public pricing is still $20/mo for Claude Code access today. But the signal is clear: the token economics on a $20 plan with heavy Claude Code usage don’t work, and some form of plan restructure (higher price, tighter caps, or a separate Claude Code tier) is likely in the next few months. If Claude Code is central to your workflow, assume your effective monthly cost could move up meaningfully by mid-year and plan accordingly.

Limitations

The agent can make mistakes, particularly on large multi-step refactors where early incorrect assumptions compound. It requires human oversight. You should be reviewing diffs before merging, not accepting changes blindly. The terminal interface, while powerful, has a learning curve for developers who haven’t worked with agentic tools before.

Documentation and onboarding materials were still maturing as of this writing. The tool rewards users who invest time understanding its capabilities and limitations; it punishes those who treat it as magic.

Who it’s for

Senior engineers working on complex existing codebases. Technical co-founders who want to move faster on architecture and refactoring work. Any developer comfortable in a terminal who wants a powerful AI agent they can trust with non-trivial tasks.

Verdict

Among developer-focused AI agents, Claude Code is one of the best available. The combination of genuine agentic capability, large context handling, and terminal-native design makes it stand out. But “best developer tool” and “useful for non-technical founders” are different categories, and this firmly only occupies the first one.

Was this helpful?
Related tools All tools →
Cline
AI coding agent

Open-source agentic coding assistant for VS Code: bring your own model, see every move

●●●●● Free · Free + your own API keys
CodeRabbit
AI coding agent

AI code review that reads every change your agent makes before it ships

●●●●● Free · $24/dev/mo (annual)
Devin Updated
AI coding agent

The first AI software engineer — now $20/mo, still built for people who can read a pull request

●●●●● Free · $20/mo