Blog

Claude Code review: what it catches and where it falls short

Claude Code review is three products: managed PR reviews averaging $15-25, /code-review and a GitHub Action. Pricing, setup, limits and benchmark scores.

Alex Mercer

Claude Code review is three products under one name: Anthropic's managed Code Review for GitHub pull requests, the /code-review command in Claude Code, and a GitHub Action you run on your own runners. The managed one costs $15-25 a review on average, and only Claude Team and Enterprise plans can turn it on.

This blog belongs to cubic (cubic.dev), an AI code review tool that competes with Claude's managed review, so weigh our comparisons accordingly.

Which Claude Code review should you use?

Use /code-review to check your own branch before you push; every Claude Code user has it. Managed Code Review gives deeper Claude AI code review on pull requests, but at $15-25 a review it only makes sense in Manual mode, on the few large PRs that need it, and only if you're already on Claude Team or Enterprise. To choose the prompt, model and triggers yourself, use the GitHub Action.


Managed Code Review

Local commands

Claude Code GitHub Action

What it is

Anthropic's agents review your GitHub PRs and post inline comments

/code-review (alias /review), /security-review and /code-review ultra, run in a Claude Code session

anthropics/claude-code-action runs Claude Code in your own workflows

Runs on

Anthropic's infrastructure

Your session; ultra runs in Anthropic's cloud

Your GitHub Actions runners

Who can use it

Claude Team and Enterprise plans (research preview)

Any Claude Code user; ultra needs a claude.ai login

Anyone with a Claude API key, or a Pro, Max, Team or Enterprise subscription

Cost

$15-25 per review on average, in usage credits

Normal Claude Code usage; ultra has 3 free runs on Pro and Max, then about $5-25 a run

GitHub Actions minutes plus API tokens or subscription usage

Code hosts

github.com and GitHub Enterprise Server

Any git repo; --comment posts to GitHub PRs and GitLab merge requests

GitHub

How to start

An Owner turns it on in Claude's admin settings

Type the command

Run /install-github-app

Prices are as of October 2026, from Anthropic's docs for Code Review, ultrareview and the GitHub Action. Team and Enterprise plan prices are on Anthropic's pricing page.

How Anthropic's managed Code Review works

Code Review is Anthropic's managed Claude Code review tool for GitHub pull requests, launched on March 9, 2026, and still a research preview for Team and Enterprise plans.

A fleet of agents reads each diff in parallel against your whole codebase, looking for logic errors, security vulnerabilities, broken edge cases and subtle regressions. By default it skips formatting and missing tests. A verification step drops candidates that don't hold up, and Claude posts the rest as severity-ranked inline comments with a summary, labeling bugs that predate the PR.

It never approves or blocks a PR: the check run always concludes neutral, and Anthropic's launch post says approval is still a human call. To gate merges on its findings, have your own CI read the check run's machine-readable severity count.

How to set up Claude Code Review

  1. An Owner or Primary Owner of your Claude organization, with permission to install GitHub Apps in your GitHub organization, opens claude.ai/admin-settings/claude-code and clicks Setup under Code Review.

  2. Install the Claude GitHub App on the repositories you want reviewed, then choose which of them Code Review covers.

  3. Set Review Behavior for each repository: Once after PR creation, After every push (which also resolves threads when you fix a flagged issue) or Manual.

  4. Open a test PR. A Claude Code Review check run should appear within a few minutes; in Manual mode, comment @claude review first.

Anyone with write access can request a review in any mode with a top-level @claude review comment on an open PR; the docs add @claude review always to review the PR on every push.

Why didn't Claude Code Review run?

In Manual mode, reviews wait for someone to comment @claude review, and fork PRs need that comment in every mode. If your GitHub organization membership is private, GitHub's default, your comment won't start a review unless you're a direct repository collaborator. Over the spend cap or out of credits, Code Review skips it and says why in a PR comment.

Customizing reviews with CLAUDE.md and REVIEW.md

Code Review reads your CLAUDE.md, flags new violations of it as nits and flags PRs that make it out of date. REVIEW.md holds review-only instructions: what counts as serious in your repo, how many nits to post, which paths to skip and what to check on every PR.

How much does Claude Code Review cost?

Anthropic puts the average cost at $15-25 per review, billed by token usage and rising with PR size, codebase complexity and the number of findings to verify. It comes out of usage credits, separate from your plan's included usage, and there's no free tier.

The charge lands on your Anthropic bill even if you run other Claude Code work through Amazon Bedrock or Google Cloud.

The trigger you pick sets the bill. After every push multiplies the cost by the number of pushes, and Manual costs nothing until someone asks. Say your team opens 200 PRs a month and each gets one Claude review: at Anthropic's average, that's $3,000-5,000 a month in usage credits, on top of your seats.

Claude's analytics show Code Review usage, admin settings show each repository's average cost, and admins can cap monthly spend at claude.ai/admin-settings/usage. For other tools' prices, see AI code review pricing.

Is Claude Code Review worth paying for?

For most teams, no. At $15-25 a review on average, and more if it reviews every push, it's the most expensive way to get an AI review on a pull request. A 10-developer team opening 200 PRs a month would pay about $3,000-5,000 a month in usage credits, on top of its Claude seats. The same team pays $300 a month for cubic on Team, or $790 on Pro, billed yearly.

Paying more doesn't buy a better score either. On Martian's online tracker (last-month view, checked October 4, 2026), cubic scores 65.3% F1 and the Claude GitHub App 60.2%. On the offline benchmark (as labeled Sep 8, 2026), cubic scores 64.9% F2 and Martian's "Claude Code Reviewer" entry 47.7%. cubic is our product, so check those numbers yourself; the caveats are in the benchmark section below.

If you're on Claude Team or Enterprise anyway, keep Code Review in Manual mode for the occasional large PR, where Anthropic's own data shows it finds the most (findings on 84% of PRs over 1,000 changed lines), and use /code-review, which runs on normal usage, for everything else.

The Claude Code /review command (now /code-review)

/review in Claude Code is an alias of /code-review, since v2.1.223 (August 6, 2026). Older guides describe it as a separate command for reviewing a GitHub PR, and /code-review itself was called /simplify until v2.1.147 (May 21, 2026).

/code-review reviews your branch's commits ahead of upstream plus any uncommitted changes, or a PR number, branch, path or ref range you pass it, at an effort level from low to max:

/code-review high 1234
/code-review high 1234
/code-review high 1234

low and medium report only its most confident findings. --fix applies the findings to your working tree, and --comment posts them as comments on the PR. It counts toward your normal Claude Code usage and follows CLAUDE.md but not REVIEW.md.

Local reviews run only when someone remembers to run them, and the findings stay in that session unless someone posts them.

/code-review ultra

/code-review ultra (alias /ultrareview) is the deep option, in research preview. It sends your branch or PR to a fleet of agents in an Anthropic cloud sandbox, and Anthropic says each finding it reports is independently reproduced and verified. A run typically takes 5 to 10 minutes. Pro and Max get three one-time free runs, then pay about $5-25 a run in usage credits.

On Pro or Max, where managed Code Review isn't offered, we'd treat ultra as the closest substitute.

Claude Code review in GitHub Actions

anthropics/claude-code-action runs Claude Code in your GitHub workflows, answering @claude mentions or running your own prompt on any GitHub event. It's MIT-licensed.

Set it up with /install-github-app from Claude Code inside the repository, as a repo admin with a signed-in GitHub CLI (gh), then merge the workflow PR it prepares.

The review workflow runs Anthropic's code-review plugin on new and updated PRs. Four agents work in parallel: two check CLAUDE.md compliance, one looks for obvious bugs and one reads git blame and history. The plugin scores each issue from 0 to 100 and by default posts only those at 80 or above. If your CLAUDE.md is thin, half of this review has little to check.

It skips drafts, closed PRs, PRs it judges trivial or automated, and any PR Claude has already commented on, so later pushes to that PR get no fresh review. Two more gotchas:

  • Workflows generated before v2.1.229 (August 12, 2026) wrote the review to the run log instead of the PR. If yours looks silent, rerun /install-github-app and update the workflow file.

  • On public repositories, GitHub withholds secrets from fork PRs, so the review runs only on branches in the same repository.

Anthropic positions the Action as the lighter, cheaper option and managed Code Review as the deeper, more expensive one. The upkeep is yours: workflow files, secrets, runner minutes and upgrades.

Claude Code security review

Claude Code security review is two separate tools. The /security-review command checks the diff between your branch and origin's default branch for risks such as injection, auth issues and data exposure. Anthropic's support article says every Claude Code user has it, as a complement to your other security practices.

For pull requests, the separate anthropics/claude-code-security-review action reviews PR diffs. Its README says it isn't hardened against prompt injection and should only review trusted PRs, so keep it off pull requests from outside contributors.

Is Claude Code review any good?

Yes, Claude is a credible reviewer. On Martian's Code Review Bench, reviews posted by the Claude GitHub App rank #6 of 14 tools at 60.2% F1 (67.6% precision, 54.2% recall) in the online tracker's last-month view, checked October 4, 2026. That puts it in the same pack as GitHub Copilot and CodeRabbit.

The tracker scores review bots on real open-source pull requests. Precision is the share of a reviewer's comments that developers acted on, recall is the share of real fixes it caught, and F1 balances the two. The middle of the board is tight: six tools, Claude among them, sit between 60.0% and 61.9% F1. cubic ranks #1 at 65.3% F1 (72.0% precision, 59.7% recall).

Anthropic's launch post reports two numbers from its own engineers: the share of PRs with substantive review comments rose from 16% to 54%, and engineers marked under 1% of findings as incorrect. A finding nobody marks incorrect can also be one nobody acts on, and acting on findings is what Martian measures.

Read the benchmark with care:

  • The tracker can't tell managed Code Review from the Action, since both post as the Claude GitHub App and every Action user picks their own prompt and model.

  • Martian says its online data can't support head-to-head comparisons, because tools see each other's comments and the repos that adopt each tool differ.

  • The board moves daily, and several vendors, cubic included, have each reported #1 at different dates and in different modes. On Martian's offline benchmark (50 hard bugs, as labeled Sep 8, 2026), Qodo's Deep configuration leads cubic by 0.2 F2 points.

Our read is that the tools differ more in what surrounds the review: team rules, feedback, approval, scans and price.

Is Claude Code better than ChatGPT for code review?

On Martian's Code Review Bench, yes. In the online tracker's last-month view, checked October 4, 2026, the Claude GitHub App ranks #6 at 60.2% F1 and the ChatGPT Codex Connector #11 at 56.0%. Martian warns against head-to-head readings of its online data, so treat the gap as a hint.

Where Claude Code Review falls short

Anthropic documents each of these limits itself:

  • Plans and data. It's Team and Enterprise only, in research preview, and closed to organizations with Zero Data Retention.

  • Code hosts. It supports github.com and self-hosted GitHub Enterprise Server, but not GitHub Enterprise Cloud with data residency (*.ghe.com), GitLab, Bitbucket or Azure DevOps.

  • Speed. A review takes about 20 minutes on average. Anthropic built it for depth over speed.

  • No conversation. Replies to a finding get no response; to act on one, you push a fix. Thumbs up and down go to Anthropic to tune the reviewer and change nothing on your PR.

  • Big PRs. On very large PRs, REVIEW.md can be cut or left out. The check run now says when that happens.

An AI review also doesn't replace a person who knows why the change exists. Our code review checklist covers what that person should still check.

Claude Code Review vs GitHub Copilot and cubic


Claude Code Review (managed)

GitHub Copilot code review

cubic

Who can use it

Claude Team and Enterprise (research preview)

Paid Copilot plans and Copilot Student

Any GitHub organization; free plan with 20 PR reviews a month; free for public repos, with fair-use limits

Price

About $15-25 per review in usage credits, on top of the plan

$10-100 a month for individuals, $19 or $39 per seat for organizations; since June 1, 2026, reviews also use AI credits and Actions minutes

$40, $99 or $200 per developer a month (Team, Pro, Max), or $30, $79 or $160 billed yearly

Code hosts

github.com and GitHub Enterprise Server

GitHub; Azure DevOps in preview

GitHub only

Team rules

CLAUDE.md, REVIEW.md

.github/copilot-instructions.md, path-specific instruction files, AGENTS.md

Custom agents with path filters, cubic.yaml; also reads CLAUDE.md and AGENTS.md

Learns from feedback

Thumbs reactions, which Anthropic uses to tune the reviewer

Copilot Memory (public preview), thumbs feedback

Learnings from replies, reactions and senior reviewers' comments, scoped to your team

Approves PRs

Never; the check run is always neutral

Comment-only by default; approvals in public preview, off by default

Optional auto-approval on Team, Pro, Max and Enterprise, with a shadow mode to test it first

Martian online F1

60.2% (#6, as the Claude GitHub App)

61.0% (#4)

65.3% (#1)

Prices are as of October 2026, from Anthropic's docs, GitHub's plans page and cubic's pricing page. Benchmark scores come from Martian's online tracker, last-month view, F1, checked October 4, 2026. For more on Copilot, see our guide to GitHub Copilot code review and cubic vs GitHub Copilot.

If your code lives on GitLab, Bitbucket or Azure DevOps, cubic isn't an option, because it's GitHub only. Copilot's review is in preview for Azure DevOps, and CodeRabbit, Greptile and Qodo all list Bitbucket support. On GitLab, /code-review --comment posts to merge requests, and Claude Code GitLab CI/CD, in beta and maintained by GitLab, runs Claude in your pipelines.

Where a dedicated reviewer adds something

If you already pay for Claude Team or Enterprise, its review is a reasonable default. cubic is our product, so read this part as a pitch; the questions apply to any reviewer.

One agent per rule. Each cubic team rule is a custom agent, a plain-English rule with optional path filters, and every comment names the rule that triggered it. Anthropic's docs advise keeping Claude's single REVIEW.md short, since a long one weakens the rules that matter most.

Learning from your team. cubic learns from your team's feedback, including the past PR comments of senior reviewers you choose, and drops a finding when the author or their coding agent replies that it's wrong.

A page for every PR. Swap github.com for cubic.dev in a PR's URL to review it in cubic, with files grouped by purpose and an AI chat about the change. Comments, approvals and merges sync back to GitHub.

Approval for low-risk PRs. cubic can auto-approve PRs that match a policy you set per repository. If you agree with Anthropic that approval should stay human, Claude's design suits you better.

Codebase scans. A PR review sees the code around each change. cubic's codebase scans, in beta and available by request, send thousands of agents through the whole repository in an isolated sandbox, then scan what changed on a schedule. Pro and Max include scans for up to 3 repositories.

Pricing per developer. cubic charges per developer, with a pooled allowance of reviewed lines. Incremental reviews count only the new changes, so reviewing every push doesn't multiply the bill.

Using Claude Code and cubic together

You don't have to choose. Connect the cubic CLI with cubic auth connect claude-code, pick opus, sonnet or haiku, and cubic review runs through your local Claude Code on your Claude plan, still applying your team's custom agents and learnings. Add cubic's MCP server with claude mcp add --transport http --scope user cubic https://www.cubic.dev/api/mcp, sign in from /mcp, and Claude Code can read and resolve cubic's findings on a PR and pull your team's learnings.

Run both on the same PRs for a week, and keep the reviewer whose comments your team acts on.

Start a 7-day cubic trial, no credit card required.

Table of contents