Blog
Automated code review: what to automate, what to keep human
Automated code review from a team that builds an AI reviewer: what each layer catches, what to keep human, and 4 numbers that show whether it works.
Alex Mercer
Automated code review is any check that inspects a code change and reports problems without a person doing the reading: formatters and linters, static analysis and security scanners, tests that must pass in continuous integration (CI), and AI reviewers that comment on pull requests. We build cubic (cubic.dev), one of those AI reviewers, and we still wouldn't automate all of it.
Automate every check a machine can do reliably. Those checks run on each change before a teammate opens it, and they don't get tired on the last PR of the day. They can't decide whether the change was a good idea, so keep that part human. Review time is too expensive to spend on semicolons.
Automated code review tools, layer by layer
Automated review comes in four layers, from the most certain to the most judgment-heavy.
Layer | What it catches | What it misses | Examples |
|---|---|---|---|
Formatters and linters | Style, likely bugs, banned patterns | Anything that depends on what the code is for | Prettier, Black, gofmt, ESLint, Ruff |
Static analysis and SAST | Known vulnerability patterns, unsafe data flows, vulnerable dependencies, leaked secrets | Logic specific to your business | SonarQube, CodeQL, Semgrep, Dependabot alerts, GitHub secret scanning |
Tests and CI gates | Regressions someone wrote a test for | Whatever nobody thought to test | Your test suite, required status checks |
AI review | Logic errors, missed edge cases, breakage across files, violations of rules you wrote in plain English | Context it wasn't given; it can be wrong | Claude Code Review, CodeRabbit, cubic, Cursor Bugbot, GitHub Copilot code review, Greptile |
cubic is our product, so weigh what we say about the AI layer accordingly. For more options in each category, see our comparison of code review tools.
Formatters and linters
Formatters such as Prettier, Black and gofmt rewrite code into one consistent style, so nobody argues about style in review; Black's documentation adds that smaller diffs make review faster. Linters such as ESLint for JavaScript and Ruff for Python flag likely mistakes and patterns your team has banned, and eslint --fix corrects many of them for you.
Run both in the editor, in a pre-commit hook and in CI as a required check. The findings are deterministic and the fixes mechanical, so block merges on this layer from day one. It can't see purpose: a linter can tell you a variable is unused, but not that the function shouldn't exist.
Static analysis and security scanning
Static analysis goes deeper than linting: it follows how data moves through the program, so it can catch user input flowing into a SQL query. The security-focused kind is SAST (static application security testing); CodeQL, GitHub's engine for automated security checks, and Semgrep are examples. SonarQube adds a quality gate, pass-or-fail conditions that can fail a pipeline or block a pull request. Dependency and secret scanning belong here too: Dependabot alerts flag packages with known vulnerabilities, and GitHub's push protection blocks pushes that contain secrets.
A scanner knows what SQL injection looks like, but not that only a workspace admin may export billing data.
Point one at an old codebase and it reports more findings than anyone will fix this quarter, and your team learns to ignore a check that's always red. Gate on new code only, the way SonarQube's default gate, Sonar way, does.
Tests and CI gates
Tests are the layer that runs the code; the others mostly read it. CI runs the suite on every pull request, and GitHub's branch protection can require those status checks to pass before merging.
Tests can pass for the wrong reason. Our post Why AI code fails differently describes an AI-written payment change whose tests passed because a mocked analytics service answered instantly; under real production load, payments timed out. If an agent writes both the code and its tests, a green build only proves the two agree, and someone still has to check the tests against the intent.
Can AI do code reviews?
Yes, for part of the job. In Atlassian's study of its own AI reviewer (ICSE 2026), 38.70% of the AI's comments led to a code change in the next commit, against 44.45% for comments written by people. AI review reads a change roughly the way a colleague would. It can't tell you whether the change should exist.
The newest layer runs a large language model over the diff plus context from the rest of the repository, so it can flag an inverted condition, a caller in another file that the change broke, or a team rule the code ignores. Most linters cover one programming language; GitHub's docs say Copilot's reviewer handles any language.
The team rules matter most. The AI-written changes that hurt look clean and break a rule that only lives in your team's heads; that was the main point of Why AI code fails differently. Write those rules down in plain English and let the reviewer enforce them on every PR.
It misses context you didn't give it: why the change exists, unless the PR or ticket says so, and your product strategy. It's also probabilistic, so it will miss real bugs and flag some that aren't. Treat its comments as a strong first pass that a person still judges.
Which AI code review tool is the best?
It depends on the test, so distrust any ranking without a date and method, ours included. On Martian's Code Review Bench online tracker (last-month view, checked October 4, 2026), cubic, our product, ranks first at 65.3% F1. On its offline benchmark of 50 hard bugs (as labeled Sep 8, 2026), Qodo's Deep configuration leads cubic by 0.2 F2 points.
Martian scores reviewers on open-source pull requests: precision is the share of a reviewer's comments developers acted on, recall is the share of real fixes it caught, and F1 balances the two. The board moves daily, and several vendors have each reported #1 at different dates and in different modes.
Start with your platform. All six AI reviewers in the table work on GitHub, and cubic works only there. On GitLab, Bitbucket or Azure DevOps, CodeRabbit and Qodo support all three. Claude Code Review, Anthropic's managed reviewer, is a research preview on Team and Enterprise plans. Then build a shortlist from benchmarks, run it on your own pull requests, and keep the reviewer whose comments your team acts on.
AI vs human code review: what to keep human
Our rule of thumb: if a human review comment could have been a lint rule, a test or a written rule for the AI reviewer, turn it into one. To find the gaps, read a week of your team's review comments; anything about formatting or a missing null check belongs to a machine. The rest belongs to people: intent, design, product trade-offs and understanding.
Intent. Should this change exist, and does it solve the problem it was opened for? An AI reviewer can check the code against the ticket, but only someone who knows the customer and the history can tell you the ticket asks for the wrong thing.
Design. Google's code review guide calls the overall design "the most important thing to cover in a review." Fit depends on where the system is going, and that lives in people's heads and roadmaps, not in the diff.
Product trade-offs. Serve slightly stale data or pay for a slower query; ship behind a flag this week or wait for the full version. An AI reviewer can point out the trade-off; the people who answer for the product pick the side.
Understanding. Someone besides the author should understand the change well enough to debug it during an incident, which matters more as agents write more of the code. Our post The real problem with AI coding calls the alternative comprehension debt: code that works and ships, and that nobody understood when it was written. With technical debt, at least someone did. Review is the last cheap place to catch it.
For the line-level questions a reviewer should still ask, use our code review checklist.
How to roll out automated code review without the noise
Turn everything on at once and developers get a wall of comments on their first PR, learn to scroll past them, and keep scrolling when one is a real bug. Add layers in order, and let each one prove itself before it can block anything.
Block on the deterministic checks first. Formatter, linter, build and tests. Reformat the codebase once, in a PR that does nothing else, as Google's guide also asks.
Add static analysis and security scanning, gated on new code. Block on leaked secrets and high-severity findings in the lines the PR changes. Report the rest without blocking, and work through the backlog separately.
Make the blocking checks required. On GitHub, add them as required status checks in branch protection or a ruleset, and require code owner review on paths like auth, billing and migrations.
Turn on AI review as comments only. Most tools start there: Copilot leaves a Comment review that doesn't count toward required approvals, Claude Code Review never approves or blocks, and Bugbot's check stays neutral unless you turn on failing for unresolved issues. Our guide to AI pull request review on GitHub covers setup.
Tune it for a few weeks. Exclude generated files, lockfiles and vendored code. Tell the AI reviewer to skip what your linter enforces, and turn off categories nobody acts on. Write your team's rules where the reviewer reads them: custom agents in cubic,
.github/copilot-instructions.mdfor Copilot,.cursor/BUGBOT.mdfor Bugbot,REVIEW.mdfor Claude Code Review. Start with one rule per recent incident.Give it a vote only after it earns one. Once developers act on its comments at least as often as on your human reviewers' comments, consider blocking on its highest-severity findings or auto-approving narrow classes of low-risk PRs. Where approvals exist, they're opt-in: Copilot approvals are in public preview, Greptile's auto-approve is a beta, and cubic's auto-approval is off by default, with a shadow mode that stays comment-only and shows which PRs it would have approved.
How to tell if automated code review is working
Record a baseline before you turn anything on, change one layer at a time, and track four numbers: acted-on rate, time to first review, time to merge and escaped bugs.
Acted-on rate
The share of automated comments that lead to a code change before merge. Count it per tool and per category, since one noisy category can hide inside a decent average. Sample recently merged PRs and check, for each automated comment, whether the lines it pointed at changed in a later commit.
Count your human reviewers' comments the same way, because the realistic bar isn't 100%: in Atlassian's study, people's comments led to a change less than half the time.
Time to first review
The time a ready PR waits for its first human review. Automated feedback should land first, so authors fix the mechanical problems before anyone else looks. If human pickup slows after a rollout, check whether reviewers are waiting for the bots to finish or for authors to clear the bots' comments.
Time to merge
From opened to merged. Use the median, since a few stuck PRs distort an average.
In a 2024 study at Beko, PR authors labeled 73.8% of an LLM reviewer's comments as resolved, yet average PR closure time rose from 5 hours 52 minutes to 8 hours 20 minutes, though one of the three projects got faster. In Atlassian's study, median PR cycle time fell 30.8% and human comments per PR fell 35.6%. The same kind of tool pushed this number in opposite directions, so track your own.
The GitHub CLI gives you the timestamps for both time measures. List your bots' logins in BOTS so their reviews don't count as the first human one:
Each row is a PR number, when it was opened, its first human review and when it merged. Draft PRs will look slower to review than they were.
Escaped bugs
Bugs found after merge, in staging, in production or by a customer, that trace back to a PR that passed review. Label them in your issue tracker when you find the cause, and count them per hundred merged PRs each month. It's the slowest signal and the one that matters most, so compare quarters, not weeks.
Common failure modes and how to fix them
Noise. One developer in the Beko study complained that every fix set off a new review whose comments were often redundant. Review only what changed since the last review, raise the severity bar, and cut categories with a low acted-on rate.
False positives. A wrong comment costs the author time and the tool trust, and trust is what gets the next correct comment read. Track which rules get dismissed, and rewrite or delete the ones that are wrong more often than right. Prefer a tool that learns from your team's replies and reactions; cubic, Bugbot, CodeRabbit and Greptile each document a version of this.
Rubber-stamping. Reviewers start treating green checks as the review. A respondent in the Beko survey worried that reviewers would skip problems, assuming the bot would have flagged them. On GitHub, dismiss stale approvals when new commits change the diff, and require the most recent push to be approved by someone other than the person who pushed it.
Approving AI code nobody understood. The newest failure, and the one we worry about most: an agent writes the code, an AI reviewer passes it, and a person approves a diff they skimmed. Every check is green and nobody can explain the change. Ask authors to explain it in their own words in the PR description, and treat "I don't know why the agent did this" as a reason not to merge.
Automate what can be checked. Keep people on what has to be judged.
If you're comparing AI code review tools for that fourth layer, cubic reviews pull requests on GitHub. It's free for public repositories, and paid plans start with a 7-day trial, no credit card needed.
