Blog

Ultrareviews find 4.7x more critical bugs in cubic reviews

We compared ultrareview, a longer multi-pass version of cubic's AI code review, with the standard review on production data.

Victor Mier

In April, we launched ultrareview, a longer, multi-pass version of cubic's AI code review that runs on our most capable models.

Teams now run more than 11,000 ultrareviews a month.

To see what the extra depth finds, we compared ultrareviews with standard reviews on production data.

What we found:

  • For the same amount of code, ultrareview finds 1.7x as many real bugs as a standard review, and 4.7x as many critical ones. The more serious the bug, the bigger the gap.

  • The biggest gap by type of bug is security and privacy, where ultrareview finds 2.3x as many bugs that developers went on to fix.

  • The 15% of PRs that cubic marks as high-risk have many more serious bugs than the rest.

  • Half of ultrareviews finish within 9 minutes, about 5 minutes longer than a standard review. The extra time goes to complicated files.

Findings by severity and category

Per 1,000 lines of code, ultrareview produces more valid findings than a standard review at every severity level, and the ratio grows with severity.

Line chart of the ultrareview vs standard ratio of valid findings per 1,000 lines of code, by severity: 1.68x for all findings, 2.61x for mid and above, 3.56x for high and above, 4.71x for critical.

By category, counting only findings that developers fixed, the gap is largest for security and privacy:

Bar chart of where ultrareview finds more, by category, in addressed findings per 1,000 reviewed lines: security and privacy 2.29x, architecture and design 1.93x, code quality 1.93x, performance 1.91x, business logic 1.84x, down to accessibility and UX 1.38x. All categories together: 1.69x.


Automatic ultrareview for high-risk PRs

With ultrareview set to Automatic and the High-risk PRs policy selected, cubic checks each PR when it's opened or marked ready for review. If the change is high risk, like auth, payments, or a database migration, the first full review runs as an ultrareview and cubic posts the reason on the PR.

cubic flags 15% of PRs as high-risk, and later commits on those PRs get standard reviews.

High-risk PRs have many more P0 and P1 bugs than other PRs, and ultrareview finds more of them.

Bar chart of valid P0/P1 findings per review, by PR risk: on PRs that are not high risk, standard review finds 0.16 and ultrareview 0.75; on high-risk PRs, standard review finds 0.84 and ultrareview 2.02.

Where the extra time goes

Ultrareviews take a median of 9 minutes, about 5 minutes longer than a standard review, and 90% finish within 18 minutes.

We let ultrareview take longer to check each possible bug, and it spends that time on complicated files, where it takes 2.3x as long per file as a standard review. On simple files, the two take about the same time.

Bar chart of median review time per file, by complexity: low complexity 27s standard vs 25s ultrareview, medium 43s vs 82s, high 66s vs 149s.

Methodology and limitations

  • We only count findings that developers fixed in a later commit on the same PR.

  • Rates are per 1,000 reviewed lines of code, which adjusts for PR size.

  • Ultrareviews don't run on a random sample of PRs. Teams start them by hand, or cubic starts them on PRs it flags as high-risk. Not-high-risk PRs that got an ultrareview are about twice as large as a typical not-high-risk PR.

Availability

Ultrareview is available on paid and trial plans. Each ultrareview uses 3x the quota of a standard review. To run one, comment @cubic-dev-ai ultrareview on a pull request, or start it from the PR page in cubic. Automatic ultrareview is off by default. You can turn it on in AI review settings.







Table of contents