Developers: Code Review F1 Is Misleading, Real PR Ceiling 19.38%
Developers: Code Review F1 Is Misleading, Real PR Ceiling 19.38%

F1 score is the harmonic mean of precision and recall, and benchmarks like SWR-Bench, ContextCRBench, and our own evidence-based scoring at Veridical all lean on it to grade AI-assisted code review. Treat it as a capability diagnostic, not a verdict: F1 ignores true negatives, swings with dataset imbalance, and depends entirely on how a benchmark defines a “match.” Read it alongside confusion matrices, context configuration, and complementary metrics before you trust a number.
TL;DR:
- F1 scores can significantly vary depending on benchmark design choices such as context configuration and matching strictness, affecting their reliability.
- Real-world F1 scores on human-verified pull requests typically range from 15% to 31%, much lower than synthetic benchmark reports often suggest.
- Evaluating models across different diff sizes shows F1 declines sharply with larger diffs, highlighting the importance of testing smaller, focused code segments.
- Complementary metrics like MCC, PR-AUC, and usefulness scores reveal that F1 alone can be misleading about a tool’s real defect detection capabilities.
- Always verify dataset provenance, context settings, matching rules, and diff size breakdowns before trusting any reported F1 score.
Table of Contents
- 1. What F1 measures and how it maps to real review outcomes
- 2. Why F1 alone can mislead, and what to measure instead
- 3. How benchmark design choices quietly change the F1 you see
- 4. Real benchmark F1 ranges and why synthetic scores mislead
- 5. A checklist for vetting F1 claims and running your own evaluation
- 6. How we turn F1 projections into evidence-based review scores
- 7. A short perspective on trustworthy F1 evaluation
- 8. Put evidence-based F1 scoring to work on your repositories
- FAQ
- Sources
1. What F1 measures and how it maps to real review outcomes
F1 is the harmonic mean of precision and recall, calculated as 2 × (precision × recall) / (precision + recall). Precision asks what share of flagged defects were real; recall asks what share of real defects got flagged. The harmonic mean penalizes imbalance between the two more severely than a simple average would, which is exactly why it became the default scalar for defect-detection tasks: a tool that finds everything but drowns you in false alarms, or one that stays quiet and misses half the bugs, both get punished.
In a code review context, the confusion matrix translates into concrete developer experience:
- True positives: comments that correctly identify a real defect, the outcome you are paying for.
- False positives: flagged issues that are not actually problems, the noise that erodes trust in the tool.
- False negatives: real defects the reviewer missed entirely, the failure mode that lets bugs reach production.
- True negatives: clean code correctly left uncommented, a category F1 never counts at all.
That last point matters more than it looks. Most pull requests are mostly correct code. F1 only scores what happened on the flagged and missed items, so it is structurally blind to how well a tool handles the easy, uneventful majority of a codebase.
Matching also gets harder than the formula suggests. A single defect in a function with duplicated logic might warrant a comment at one of three plausible line locations, and two reviewers describing the same root cause in different words can look like a mismatch to an automated scorer. SWR-Bench addresses this by using objective LLM evaluators to match generated findings against ground-truth change-actions rather than relying on raw text similarity, because unstructured natural-language output makes direct string overlap a poor proxy for correctness. How a benchmark resolves multi-location findings and wording variance changes the resulting F1 before any model capability is even measured.
2. Why F1 alone can mislead, and what to measure instead
Code review datasets are heavily imbalanced: most files in most pull requests contain no defects at all, and the interesting class, real bugs, is a small minority. F1 was built for exactly this kind of imbalance, but it still has a blind spot, since it ignores true negatives entirely. Two tools can post identical F1 scores while one is dramatically better at staying quiet on clean code, and you would never see that difference in the headline number.
The statistical risk is not theoretical. A systematic review comparing F1 against the Matthews Correlation Coefficient found that switching the evaluation metric changed which approach looked best in roughly 23% of paired comparisons. That is not noise. It means that in nearly a quarter of cases, a decision made on F1 alone would have favored a different tool than one made on a metric that accounts for all four quadrants of the confusion matrix.
A quarter of paired comparisons flip when you switch from F1 to MCC, according to the systematic review on biased performance metrics, which means a vendor claim built on F1 alone carries a real chance of pointing you at the wrong conclusion.
F1 also says nothing about whether a correctly matched comment was actually useful to the developer who received it. CR-Bench makes this gap explicit: in one reported case, a model scored just 6.30% F1 (27.01% recall, 3.56% precision) yet still achieved 83.63% usefulness and a signal-to-noise ratio of 5.11. A reviewer can be statistically weak on exact-match defect detection while still producing comments developers find worth reading, and the reverse is just as possible.
A few additions close most of the gap:
- Matthews Correlation Coefficient (MCC) uses all four confusion-matrix quadrants and stays stable under class imbalance, unlike F1.
- PR-AUC shows how precision and recall trade off across confidence thresholds instead of collapsing them to one operating point.
- Severity-weighted F1 counts a missed SQL injection very differently from a missed variable-naming nit, which raw F1 treats identically.
- Usefulness and signal-to-noise metrics, as used in CR-Bench, measure whether developers actually act on a finding rather than whether it matched a ground-truth label.
None of these replace F1. They sit next to it, and together they describe a tool’s behavior in a way no single scalar can.
3. How benchmark design choices quietly change the F1 you see
Two benchmarks can evaluate the same model and report wildly different F1 scores, and the gap often comes from design decisions that have nothing to do with the model’s underlying skill. Before comparing F1 numbers across sources, check four things.
Context configuration is the first lever. ContextCRBench built a dataset of 67,910 enriched entries drawn from more than 153,700 crawled issues and pull requests specifically to test this, and found that adding textual context, issue descriptions and surrounding function bodies, improved detection more than adding raw code context alone. A model evaluated on a diff-only slice and the same model evaluated with full repository context are not comparable, even though both results might get reported simply as “F1.”
Matching strictness is the second. Exact-text matching punishes a reviewer for phrasing a correct finding differently than the ground truth, while semantic or location-tolerant matching gives credit for the same underlying defect description. CodeFuse-CR-Bench applies a defect-match threshold rather than requiring exact text overlap specifically to avoid undercounting findings that are substantively correct but worded differently. A benchmark’s matching rule is effectively a dial that can raise or lower every reported F1 without touching the model at all.
Diff size is the third, and probably the most consequential. The comparative evaluation of LLMs for automated code review found F1 falling from roughly 0.657 on diffs under 10 lines to roughly 0.043 on diffs over 150 lines, a collapse that tracks diff size far more tightly than it tracks model choice. Any benchmark that reports a single blended F1 across diff sizes is hiding this curve from you.
Annotation reliability is the fourth check, and it determines whether the ground truth itself deserves trust. SWE-PRBench reports inter-annotator agreement of κ = 0.75 on its LLM-as-judge validation across 350 human-annotated pull requests, a reasonable but not perfect agreement level that puts a ceiling on how much precision any downstream F1 claim can really carry.

Pro Tip: Ask every benchmark or vendor to report F1 broken out by diff-size bucket and context configuration rather than as a single blended number; a blended F1 is nearly always hiding a steep internal curve.
A quick diagnostic checklist for any reported F1:
- Was it computed diff-only, diff plus file, or with full repository context?
- Does the matching rule tolerate semantic and location variance, or require exact text overlap?
- Is the number broken out by diff-size bucket, or blended across all sizes?
- What inter-annotator agreement backs the ground-truth labels?
4. Real benchmark F1 ranges and why synthetic scores mislead
The gap between synthetic and real-world benchmarks is the single most important calibration point for anyone reading an F1 claim. SWR-Bench evaluated 1,000 manually verified GitHub pull requests and found that the best evaluated combination of model and configuration achieved an overall F1 of just 19.38%, a result built entirely on real, messy, human-authored code changes.
The best result on 1,000 manually verified real-world PRs was an F1 of 19.38%, according to SWR-Bench, a figure worth holding in mind every time a vendor quotes a dramatically higher score without specifying the dataset behind it.
That modest ceiling is not an outlier. SWE-PRBench evaluated 350 human-annotated pull requests and found frontier models detecting only 15% to 31% of human-flagged issues, with mean detection around 26%, far below what a human reviewer catches on the same set. The comparative LLM evaluation puts a number directly on the synthetic-versus-real gap: one reported comparison showed a synthetic benchmark F1 of 0.847 against a real-PR F1 of 0.066 on comparable tasks, a drop of roughly 92%.

Task type widens the spread further. Security-defect detection, general logic-bug detection, and logging or PII-leak detection do not share a difficulty profile, and benchmarks that blend them into one headline F1 obscure which category is dragging the average down. A tool that excels at catching missing null checks may be unremarkable at catching a subtle authentication bypass, and a single F1 figure will not tell you which is which.
The diff-size curve from the previous section explains much of this gap. Synthetic benchmarks tend to use short, clean, single-purpose diffs, exactly the regime where the comparative evaluation found F1 near 0.657. Real pull requests routinely mix refactors, feature code, and incidental formatting changes in diffs well over 150 lines, the regime where the same study found F1 collapsing to roughly 0.043.
Three practical calibration points follow:
- Expect real-PR F1 in the high-teens to low-20s as a realistic ceiling for current tools, not the 80%-plus figures synthetic benchmarks sometimes report.
- Expect detection rates well below human reviewer performance, in line with the 15% to 31% range SWE-PRBench reports.
- Expect smaller, more focused diffs to produce meaningfully higher F1 than large, mixed-purpose ones, which argues for evaluating and reviewing in smaller units.
5. A checklist for vetting F1 claims and running your own evaluation
Most vendor F1 claims arrive without the context needed to interpret them. Before you act on one, work through this sequence:
- Verify dataset provenance. Ask whether the benchmark used real, manually verified pull requests, like the 1,000 PRs behind SWR-Bench, or a synthetic or curated set, and whether that dataset resembles your own codebase’s language mix and PR size distribution.
- Request the confusion matrix, not just the F1. Precision and recall separately tell you whether a tool fails by noise or by omission, a distinction the single F1 number erases.
- Check the context configuration. Confirm whether the reported number reflects diff-only evaluation or full repository context, per the configuration effects ContextCRBench documents.
- Check the matching rule. Exact-text matching and semantic or location-tolerant matching, like the threshold approach in CodeFuse-CR-Bench, can move F1 substantially without any change in underlying capability.
- Ask for diff-size buckets. A single blended F1 hides the steep decline the comparative evaluation documented from short to long diffs.
- Recompute MCC where you have the raw counts. It is a better stability check under imbalance and can reveal the kind of discordance the systematic review found in 23% of comparisons.
- Compute severity-weighted F1 for your own risk profile. A missed authentication flaw should not count the same as a missed style nit in your evaluation, whatever the benchmark’s default weighting was.
- Sample actual generated comments for usefulness. Borrow the spirit of CR-Bench’s usefulness and signal-to-noise metrics by having a developer rate a sample of flagged comments as actionable or not.
- Track acceptance and rejection rates operationally. A large empirical study of agentic code review across tens of thousands of review pairs found that roughly one-third of suggestions were accepted, a small share triggered discussion, and the majority were rejected, illustrating real-world utility beyond F1 scores.
- Chunk large PRs before scoring or reviewing them. Given how steeply F1 drops with diff size, per-function or per-hunk evaluation is the single highest-leverage change you can make to both your in-house benchmark and your actual review workflow.
- Keep a human-in-the-loop gate on anything merge-blocking, and track cost per true positive over time as your real measure of return on the tool.
Pro Tip: Run your own 50 to 100 PR sample through any tool you are evaluating before trusting a vendor’s published F1, since diff-size mix and context availability in your repositories will differ from whatever benchmark produced that number.
6. How we turn F1 projections into evidence-based review scores
We built our approach around a simple frustration with the F1 claims circulating in this space: a number without evidence behind it is not a number you can act on. Every finding we surface comes with inline evidence tying the flagged defect to a concrete line, call path, or interface contract, so a calibrated score is never the only thing standing between a reviewer and a merge decision.
Our F1 projections are not computed in isolation from the code around the diff. We review repository context, including callers, interfaces, and dependencies, because the benchmarks above make clear that context configuration is one of the biggest levers on whether a flagged defect is real or a false alarm. A finding that looks suspicious in a diff-only view often resolves cleanly once you see how a function is actually called elsewhere in the repository, and our review process accounts for that before a score is ever generated.
Every review ends with a calibrated advisory score built on real-defect F1 projections, verified findings, and detailed summaries, the same kind of evidence-based layer the checklist above asks you to demand from any tool. That score is designed to be wired directly into merge gating rather than treated as a decorative badge, which means it has to hold up under the same scrutiny we are recommending you apply to any vendor’s number. We detail the engineering choices behind this in how we beat the leaderboard and find more real bugs, and our field notes on code review research track how these benchmarks evolve.
Teams can apply the checklist from the previous section directly against our own outputs: ask for the confusion matrix behind a score, check whether a finding is backed by evidence you can independently verify, and sample flagged comments for usefulness before trusting the number at a glance.
7. A short perspective on trustworthy F1 evaluation
The biggest risk in this field right now is not that F1 is the wrong metric. It is that most reported F1s are not comparable to each other, and the gap between a synthetic benchmark and a real pull request can be the difference between a 0.847 and a 0.066 on the same underlying task, as the comparative LLM evaluation shows. Research priorities should shift toward richer context representation, semantic rather than exact-text matching, and severity weighting as defaults rather than afterthoughts.
For engineering teams, the operational answer is less about chasing a better metric and more about changing how you consume any metric at all. Chunk large pull requests before you evaluate or review them. Keep a human gate on anything that blocks a merge. Track acceptance rate over time as your real signal of whether a tool is earning its place in the workflow, not just its headline F1.
The synthetic benchmarks that produce the most impressive-looking numbers are also the ones least representative of what you will actually ship. Weight your trust accordingly.
— Łukasz
8. Put evidence-based F1 scoring to work on your repositories
We think the lesson from every benchmark cited above points in the same direction: a single F1 number is never enough, and the tools worth trusting are the ones that show their evidence alongside the score. That is the gap we built to close. Every review we run produces a calibrated advisory score grounded in real-defect F1 projections, backed by inline evidence and repository context, so you are never asked to take a number on faith.

If you want to see how that plays out on your own codebase, our AI code review for GitHub page walks through how verified findings and advisory scores work in practice, and open-source projects can run it at no cost through our free tier for open source. Teams ready to wire evidence-based scoring into merge gating can check plan details, including Standard at $39 per month and Pro at $69 per month, on our pricing page.
FAQ
What is a good F1 score for AI code review?
There is no single universal threshold, because F1 depends heavily on dataset, context configuration, and matching rules. On real, manually verified pull requests, SWR-Bench reports a best result of 19.38% F1, which is a more realistic benchmark for production expectations than the much higher scores synthetic benchmarks sometimes report.
Why does F1 drop so much on larger pull requests?
Larger diffs mix more unrelated changes, dilute reviewer attention, and make exact matching harder, and the effect is measurable: the comparative LLM evaluation found F1 falling from about 0.657 on diffs under 10 lines to about 0.043 on diffs over 150 lines. Chunking large pull requests into smaller units before review is the most direct way to counter this.
Is MCC better than F1 for code review evaluation?
MCC accounts for true negatives and stays more stable under class imbalance, which F1 ignores entirely. A systematic review found that switching from F1 to MCC changed the favored approach in about 23% of paired comparisons, so the two metrics are worth computing side by side rather than treating F1 as sufficient on its own.
Can a tool have low F1 but still be useful to developers?
Yes. CR-Bench documented a case where a model scored just 6.30% F1 yet reached 83.63% usefulness and a signal-to-noise ratio of 5.11, showing that exact-match defect detection and developer-perceived usefulness can diverge sharply.
Does Veridical publish an F1 score for its reviews?
We generate a calibrated advisory score from real-defect F1 projections on every review, tied to verified findings and inline evidence rather than presented as a standalone figure. Pricing and plan details, including the Standard and Pro tiers, are listed on our pricing page.
Sources
- SWR-Bench (arXiv)
- CR-Bench (arXiv)
- ContextCRBench (ACM / arXiv)
- Bigger isn’t always better: comparative evaluation of LLMs for automated code review (arXiv)