Veridical17 min readArticle

PR Reviews: Find False Negatives With Evidence, Not Guesswork

PR Reviews: Find False Negatives With Evidence, Not Guesswork

Abstract review path reveals an overlooked defect

The fastest way to reduce false negatives in pull request reviews is to combine inline automated scanning with focused, evidence-seeking human review and a measurable feedback loop. No single layer catches everything on its own. Automation handles mechanical checks at scale, reviewers apply judgment to data flow and business logic, and metrics tell you where the first two are still failing.


TL;DR:

  • Changes across many directories and overloaded reviewers were associated with more missed defects; mutual reviews and longer review time correlated with better detection.
  • Run static analysis, dependency, and secret scans inline on every pull request, then inspect authentication, configuration, and dependency changes with extra care.
  • A 2024 study found one SAST tool warned on only 52% of known vulnerability contributing commits, so a clean scan cannot establish safety.
  • Semantic issues accounted for 51.34% of missed bugs, versus 15.5% for build problems; test input validation, error paths, compatibility, and concurrency accordingly.
  • Track escaped defects by category, risky change coverage, reviewer participation, and time to first review; use an F1 score to compare process changes.

Veridical
Catch More Defects Before Merge
Veridical reviews GitHub pull requests with verified findings, detailed summaries, and inline evidence to help teams spot critical defects.
Explore evidence-based reviews

Table of Contents

Reviewer Behaviors and Team Practices That Reduce Misses

Most missed defects are not technical mysteries. They are process gaps that specific, repeatable habits close.

GitHub’s own review documentation recommends reviewing one file at a time and marking each file as Viewed, which keeps reviewers from losing their place in large diffs and skipping sections under time pressure. Tests belonging to the change should land in the same pull request as the code, not a follow-up, so reviewers can check behavior directly instead of taking the author’s word for it.

Size and workload matter more than most teams admit. A case-control study of the Chromium OS project found that changes spanning many directories were significantly more likely to have security defects slip through, and reviewer workload independently increased the odds of a miss. The same study found that mutual reviews and longer review time correlated with higher detection rates, and that bug-fix commits tagged as such received sharper scrutiny and caught more.

  1. Review one file at a time and mark files Viewed to track real progress through a diff.
  2. Require tests or reproduction steps in the same PR, and ask authors to demonstrate edge cases rather than describe them.
  3. Set a size guideline for pull requests and split any change that touches unrelated areas of the codebase.
  4. Rotate and pair reviewers deliberately, and watch individual review load so no one person is reviewing when fatigued.
  5. Flag bug-fix commits and multi-directory changes for extra review time rather than treating every PR the same.

Workload balance deserves particular attention because it is the easiest pillar to ignore. A reviewer handling ten pull requests a day will skim the eleventh, regardless of skill.

Pro Tip: Assign a second reviewer specifically for any PR that touches more than three directories; the added perspective catches context the primary reviewer missed.

Integrating Automated Tools Without Overwhelming Reviewers

Automated scanning should run on every pull request, but it only works as a false-negative reducer when reviewers treat its output as a starting point rather than a verdict.

OWASP’s guidance on code review and scan integration recommends running static application security testing, software composition analysis, and secret scanning on every PR, with findings surfaced inline so reviewers see them next to the code. The same guidance is explicit that a clean scan marks the beginning of review, not the end: human attention still needs to go to architecture, data flow, authentication, and business logic that static tools cannot reason about.

  • Run SAST, SCA, and secret scanning automatically on every pull request and show findings inline, not in a separate dashboard.
  • Treat a passing scan as a signal to start reviewing context, not proof the change is safe.
  • Triage recurring human-confirmed findings and convert them into custom rules so the same mistake never needs a human to catch it twice.
  • Flag dependency changes, configuration edits, and authentication logic for mandatory extra review, since these categories carry disproportionate risk.

One SAST tool warned in only 52% of known vulnerability-contributing commits in a 2024 empirical study of C/C++ projects, while 22% of those commits produced no warning at all in the changed code. Static analysis narrows the search space; it does not replace the search.

A Review Workflow and Checklist Built Around Real Miss Patterns

A workflow only reduces false negatives if it points reviewers at the categories where defects actually hide. A taxonomy study using the SmartSHARK dataset found that bugs missed before merge were dominated by semantic issues, at 51.34% of cases, followed by build problems at 15.5% and analysis-check misses at 9.09%. That distribution should shape where reviewer attention goes.

  1. Automated pre-scan runs and surfaces findings inline.
  2. Author completes a self-review pass before requesting reviewers.
  3. PR size is checked against the team’s guideline; oversized or tangled changes get split.
  4. Reviewer completes a one-file-at-a-time pass, marking files Viewed.
  5. Reviewer requests evidence, tests, or benchmarks for anything ambiguous.
  6. Reviewer approves, requests changes, or approves with documented notes.

A short checklist keeps the reviewer’s eye on the categories most likely to escape:

  • Data flow and input validation at every boundary the change touches.
  • Error paths and exception handling, not just the happy path.
  • Build and compatibility risk, including anything that touches shared configuration.
  • Concurrency behavior where the change introduces shared state.
  • API or interface compatibility for any caller outside the immediate diff.

New commits pushed after review started should reset the Viewed status on affected files, so nothing merges on the strength of an earlier version of the code.

Metrics That Turn Missed Defects Into Permanent Fixes

Reviewer habits and tooling only improve over time when a team measures what slips through and closes the loop deliberately.

Research on code review coverage and participation links review coverage and reviewer expertise with fewer reported security bugs over time, and recommends tracking escaped defects by category, review coverage on risky changes, and time-to-first-review as core team metrics. An F1-style advisory score, which weighs how many real defects a process catches against how many false alarms it raises, gives teams a single number to compare reviewer and tool configurations against each other.

Metric What it tells you
Escaped defects by category and severity Where the process is still missing real issues after merge
Review coverage on risky changes Whether auth, config, and dependency changes get reviewed at all
Time-to-first-review Whether review delay is pushing authors to merge without full feedback
Reviewer participation Whether review load is concentrated on too few people
Reopened findings Whether fixes are holding or recurring

When the same category of miss repeats, the fix belongs in a static rule, a CI gate, or a test, not in another reminder email.

How an Evidence-Based Review Process Supports These Practices

Our review process addresses the gap the research points to: tools surface candidates, but verification separates real defects from noise. Findings are tied to concrete evidence and reproducible checks, and a calibrated advisory score is calculated from real-defect F1 metrics so teams can assess confidence before merge gating.

We also review repository context beyond the diff itself, including callers, interfaces, and dependencies, which is where the Chromium OS research found multi-directory changes most often slip past reviewers.

A finding is only as useful as the evidence a reviewer can check against it.

Pro Tip: If you are evaluating an evidence-based review tool, start with a narrow pilot on your highest-risk repository and compare its flagged findings against your own escaped-defect log for a few weeks before wiring it into a merge gate.

Techniques That Target False Negatives Directly

Most review practices reduce false negatives as a side effect of improving general review quality. A smaller set of techniques exists specifically to measure whether your detection process is actually working, and mutation testing is the clearest example.

Mutation testing works by introducing small, deliberate faults into the codebase, flipping a conditional, changing a boundary value, and then checking whether the existing test suite catches the change. A test suite that passes against a mutated version of the code is not actually testing that logic, even if coverage reports show the line as exercised. This matters directly for false negatives because high line coverage can hide weak assertions that would never catch a real regression.

Mutation testing exposes surviving defects despite coverage

Running mutation testing periodically, rather than on every PR, gives a team an honest read on which parts of the codebase have tests that only look thorough. Pair the results with the SmartSHARK taxonomy findings on where bugs are missed, semantic issues and build problems, and prioritize mutation testing on the modules most likely to carry that kind of risk.

Prioritizing Review Focus by Risk and Defect History

Not every pull request deserves the same scrutiny, and treating them equally wastes reviewer attention on low-risk changes while high-risk ones get the same cursory pass.

Risk-based prioritization starts with the change type. Authentication logic, dependency updates, and configuration edits carry outsized risk relative to their size, a pattern the OWASP scan integration guidance reflects by recommending these categories get flagged automatically rather than relying on a reviewer to notice. Past defect history adds a second layer: a module that has produced repeated escaped defects deserves a standing rule requiring a second reviewer or extra review time, regardless of how small the current change looks.

The Chromium OS research reinforces this from the opposite direction. It found that large, multi-directory changes were associated with more missed security defects, which means risk assessment should weigh the shape of a change, not just its labeled category. A one-line fix to an authentication module and a fifteen-file refactor touching five directories need different review budgets even if both are technically “small” by line count.

Practical prioritization means building a short list, authentication, payment logic, dependency changes, and any module with a history of escaped defects, and routing those PRs to your most experienced reviewers first, regardless of queue order.

Using Metrics and Analytics to Predict Misses Over Time

Measurement only reduces false negatives when it feeds back into the process rather than sitting in a dashboard nobody checks. The goal is to move from counting what happened to predicting where the next miss is likely.

Tracking escaped defects by category over several months reveals patterns that a single postmortem never would. If analysis-check misses or build problems keep recurring in the same subsystem, that is a signal to add a targeted CI gate rather than hope reviewers catch it next time. Research on review coverage and participation found a measurable relationship between review activity and fewer reported security bugs over time, which supports using participation and coverage trends as leading indicators rather than waiting for the lagging signal of an incident.

An F1-style projected score gives this trend a single comparable number: as recall improves (fewer real defects missed) without precision collapsing (too many false alarms), the score moves in a way a team can track release over release. This is more useful than tracking raw finding counts, which can rise simply because a tool got noisier, not more accurate.

Treat your own historical data as a training signal for prioritization, not just a report card: a subsystem with three escaped defects in two months is telling you exactly where to add tests, add a reviewer, or change the CI gate before the fourth one ships.

Checklists Built for the Scenarios Reviewers Actually Miss

A generic code review checklist tends to ask broad questions: “Is the code clean?”

A checklist tuned to false negative scenarios should map directly to the SmartSHARK taxonomy distribution of missed bugs: semantic issues, build problems, and analysis-check gaps account for the large majority of what slips through. That means the checklist items worth keeping are narrow and specific rather than broad and aspirational.

  • Does every new input path have explicit validation, not just the one the author tested?
  • Does the change compile and pass checks in configurations other than the author’s local environment?
  • Are error paths and exception handling covered by a test, not just the success path?
  • Does any caller outside this PR’s diff depend on the changed interface or return type?
  • Does the change introduce shared state that could behave differently under concurrent access?

Practitioner guidance on evidence-based review recommends asking direct, verifiable questions rather than relying on assertion: “What are the invalid inputs?” or “Show a failing test for the edge case” convert a reviewer’s hunch into something checkable. A checklist that forces this kind of evidence request catches far more than one that just asks whether the code “looks right.”

Training Reviewers to Spot Non-Obvious Defects

Experienced reviewers do not catch more defects because they read faster. They catch more because they know which parts of a diff are statistically more likely to hide a problem, and they ask pointed questions instead of skimming for style issues.

Training that improves detection focuses on pattern recognition over rule memorization. A reviewer who has seen a race condition introduced by a seemingly harmless state change will recognize the shape of that risk again, even in unfamiliar code. The categories worth drilling are the ones the SmartSHARK research identifies as most commonly missed: semantic logic errors, build and compatibility issues, and gaps in analysis checks, since these rarely announce themselves with an obvious syntax problem.

Practitioner guidance on evidence-based review points to a specific habit worth teaching directly: asking for reproduction steps or a failing test converts a reviewer’s suspicion into verified fact, and reviewers who make this a reflex catch edge cases that confident-sounding code review comments usually miss. Pairing a newer reviewer with someone experienced on a few real pull requests, rather than running a one-time training session, builds this instinct faster because it happens on live code with real stakes.

For teams without in-house security expertise, periodic third-party assessments, such as the penetration testing services offered by Solusec, can surface the kind of subtle, exploit-chain defects that internal reviewers are least likely to catch on their own.

Training Reviewers to Spot Non-Obvious Defects — overview diagram

Where AI and Machine Learning Tools Fit In

AI-assisted review tools can meaningfully raise recall, the share of real defects actually caught, when they are built to produce verifiable findings rather than plausible-sounding ones. The risk with many automated tools, AI-based or otherwise, is the same one OWASP documents for static analysis generally: tools can miss complex or context-dependent vulnerabilities while also generating false positives that train reviewers to ignore their output.

The practical fix is to demand evidence alongside any automated finding. A tool that flags a potential issue with a concrete reproduction path or a specific code reference gives a reviewer something to verify quickly. A tool that only offers a confidence label gives the reviewer nothing to check, which is how teams end up either over-trusting or ignoring AI output entirely.

Teams evaluating or building these tools should also track their own truth sets, labeled examples of confirmed and rejected findings, so the tool’s precision and recall can be measured against real outcomes rather than taken on faith. Resources like Brainiac Consulting’s guidance on building evaluation truth sets outline how to structure this kind of ground-truth comparison for AI systems more broadly, and the same discipline applies directly to measuring an AI code review tool’s real-defect F1.

What the Evidence Actually Supports

The research behind this article supports a narrower conclusion than most review guidance admits: the biggest gains come from reducing reviewer cognitive load on mechanical checks, not from asking reviewers to try harder on everything at once. Conventional advice treats “more thorough review” as the fix, but the Chromium OS data shows reviewer workload and change size predict misses better than reviewer effort does.

That means the first priority is not a longer checklist. An F1-style score matters less as a vanity metric and more as a way to notice when a process change actually moved the needle, rather than just felt more rigorous. Teams that measure before and after a process change will know; teams that do not will keep guessing.

— Łukasz

Fewer Missed Defects, Measured Instead of Assumed

We built our review process around the same three pillars this article recommends: automated scanning on every pull request, evidence-backed human review focused on the categories tools miss, and a calibrated score so you can measure whether changes to your process are actually working.

Veridical

Our AI code review for GitHub surfaces verified findings with inline evidence so your reviewers spend time checking, not searching. Open-source maintainers can pilot it through our free program for open-source projects, and teams ready to scale can compare the Launch tier, Standard, and Pro plans against their current review volume.

FAQ

What is the main cause of false negatives in PR reviews?

Large, multi-directory changes and heavy reviewer workload are strongly associated with missed defects, according to a case-control study of the Chromium OS project. Smaller, focused pull requests and balanced reviewer load reduce both risk factors directly.

Can automated tools alone eliminate false negatives?

No single automated tool catches everything. One study found a SAST tool warned in only 52% of known vulnerability-contributing commits, so automated scanning needs to be paired with human review focused on logic and data flow.

What bug categories do reviewers miss most often?

A taxonomy study found missed bugs were dominated by semantic issues at 51.34%, build problems at 15.5%, and analysis-check misses at 9.09%. Checklists and training are most effective when they target these specific categories rather than generic quality questions.

How does mutation testing help reduce false negatives?

Mutation testing introduces small deliberate faults into code and checks whether the existing test suite catches them, revealing tests that pass without actually verifying the logic they claim to cover. It is a direct way to measure whether high coverage numbers reflect real test effectiveness.

What metrics should teams track to improve review accuracy?

Useful metrics include escaped defects by category, review coverage on risky changes, reviewer participation, and time-to-first-review, all linked to better software quality in research on review coverage and participation. An F1-style advisory score calculated from real-defect projections helps compare whether a process change actually improved detection.

Sources