Cut 94–98% of GitHub PR False Positives with Evidence First LLMs
Cut 94–98% of GitHub PR False Positives with Evidence First LLMs

To cut false positives in GitHub PR checks, combine documented triage (dismissals plus an audit trail), prioritized filtering, and evidence driven automated elimination through hybrid LLM and static analysis methods, while tracking F1 and recall so you never trade noise reduction for missed bugs. Done well, this combination can remove most of the noise your team currently tolerates. Evidence based review tools like Veridical build on the same principle: verify before you suppress.
TL;DR:
- Prioritize discarding false positives only after auditing dismissal reasons and attaching structured explanations for effective future automation.
- Hybrid LLM and static analysis methods can eliminate up to 98% of false positives if supplied with comprehensive code context.
- Filtering alerts by severity and focusing only on current changes significantly reduces noise and review time.
- Implement layered automated validation techniques, including input-agnostic dynamic validation and symbolic execution, for a balanced trade-off between precision and coverage.
- Track key metrics such as false positive rate and inspection time to maintain high recall and avoid missing genuine bugs during noise reduction efforts.
Table of Contents
- Root causes of false positives in static analysis and PR checks
- Practical, day-one steps to cut false positives in GitHub PR reviews
- Higher-investment automated techniques: LLM hybrids, AFPE, and dynamic validation
- Wiring triage, suppression, and verification into GitHub PR pipelines
- What to measure so you cut noise without missing real bugs
- A straight answer on what actually works
- How Veridical applies this guide in practice
- Key papers, industrial studies, and documentation
- FAQ
- Sources
Root causes of false positives in static analysis and PR checks
Most false positives trace back to a small set of predictable failure modes. Static analyzers are built to be conservative: when a tool cannot prove a path is safe, it often flags it anyway, a design choice known as over-approximation. Because the general problem of verifying arbitrary program behavior is undecidable, no configuration removes every false alarm permanently, as the Oxford research on false-positive trade-offs points out.
The practical consequences show up in a few recurring patterns:
- Infeasible execution paths get flagged because the analyzer lacks the context to rule them out.
- Missing information about sanitizers, validation logic, or test harnesses causes the tool to treat safe code as risky.
- Configuration gaps and unsupported library signatures generate alarms that have nothing to do with real defects.
- Test-only code paths get scanned with production rules, producing alerts nobody will ever act on.
The scale of the problem is well documented. Static analysis tools commonly report false positive rates between roughly 35% and 91%, with around 40 alarms per 1,000 lines of code in some studies. Each one of those alarms costs a reviewer time to triage, and that cost compounds across every pull request in a busy repository.
Practical, day-one steps to cut false positives in GitHub PR reviews
You do not need new infrastructure to start reducing noise. The fastest wins come from disciplined triage and a few configuration changes you can ship this week.
- Use the GitHub code scanning API’s
dismissed_reasonfield for every dismissal, paired with a short comment explaining the evidence behind the call. The REST API documentation supports structured reasons such as “false positive,” “used in tests,” and “mitigated,” which turns one-off judgment calls into reusable signal. - Gate alerts by severity and filter by delta or blast radius so reviewers only see findings tied to the current change, not historical noise reintroduced by a rebase.
- Standardize triage rules so test-only alerts dim automatically and known library patterns get suppressed with a documented rationale attached, not a silent override.
- Automate repetitive dismissals into suppression rules only after you have audited their history. A pattern dismissed ten times with the same stated reason is a strong candidate for automation; a pattern dismissed once is not.
- Run a short pilot: pick one repository, measure your baseline false positive rate and triage time, then iterate weekly.
Pro Tip: Treat every dismissal comment as a training label, not a throwaway note. Six months from now, that history is what tells you which suppression rules are safe to automate.
A related discipline worth borrowing from incident response is designing alert thresholds so they do not page people on noise they will ignore. The same filtering logic that keeps an on-call rotation sane, as described in this guide to designing alert rules, applies almost directly to PR check configuration.
Higher-investment automated techniques: LLM hybrids, AFPE, and dynamic validation
Once triage discipline is in place, the next gains come from automated elimination methods, each with a different cost and risk profile.
- LLM hybrids (approaches in the literature labeled LLM4SA and LLM4PFA) combine a static analyzer’s facts with a language model’s contextual judgment. An industrial empirical study found hybrid methods eliminating 94 to 98% of false positives while preserving high recall, though results depend heavily on supplying the model with caller, callee, and surrounding code context rather than the diff alone.
- AFPE methods, meaning automated false-positive elimination through symbolic execution, SMT solving, or model checking, deliver high precision but run into scaling limits on large codebases. The ACM Computing Surveys taxonomy recommends applying AFPE after a pruning step narrows the alarm set to something tractable.
- Rican-style input-agnostic dynamic validation uses existing test coverage plus targeted fuzzing to soundly confirm or eliminate alarms. In evaluation, Rican eliminated roughly 45% of certain pointer-related false positives without removing a single real alarm, making it a strong fit for codebases with mature test suites.
A very large share of false positives were eliminated by hybrid LLM and static-analysis methods in an industrial study, a reduction large enough to change how much manual triage a team needs per release cycle.
None of these techniques work best alone. Pruning to shrink the candidate set, then applying AFPE or an LLM pass, then verifying with dynamic validation where tests exist, gives you a layered pipeline that balances cost against accuracy better than any single method on its own.
Wiring triage, suppression, and verification into GitHub PR pipelines
The mechanics matter as much as the methods. GitHub’s code scanning API stores dismissed_reason and dismissed_comment on each alert through its create_request flow, which means your dismissal history lives in a queryable, auditable place rather than scattered across PR comment threads.
- Store every dismissal with its reason and comment through the API so the record survives beyond the reviewer who made the call.
- Gate merges on severity tiers and keep a separate, visible bucket for unresolved low-risk alerts instead of forcing a binary pass or fail.
- Feed accumulated dismissal records back into suppression rules or model fine-tuning, with audit logs intact, so the system improves rather than drifting.
- Assign an owner for the triage ruleset, set a review cadence, and define a rollback plan before anything runs unattended.
Pro Tip: Keep the unresolved low-risk bucket visible on the PR, not hidden in a separate dashboard. Alerts that disappear from view get forgotten, not fixed.
What to measure so you cut noise without missing real bugs
Noise reduction only counts if recall holds. Track false positive rate, inspection time per alert, reviewer throughput, and a projected F1 score against a sample of confirmed defects.
| Metric | What it tells you | How to collect it |
|---|---|---|
| False positive rate | Share of flagged issues that are not real defects | Sample dismissed alerts and confirm manually |
| Inspection time per alert | Reviewer minutes spent per triage decision | Time-stamp triage events in your tracker |
| Reviewer throughput | Alerts cleared per reviewer per week | Count resolved and dismissed alerts over time |
| F1 projection | Balance of precision and recall on a gold set | Score tool output against confirmed defects |
Build or sample a gold set of confirmed defects to check recall before and after any change, run staged rollouts rather than flipping a switch repository-wide, and keep that safety bucket for unresolved high-risk alerts until confidence builds. Report the four metrics above on a regular cadence so the team can see whether noise is actually dropping or just moving somewhere less visible.
A straight answer on what actually works
Reducing false positives is less about finding one clever tool and more about building a pipeline where every suppression decision leaves a trail. Favor dismissals you can audit and reverse over opaque filters that quietly remove signal nobody questioned. Start with triage discipline, let the history validate itself, then automate. Hybrid LLM methods are genuinely promising, but they still earn spot checks, not blind trust. The real warning signs are over-suppression, missing audit trails, and a team that stops measuring recall because the dashboard finally looks quiet.
— Łukasz
How Veridical applies this guide in practice
Veridical builds pull request reviews around the same principle this guide argues for: nothing gets flagged without evidence, and every finding carries a calibrated advisory score based on real-defect F1 projections rather than a raw alert count. 

Reviews look beyond the diff into repository context, callers, interfaces, and dependencies, which is the same context richness that makes hybrid LLM methods effective rather than noisy. Teams can start with the free tier for open-source projects or move straight to a paid plan: Standard runs $39 per month and Pro runs $69 per month, with the Launch tier priced on request. If you want to see how evidence-based PR review handles a real codebase, visit Veridical to start a pilot.
Key papers, industrial studies, and documentation
- Static analysis false positive rates, LLM hybrid industry study, ACM postprocessing survey, Rican dynamic validation paper, and GitHub code scanning API docs.
- Further reading on applied examples: Veridical’s research blog.
FAQ
What is the main cause of false positives in static analysis?
Over-approximation is the main driver: analyzers flag code as unsafe whenever they cannot prove a path is safe, which produces alarms on correct code. Missing context about sanitizers, test harnesses, and configuration adds further noise, as described in the ACM postprocessing survey.
How much can hybrid LLM methods reduce false positives?
An industrial empirical study found hybrid LLM and static-analysis methods eliminating 94 to 98% of false positives while maintaining high recall. Results depend on supplying the model with caller and callee context rather than just the code diff.
What should I document when dismissing a GitHub code scanning alert?
Record the dismissal reason using GitHub’s structured dismissed_reason field, plus a short comment describing the evidence behind the decision. The GitHub code scanning API stores this history so it can later inform suppression rules or model training.
Can automated tools eliminate false positives without missing real bugs?
Input-agnostic dynamic validation, as shown in the Rican evaluation, eliminated roughly 45% of certain pointer-related false positives without removing any real alarms when test coverage was sufficient. No method removes every false positive permanently, since program verification remains undecidable in the general case.
Does Veridical replace manual PR triage entirely?
Veridical reduces manual triage by attaching verified evidence and an F1-based advisory score to each finding rather than eliminating human review altogether. Reviewers still confirm findings before merge, which keeps the audit trail intact and recall measurable.
Sources
- Static analysis false positive rates (arXiv)
- Reducing false positives in static bug detection with LLMs: An empirical industry study (arXiv)
- Survey of approaches for postprocessing static analysis alarms (ACM Computing Surveys)
- Bridging coverage and confidence: reliable static false alarm elimination via input-agnosticity (Proceedings of the ACM on Programming Languages)
- REST API endpoints for code scanning (GitHub Docs)