Up to 10x Faster PR Scans: Data Flow Analysis for Engineering Teams
Up to 10x Faster PR Scans: Data Flow Analysis for Engineering Teams

Security-focused data flow analysis traces how untrusted data such as user input, file contents, or API responses moves through a codebase from its source to a sensitive sink, flagging paths where that data reaches a database query, log, or network call without proper handling. In pull-request review, it works as a fast, automated first pass that narrows reviewer attention to likely leaks. It does not replace a human reviewer. With diff-informed and overlay techniques, teams can see scan times drop by up to 10 times on pull requests.
TL;DR:
- Diff-informed and overlay analysis can cut pull request scan times by up to 10 times by reusing cached data and focusing only on changed lines.
- The analysis accurately tracks untrusted data from sources like user input and API responses to sinks such as database queries and logs, but may miss vulnerabilities spanning outside diffs.
- Static data flow tools excel at flagging injection and dangerous API calls but struggle with runtime-dependent or obfuscated code, leading to limited reliability in complex cases.
- A fall back to full scans is necessary for large diffs, renamed files, or truncated outputs, and filtering results by diff range helps reduce false positives in reviews.
- Combining automated inline annotations with human confirmation and evidence-based alerts enhances security review speed without sacrificing accuracy.
Table of Contents
- What data flow (taint) analysis does in security reviews
- How data flow analysis fits into pull-request workflows
- Strengths and common limits of automated data flow analysis
- Concrete steps to implement PR-focused data flow scanning in CI/CD
- How human reviewers and teams should combine manual review with automated data flow results
- Why evidence changes how teams act on data-flow findings
- Put evidence-based review in front of every pull request
- Sources
- FAQ
What data flow (taint) analysis does in security reviews
Taint analysis starts at a source: user input, a file read, an environment variable, or a parameter coming from an external API. It ends at a sink, a location where that data does something consequential.
- Sources include form fields, query parameters, uploaded files, and third-party API responses.
- Sinks include SQL queries, log statements, outbound HTTP calls, and file writes.
- Propagation happens through variable assignments, function arguments and return values, and across files when data is passed through modules or serialized and deserialized between services.
The harder a call chain gets, the harder tracing becomes. Data that passes through several functions, gets wrapped in an object, or crosses a serialization boundary can lose its taint marking if the analysis engine cannot follow the transformation. This article uses data flow analysis in its security sense, the tracking of untrusted data toward dangerous operations, not the compiler-optimization sense used in language design coursework. Those are different disciplines built for different audiences, and conflating them leads reviewers to expect the wrong kind of output from their scanning tools.
How data flow analysis fits into pull-request workflows
Running a full data flow analysis on an entire repository for every pull request is slow and often unnecessary, since most changes touch a small fraction of the codebase. Two techniques make PR-scoped scanning practical.
- Diff-informed analysis restricts the alerts a scan reports to the line ranges actually added or modified in the pull request, using the git diff or the pull request’s API-provided range as input.
- Overlay analysis reuses a cached base database built from the target branch, layering only the changed files on top instead of rebuilding the entire database for each pull request.
Combined, these two approaches can cut pull-request scan times by up to 10x compared to a full rebuild on every run, since the overlay step avoids reanalyzing unchanged code and the diff filter avoids reporting on it.
The scan-time gain hides real edge cases. Overlay databases depend on comparing Git object IDs between the base and the pull request branch to determine which files actually changed, and a mismatch or a truncated diff can cause files to be skipped entirely. Renamed files, PRs that only delete lines, and very large diffs often need a fallback to a full scan rather than an incremental one. More importantly, a code scanning alert typically surfaces in a pull request only when every line in the flagged path falls inside the diff, which means a vulnerability spanning an unchanged file, or a long call chain that starts outside the diff, can go unreported. Filtering results through SARIF location data against the diff range keeps PR annotations stable, but it is also exactly where these gaps appear.
Strengths and common limits of automated data flow analysis
Data flow tools are strongest on well-defined, mechanical patterns: SQL injection, command injection, path traversal, and unsafe deserialization all follow a recognizable source-to-sink shape that taint tracking handles well. They are also reliable at flagging unsafe use of known dangerous APIs, since the pattern does not depend on business context.
- They catch injection-style bugs with clear, traceable paths from input to execution.
- They flag known-dangerous API calls even when a human reviewer might miss the call in a large diff.
- They struggle with anything that depends on runtime state, third-party library internals, or business logic that a static scan cannot see.
False positives come from the same limits. Missing runtime context, unfamiliar third-party components, and code that obfuscates its own data flow through indirection all push a tool toward flagging something that is not actually exploitable. OWASP describes this high false-positive rate as a primary limitation of static and data-flow analysis tools, one that keeps them from replacing a human reviewer on anything involving business logic. Statically typed languages tend to give these tools more to work with, since variable types are known before the code ever runs; dynamically typed languages remove that anchor and push precision down.
Pro Tip: Treat a data-flow alert as a lead for investigation, not a verdict, and confirm the sink is actually reachable with attacker-controlled input before you act on it.

Concrete steps to implement PR-focused data flow scanning in CI/CD
Turning diff-informed and overlay analysis into a working CI pipeline takes a short list of deliberate choices rather than a large engineering effort.
- Compute the diff range from the pull request’s base and head commits, not just the latest push, so the scan sees the full set of changes under review.
- Use a sentinel range to mark files that should be treated as entirely new when a rename or a large rewrite makes line-by-line diffing unreliable.
- Record Git OIDs for both the base and the overlay build so the overlay database can correctly detect which files actually changed.
- Design a cache key around the base branch commit so a stale overlay never gets reused against an outdated snapshot.
- Fall back to a full scan for binary files, unusually large diffs, or truncated diff output rather than silently skipping them.
On the triage side:
- Filter SARIF results by comparing each finding’s file and line locations against the diff range before posting a PR comment.
- Annotate findings inline rather than as a single aggregate report, since reviewers act faster on comments attached to the exact line.
- Require human confirmation on data-flow findings before they block a merge, and reserve automatic gating for the highest-confidence, highest-impact categories.
This matches the direction in CISA’s guidance on secure software development, which recommends automatic scanning on every committed change as part of validating release readiness, paired with more thorough scanning in the central build environment.
How human reviewers and teams should combine manual review with automated data flow results
The most efficient workflow layers automated scanning at three points: a fast local or IDE-level check during development, an automated scan on the pull request itself, and human verification before anything gets merged or blocked.
- Run lightweight checks locally or in the IDE to catch obvious issues before a pull request is even opened.
- Let the PR-scoped scan generate inline annotations tied to the exact changed lines.
- Have a human confirm the finding, especially anything touching a sensitive-data sink or a path where data could leave the system entirely.
Prioritization matters as much as coverage. Sinks that write to logs, external APIs, or storage that leaves the system deserve review before cosmetic or low-impact findings, and a tool that provides a confidence or F1-style score on its findings lets reviewers sort by likely impact instead of reading every alert in the order it was generated. None of this replaces threat modeling. A reviewer who understands the system’s architecture will catch business-logic vulnerabilities, like an authorization check that is technically present but checks the wrong field, that no automated data-flow tool can reason about on its own.
Why evidence changes how teams act on data-flow findings
Most of the friction with automated scanning comes down to trust: a reviewer who has been burned by false positives starts ignoring the tool altogether, which defeats the purpose of running it. Veridical’s approach to pull-request review is built around closing that gap. Every finding comes with verified evidence and an inline explanation tied to the specific lines in the diff, rather than a bare alert asking the reviewer to take the tool’s word for it. Reviews end with a calibrated advisory score projected from real-defect F1 metrics, giving teams a quantitative signal they can route into merge gating instead of relying on gut feel. That maps directly onto the triage workflow above: verified evidence supports faster human confirmation, and the advisory score gives maintainers a defensible basis for deciding which findings justify blocking a merge and which do not.
— Łukasz
Put evidence-based review in front of every pull request
Static scanners tell you something might be wrong. Veridical tells you what is wrong, with the evidence attached, so your team spends its time confirming real defects instead of chasing alerts that turn out to be nothing.

The Standard plan runs $39 per month and the Pro plan runs $69 per month, with a Launch tier available for teams getting started; open-source projects can use Veridical’s free offering at no cost. See how it works on pull requests on GitHub and start your first review today.
Sources
- Get faster CodeQL results on pull requests by analyzing only what changed
- Static Code Analysis | OWASP Foundation
- Securing the software supply chain: recommended practices guide for developers (CISA)
FAQ
What is data flow analysis in the context of code security?
Data flow analysis, in the security sense, tracks how untrusted data moves from a source like user input to a sink like a database query or log statement. It flags paths where that data reaches a sensitive operation without proper validation or sanitization, which is a distinct discipline from compiler-level data flow analysis used in language optimization.
How does data flow analysis differ from control flow analysis?
Data flow analysis follows what happens to a specific piece of data as it moves through a program, while control flow analysis maps the possible execution paths a program can take. In security reviews, data flow answers “where does this input end up,” while control flow answers “which branches can execute.”
Can data flow analysis replace manual code review?
No. OWASP’s guidance on static and data-flow analysis describes these tools as valuable aids with a high false-positive rate that cannot replace human reviewers for business-logic vulnerabilities. Automated scans work best when paired with human verification before a finding blocks a merge.
Why do pull-request scans sometimes miss real vulnerabilities?
A code scanning alert typically only appears in a pull request when every line in the flagged path exists within the PR diff, so a vulnerability that spans an unchanged file or a long call chain starting outside the diff can go unreported. Truncated diffs and renamed files create similar gaps unless the pipeline falls back to a full scan.
How much faster is diff-informed and overlay analysis for pull requests?
Diff-informed and overlay analysis combined can reduce pull-request scan times by up to 10 times compared to rebuilding a full analysis database on every run. The speedup comes from reusing a cached base database and limiting alerts to changed line ranges.