Veridical16 min readArticle

5 Code Review Metrics Managers Must Track to Predict DORA Delivery

5 Code Review Metrics Managers Must Track to Predict DORA Delivery

Isometric code review metrics title card

Track five metrics first: time to first response (TTFR), time to approval or merge (TTA/TTM), defect escape rate, review thoroughness (quality ratio), and reviewer load. Measure each as P50 and P90, not just an average, and pair speed metrics with escape rate so you can tell whether faster reviews are catching problems or missing them. Instrument these five before adding anything else. Done correctly, this combination maps to the delivery outcomes DORA has spent years validating the linked delivery outcomes (validating).


TL;DR:

  • Measure lead time for changes alongside turnaround metrics to ensure faster reviews do not lead to higher defect escape rates.
  • Compare P50 and P90 values for review speed with defect escape rates to detect rubber-stamping or missed issues effectively.
  • Use composite health scores combining quality ratio, turnaround, and escape rate to avoid optimizing single metrics at the expense of overall quality.
  • Set SLA targets based on PR priority and team context, recognizing that larger teams and mature codebases naturally have longer review cycles.
  • Regularly audit automated metrics, especially those influenced by AI, to prevent gaming and verify that speed improvements do not compromise review thoroughness.

Veridical
Make Review Quality Measurable
Veridical provides evidence-based pull request reviews, verified findings, detailed summaries, and transparent scoring for GitHub repositories.
Explore Veridical

Table of Contents

Checklist of essential code review metrics

Before building dashboards or setting service-level targets, decide which signals earn a place on the board. Most teams track too many metrics and act on none of them.

  • Turnaround: time to first response, time to approval, time to merge, each split by P50 and P90.
  • Quality: defect escape rate, count of blocking findings per pull request, and a quality ratio of substantive comments to total comments.
  • Throughput and size: pull requests reviewed per period and median PR size in lines changed.
  • Load balancing: reviews completed per person, current queue depth, and a Gini coefficient to flag uneven distribution.
  • Delivery linkage: which DORA metrics, particularly lead time for changes, move when review speed changes.

Turnaround tells you if reviews are a bottleneck. Quality tells you if speed is coming at the cost of catching real defects. Load balancing tells you if the same two senior engineers are quietly absorbing every hard review while the rest of the team clears easy ones. Each category needs at least one number before you touch a dashboard tool.

Exact definitions and measurement recipes with suggested thresholds and percentiles

Vague definitions produce numbers nobody trusts. Use these recipes so your metrics are comparable across repos and over time.

  1. Time to first response (TTFR): the interval from PR creation (or “ready for review” if drafts are used) to the first substantive comment or approval, excluding bot comments.
  2. Time to approval (TTA): from ready-for-review to the first approving review, even if further changes follow.
  3. Time to merge (TTM): from ready-for-review to merge, capturing the full review-and-rework cycle.
  4. Defect escape rate: post-merge defects (bugs filed or incidents traced to a merged PR) divided by total merged PRs in the same window, linked by commit SHA or PR ID.
  5. Quality ratio: substantive comments (those requesting a change or flagging a defect) divided by total comments, excluding style nits handled by linters.

Report each turnaround metric as P50 and P90, and add P95 for high-severity paths. A P50 of four hours with a P90 of three days tells you the tail is where your bottleneck actually lives. Retain raw data for at least one quarter and recompute weekly.

A single automated review-evaluation metric is not enough on its own: a robustness study of 29 peer-review evaluation metrics found that most metrics are sensitive to meaning-preserving rewrites and only a few hold up under that stress test, which is why escape rate has to sit alongside anything automated.

Set SLA expectations by priority: a P0 security fix might carry a four-hour TTFR target, while a P3 documentation change can tolerate two business days. Source your defect data from your incident tracker or bug tracker, joined to PR IDs through commit metadata, not from memory or anecdote.

How to read metrics together and build composite health signals

No single metric tells the full story, and reading them in isolation is how teams convince themselves that faster is automatically better.

  • Pair TTFR with defect escape rate: rising speed alongside rising escapes means reviews are rubber-stamping, not reviewing.
  • Combine quality ratio, TTFR, and escape rate into a composite health score rather than ranking teams on any single input.
  • Run small experiments (one team, one sprint) before changing SLAs organization-wide, so you isolate cause from noise.
  • Check whether review-speed changes track with lead time for changes, the DORA metric most directly downstream of review practice.

Pro Tip: Before declaring a review process “faster,” confirm escape rate held flat or dropped. Speed without that check is a vanity number.

DORA’s own data shows the relationship is not automatic. Faster code review has been linked to delivery performance about 50% higher than slower-reviewing teams, but the 2024 DORA dataset also found that AI adoption produced a small increase in review speed and approval speed, alongside a modest decrease in code complexity with possible tradeoffs in stability. Small gains in speed do not guarantee proportional gains in delivery performance, which is exactly why the composite view matters more than any single number on its own.

Dashboard panels and rollout plan for team and org levels

A dashboard earns its place only if it changes what someone does on Monday morning.

  • Time series with P90 markers for TTFR, TTA, and TTM, refreshed weekly, so trend direction is visible before it becomes a crisis.
  • Heatmap of reviewer load by person and week, making an uneven Gini coefficient visible at a glance.
  • SLA compliance panel by PR priority (P0 through P3), showing percent within target rather than a raw average.
  • Escape-rate trend linked to PR IDs, so a spike is traceable to specific merges, not just a quarter-level number.

Keep team dashboards granular (per-repo, per-sprint) and reserve org-level dashboards for rolled-up trends across quarters. Every panel should link back to the underlying PRs, not just display an aggregate, so anyone questioning a number can audit it directly. Roll out in three steps: pilot on one team for a sprint, review whether the SLA targets held up against reality, then widen scope and adjust thresholds rather than treating the first numbers as permanent.

Common traps (Goodhart) and actionable defenses

Any metric turned into a target invites people to optimize the number instead of the outcome, a pattern well described by Goodhart’s law. Code review is a textbook case: reviewers under a TTFR target learn to leave a quick “LGTM” within the window and do the real reading later, or never.

  • Watch for approval times dropping while defect escape rate climbs, the clearest sign of rubber-stamping.
  • Watch for PR size creeping down artificially just to hit a review-time SLA, without any real decomposition benefit.
  • Defend with composite scores that pair speed against quality, so gaming one input shows up in the other.
  • Defend with periodic audits of a sample of “fast” reviews to confirm substantive comments were actually made.

Many automated review-evaluation metrics are sensitive to surface-level rewrites, and robustness must be validated before relying on an LLM-judged score. ArXiv robustness study, 2026

If you introduce an LLM-based reviewer or scorer into your pipeline, validate it against known defects before trusting its output as a metric input.

Evidence-backed review scoring and why it matters for metric fidelity

Metric fidelity depends on the reviews behind the numbers being trustworthy. Veridical focuses on catching critical defects before merge and backs every finding with concrete evidence rather than a generic comment, which reduces the ambiguity that erodes trust in quality ratio and escape-rate figures. Its scoring is built on F1 metric projections of real defects identified, giving teams a calibrated number instead of a guess, and its summaries make the underlying evidence checkable rather than asserted.

Iteration and cycle metrics: review rounds and rework time

Turnaround and quality metrics describe one pass through review. Iteration metrics describe what happens when a PR bounces back and forth, and that bouncing is often where the real time goes.

Track review rounds per PR, the number of times a reviewer requests changes before approval, and rework time, the interval between a change request and the author’s next push. A PR with three rounds and short rework gaps signals healthy back-and-forth on a genuinely complex change. The same three rounds with multi-day gaps between each usually signals context-switching cost or an author who has moved on to other work between reviews.

Review rounds and rework time comparison

High round counts paired with small PRs often point to unclear requirements rather than a review problem, since reviewers keep discovering scope that should have been settled before the PR opened. Track rounds and rework time by PR size band (small, medium, large) rather than in aggregate, because a two-round cycle on a 400-line PR means something different than the same count on a 20-line one. When rework time consistently exceeds the original review’s TTFR, the bottleneck has moved from “getting eyes on the code” to “getting the author back to the keyboard,” and that calls for a different fix, like protected focus blocks rather than faster reviewer assignment.

Impact of code review metrics on code quality and team morale

Metrics change behavior, for better or worse, and code review is a place where that effect shows up fast. A well-designed set of review metrics, tracked transparently and tied to outcomes rather than individual blame, tends to raise the floor on review consistency: reviewers know what “thorough” looks like because the quality ratio makes it visible, and authors get feedback within a predictable window because TTFR is watched.

The same metrics, applied carelessly, do the opposite. Publishing individual reviewer speed leaderboards without context turns review into a race, and GitHub’s own research on review practice emphasizes timely feedback and small PRs precisely because slow, oversized reviews create friction that erodes morale on both sides of the review. Reviewers rushed by a visible clock produce shallower comments, and authors watching a public queue depth number feel judged rather than supported.

The fix is not fewer metrics but better framing: report team-level trends more prominently than individual rankings, and use load-balancing data (queue depth, Gini coefficient) to redistribute work rather than to single anyone out. Morale holds up when metrics visibly drive better staffing and clearer expectations, and it degrades when the same numbers are used only to apply pressure.

Contextualizing metrics based on project size and team structure

A five-person startup team and a five-hundred-engineer platform organization cannot use the same SLA targets, even though both benefit from tracking the same core metrics. Team size changes what “normal” TTFR looks like: a small team with full context on every PR might sustain a two-hour P50, while a large org with cross-team dependencies routinely waits longer simply because the right reviewer has to be found first.

Codebase maturity matters as much as headcount. A greenfield project moving fast will show higher PR volume and shorter review cycles almost by default, while a mature system with strict change-control processes will show longer TTA on paper without that being a problem, since more of that time is deliberate risk review rather than a bottleneck.

Monorepos and polyrepos also shift what “reviewer load” means: a monorepo with a small set of expert reviewers for a critical module will show a high Gini coefficient by design, and treating that as a red flag without accounting for structure leads to reflexively adding reviewers who lack the context to catch anything. Set baseline targets per team, revisit them quarterly, and never import another team’s SLA numbers wholesale.

Benchmarking code review metrics across industries or comparable teams

External benchmarks are useful for calibration but easy to misuse if treated as universal targets. DORA’s research remains the most cited anchor point: teams with faster code review report software delivery performance around 50% higher than slower-reviewing peers, a gap wide enough to justify investing in review-speed improvements regardless of industry. The 2022 Accelerate State of DevOps Report ties review speed and related engineering practices to the four core DORA metrics: lead time for changes, deployment frequency, change failure rate, and time to restore service.

Beyond that broad finding, cross-industry benchmarking gets thin. Regulated industries with mandatory review gates (finance, healthcare, aviation software) will show structurally longer TTA than a consumer app team, and that difference reflects compliance requirements rather than review inefficiency. Use DORA’s performance tiers as a rough compass for where your delivery outcomes sit, then benchmark your own metrics against your own historical baseline quarter over quarter. That comparison is more actionable than any cross-company number, because it controls for your codebase, your team, and your constraints.

Best practices for collecting and analyzing code review metrics

Collection discipline determines whether any of this analysis is trustworthy six months from now.

  • Pull raw data from your version control platform’s API (PR timestamps, comment metadata, approval events) rather than manual logs.
  • Join post-merge defect data to PR IDs through commit SHAs, so escape rate is traceable rather than estimated.
  • Recompute P50/P90 figures weekly, retain raw data for at least a full quarter, and never overwrite historical numbers when SLAs change.
  • Exclude bot-generated comments and draft-PR time from TTFR calculations, since including them distorts the true human response window.
  • Decompose large PRs into a reviewable stack before measuring cycle time, since one giant PR will skew size and round-count metrics for the whole team.

Analyze trends before absolute values. A P90 that moved from six hours to nine hours over a month matters more than knowing today’s number in isolation, and a single outlier PR should never be allowed to reset a team’s baseline expectations.

Limitations and potential biases in common code review metrics

Every metric in this playbook has a blind spot, and pretending otherwise is how dashboards lose credibility.

TTFR rewards a quick acknowledgment, not a thorough one, so a reviewer can hit a strong TTFR by leaving a one-line comment and returning hours later to do the real work; the metric alone cannot see that gap. Defect escape rate is only as good as your incident tracking: a team with weak post-merge monitoring will show an artificially low escape rate that reflects poor detection rather than good prevention. Quality ratio can be gamed by padding substantive-looking comments that do not actually block a real problem, which is why periodic manual audits still matter even with the metric in place.

Automated scoring carries its own bias risk. The arXiv robustness study found that many LLM-based evaluation metrics respond to surface-level phrasing rather than substance, which means a metric that looks precise can still be measuring the wrong thing. GitHub’s research reaches a similar conclusion from a different angle: AI-assisted review can strip out trivial nit-picks and raise the baseline, but architectural judgment and ethical trade-offs still require a human reviewer, and no current metric captures that judgment directly. Treat every number here as a signal to investigate, never as a verdict on its own.

Limitations and potential biases in common code review metrics — overview diagram

What managers should do first

Start with TTFR, defect escape rate, and a hard PR size limit. Measure P90 for a full sprint and check it against lead time for changes before you touch any SLA. Keep engineers on the judgment calls; let automation handle the mechanical checks.

— Łukasz

How Veridical helps you measure and improve review quality

Evidence-backed reviews make your metrics worth trusting. Veridical attaches verified findings and inline evidence to every pull request, scored with F1-based projections of real defects rather than a generic pass or fail, so your quality ratio and escape-rate numbers reflect what was actually caught, not what was asserted. That advisory score wires directly into merge gating, giving your dashboard a defensible input instead of a guess. Compare plans on the pricing page.

Veridical

Primary sources to consult

  • DORA: research connecting review speed to delivery performance.
  • GitHub Blog: findings on AI’s role alongside human reviewers.
  • ArXiv study: robustness warnings for automated review-evaluation metrics.
  • Secure AI governance guidance from MARFI for teams validating automated review tools.

Sources

FAQ

What are some metrics for measuring code quality?

Common code quality metrics include defect escape rate, code review quality ratio, cyclomatic complexity, and post-merge incident counts. In review specifically, blocking findings per PR and defect escape rate are the two most direct signals of whether review is actually catching problems.

What is the 40/20/40 rule in software engineering?

Definitions of this rule vary across teams and are not tied to a single authoritative source, so treat any specific version with caution. A common interpretation splits effort across planning, building, and testing or review phases, but no standardized industry definition governs the exact percentages.

What are the top 5 code review tools?

Tooling choices depend on your stack, team size, and whether you need automated defect detection or just workflow management, so no single ranked list applies universally. Evaluate options by how well they surface verifiable findings and integrate with your existing merge gate, rather than by popularity alone.

What should be checked in code review?

A thorough review checks correctness, security implications, performance impact, and whether the change fits the surrounding architecture, not just style and formatting. GitHub’s guidance for effective review emphasizes prioritizing timely feedback and small, focused pull requests to keep that checklist manageable.