Veridical14 min readArticle

Code Quality Scores: Why Metrics Need Verified PR Reviews

Code Quality Scores: Why Metrics Need Verified PR Reviews

Geometric score path passing through verification

A code quality score is a composite rating, usually expressed as 0 to 100 or as Excellent, Good, Fair, or Poor, that combines maintainability, reliability, security, and test coverage signals into a single number engineering teams can act on. It exists to answer one question fast: is this code safe to merge, and where should review effort go next? Teams use it immediately in two ways: as a gate that blocks risky pull requests and as a prioritization signal for where to focus refactoring.


TL;DR:

  • The Maintainability Index rates 0–9 red, 10–19 yellow, and 20–100 green; prioritize red modules, but distinguish isolated hotspots from a falling repository average.
  • Prioritize files where churn meets low coverage or high complexity; either signal alone is less informative, and static metrics miss runtime and concurrency failures.
  • Run pull request scans and coverage checks on changed files first, then expand merge gates repository wide when teams trust the signal’s false positive rate.
  • Defect calibrated weights can outperform arbitrary sums, but historical models may not transfer between codebases, and correlated measures can double count one underlying problem.

Veridical
veridical.dev
Ground Code Quality Scores in Evidence
Veridical reviews GitHub pull requests with verified findings, detailed summaries, and inline evidence to help teams assess code quality.
Visit Veridical

Table of Contents

What a code quality score measures

A quality score typically rolls up four pillars: maintainability (how easily the code can be changed), reliability (how likely it is to fail in production), security (exposure to known vulnerability patterns), and testing adequacy (how much of the behavior is verified). Each pillar matters for a different reason, and all contribute to assessing code quality. Maintainability predicts the cost of future changes. Reliability and security predict the cost of doing nothing. Testing adequacy predicts whether a regression gets caught before a user does.

Most scores blend two kinds of input. Quantitative metrics, like complexity counts or coverage percentages, come from static analysis and are cheap to compute at scale. Qualitative findings come from human or evidence-based review and catch issues that counting alone misses, such as a flawed assumption about how a caller invokes a function. A score built only from quantitative metrics will miss dynamic runtime errors, infrastructure misconfiguration, and concurrency bugs that only appear under load. Treat the number as a starting point, not a verdict.

Core quantitative metrics that feed a code quality score

Most scoring systems draw from a shared pool of metrics that have been tracked in software engineering for decades, each correlating with maintenance cost or defect risk in a different way.

  • Cyclomatic complexity counts independent paths through a function; higher values generally make a module harder to test exhaustively and harder to reason about, though experienced teams often tolerate higher complexity in well-covered modules rather than applying a blanket limit, per Microsoft’s complexity guidance.
  • Maintainability Index combines Halstead volume, cyclomatic complexity, and lines of code into a single 0 to 100 figure, with bands at 0 to 9 (red), 10 to 19 (yellow), and 20 to 100 (green), as documented by Microsoft Learn.
  • Test coverage, including line and branch coverage, measures the extent to which code is executed during tests; coverage has been shown to be a strong indicator of defect proneness according to recent defect-prediction research.
  • Code churn and file age track how frequently a file changes and its recency; files with high churn combined with low coverage are often more likely to contain undiscovered bugs.
  • Duplication, class coupling, lines of code, and bug rate round out the picture, flagging copy-pasted logic, overly entangled classes, oversized files, and historical defect density.

No single metric tells the whole story. A short file with low complexity can still be fragile if it churns constantly and has no tests covering its branches.

How code quality scores are composed: normalization, weighting, and banding

Before they can combine into one score, each metric gets normalized, typically rescaled to a common 0 to 100 range using fixed thresholds or percentile ranks against a baseline codebase.

Once normalized, metrics are weighted and summed. Equal weighting is simple but naive: it treats a cosmetic duplication issue the same as a security-relevant complexity spike. Defect-calibrated weighting instead sets each metric’s influence based on how strongly it predicted real defects in historical data, an approach supported by research showing effort-related and coverage metrics outperform purely structural ones for defect prediction.

Banding, grouping scores into labels like Excellent, Good, Fair, and Poor, makes the number easier to communicate to non-engineering stakeholders than a bare figure. The main pitfall is correlated metrics: lines of code, complexity, and coupling often move together, so naive weighting can double-count the same underlying problem and inflate its apparent severity.

Measuring code quality in CI/CD and gating pull requests

Scoring only matters if it runs where decisions happen: inside the pull request, before code reaches the default branch. GitHub’s code quality tooling runs deterministic scans on both pull requests and the default branch, posts findings inline on the diff, ingests coverage reports such as Cobertura or JaCoCo, and lets teams define rulesets that block a merge when thresholds aren’t met, according to GitHub’s documentation.

A practical setup follows a few steps:

  1. Add a static analysis and coverage step to the pull request workflow, not just the nightly build.
  2. Import coverage reports in a standard format so line and branch coverage feed the same dashboard as complexity and duplication metrics.
  3. Define a ruleset that blocks merges below a coverage floor or above a complexity ceiling, scoped to changed files first.
  4. Expand the ruleset to the full repository once the team trusts the signal and false-positive rate.

Keeping PR checks fast and scoped to the diff, rather than rescanning the entire repository on every push, preserves developer trust in the gate. Static analysis platforms like SonarQube can plug into this same pipeline alongside the source control layer, an integration pattern Opsphere covers in the context of broader delivery monitoring.

Interpreting scores and practical thresholds for action

A number only helps if the team agrees on what it means. Using the Maintainability Index bands as a reference, a score in the 0 to 9 red zone signals a module that is difficult to maintain and worth immediate attention, 10 to 19 yellow suggests moderate risk worth monitoring, and 20 to 100 green indicates code that’s reasonably maintainable, per Microsoft’s archived documentation on the formula.

Maintainability score bands and action thresholds

A single red file in an otherwise green codebase is a localized hotspot: fix it and move on. A codebase where the average score trends downward release over release points to systemic risk, often from accumulated shortcuts under deadline pressure, and calls for a broader remediation plan rather than a single patch.

Use cases follow naturally from this logic: release gating blocks a build when the aggregate score drops below an agreed floor, audit prep uses scores and their history to demonstrate due diligence, and backlog prioritization ranks refactoring tickets by which files carry the worst scores combined with the highest churn.

Concrete remediation steps to improve a code quality score

Improving a score is rarely about chasing the number itself. It’s about fixing the underlying conditions the metrics are detecting.

  • Expand tests where coverage is thin, starting with files that combine low coverage and high churn, since those are the modules most likely to regress.
  • Triage complexity and churn hotspots first: a file with both high cyclomatic complexity and frequent recent changes carries more risk than either factor alone.
  • Refactor duplication and reduce coupling incrementally, in small commits that each keep the test suite green, rather than one large rewrite.
  • Introduce linting and automated rule checks, then tune the ruleset to cut noise, since a gate that produces too many false positives gets ignored or disabled.
  • Enforce smaller, focused pull requests, since smaller diffs are easier to review thoroughly and less likely to hide a defect in volume.

These practices align with established code review best practices that treat review discipline as a quality lever in its own right, not just a formality before merge.

Pro Tip: Fix the two or three files with the worst combination of complexity and churn before touching anything else; they typically account for a disproportionate share of a repository’s defect risk.

Why evidence-based scoring and verified PR reviews improve prediction

Supervised research on defect prediction has found that adding effort-related features, developer activity, code churn, file age, alongside test coverage metrics meaningfully improves prediction accuracy over structural metrics alone, with reported AUC improvements reaching roughly 0.932 in recent peer-reviewed work. Adding more structural metrics on top of these features tends to yield diminishing returns.

This is the logic behind F1-projected scoring: instead of an arbitrary weighted sum, the score is calibrated against how well it would have identified real, confirmed defects in historical data, expressed through F1, a balance of precision and recall. Pairing that calibration with verified findings, evidence tied to each flagged issue rather than an unexplained number, gives teams something they can act on and defend in an audit, not just a score to watch trend lines.

Confirmed defects calibrate evidence-linked review findings

Comparison of code quality score approaches and methodologies across different tools and vendors

Scoring approaches across the industry fall into roughly three categories. Rule-based rating systems, the kind built into many static analysis platforms, summarize findings by severity (error, warning, note) into bands like Excellent, Good, Fair, or Poor, as described in GitHub’s code quality documentation. These are fast to compute and easy to explain, but they weight a style warning and a security-relevant error using the same severity taxonomy unless a team customizes it.

Formula-based metrics, like the Maintainability Index, compute a single number from a fixed mathematical combination of Halstead volume, cyclomatic complexity, and lines of code. They’re transparent and reproducible, since the formula is public, but they don’t incorporate historical defect data and can rate two structurally similar files identically even when one has far better test coverage.

This category is newer and depends heavily on the quality and size of the historical defect data used for calibration; a weighting scheme trained on one kind of codebase may not transfer cleanly to another without recalibration.

Most engineering teams end up blending approaches: a formula-based metric for a quick maintainability read, rule-based ratings for day-to-day linting, and defect-calibrated scoring where the cost of a missed defect is high enough to justify the extra setup.

How to integrate code quality scores with team workflows and developer collaboration

A score that only exists on a dashboard doesn’t change behavior. It needs to show up where developers already work: inline on the pull request diff, as a required check before merge, and as a trend line reviewed in sprint retrospectives. Posting findings directly on the lines they affect, rather than in a separate report, shortens the loop between a flag and a fix.

Collaboration improves when the score comes with explanation rather than a bare number. A reviewer who sees why a file dropped from Good to Fair, which metric moved and by how much can have a concrete conversation with the author instead of a vague one about “code quality.” Teams that adopt evidence-based pull request review practices tend to build this habit naturally, since each finding already carries supporting evidence rather than an opaque severity label.

Ownership matters too. Assigning score trends to specific services or teams, rather than tracking one organization-wide number, makes accountability clear and avoids diluting a serious regression in one module with improvements elsewhere. Pairing automated scoring with a lightweight manual review step for anything flagged red keeps the human judgment that pure automation still misses, particularly around business logic and edge cases a static analyzer has no way to understand.

Author perspective on practical use of code quality scores

Treat a code quality score as a prioritization signal, not a verdict. The number indicates where to begin investigation, not a definitive assessment of code safety. Combining automated scoring with evidence-based human review provides the best results, starting with gating your highest-risk repositories before expanding thresholds.

— Łukasz

How Veridical implements evidence-based scoring and PR reviews

Veridical

We built Veridical around the gap this article keeps coming back to: a quantitative score is only as useful as the evidence behind it. Our evidence-based pull request reviews analyze repository context, not just the diff, pulling in callers, interfaces, and dependencies so a flagged issue reflects how the code is actually used. Every finding ships with inline evidence, and every review ends with an advisory score calibrated against real-defect F1 projections, the same defect-calibrated logic this article describes, so the number maps to expected detection performance rather than an arbitrary formula.

That score wires directly into your merge gate, giving you a concrete, defensible threshold instead of a dashboard nobody checks. Teams get clearer remediation guidance, safer merges, and an audit trail of verified findings for every pull request we touch, with security and data handling documented for private repositories.

Open-source maintainers can try Veridical free, and teams ready to gate PRs at scale can compare the Launch tier, Standard at $39 per month, and Pro at $69 per month on our pricing page.

FAQ

How do you measure code quality?

You measure code quality by combining quantitative metrics, cyclomatic complexity, maintainability index, test coverage, duplication, and churn, into a single score, then validating flagged issues with evidence-based review. Static analysis tools compute the metrics automatically, while defect-prediction research shows that weighting coverage and effort-related metrics heavily improves how well the score predicts real defects.

What is the maintainability index and what do its bands mean?

The Maintainability Index is a 0 to 100 score computed from Halstead volume, cyclomatic complexity, and lines of code, with 0 to 9 flagged red, 10 to 19 yellow, and 20 to 100 green, per Microsoft’s documentation. Red indicates a module likely to be difficult to maintain and worth prioritizing for refactoring.

Can ChatGPT do a code review?

A general-purpose language model can read code and suggest changes, but it typically lacks repository context, like callers and dependencies, and doesn’t verify its findings against reproducible evidence. Evidence-based review tools that check claims against the actual codebase and attach concrete evidence to each finding, like Veridical’s approach, are built specifically to close that gap before code merges.

What is the 40/20/40 rule in software engineering?

This isn’t a standardized or widely documented metric in code quality literature, and definitions vary depending on the source. If you’ve seen it referenced for test effort or review time allocation, treat it as a team-specific heuristic rather than an industry standard.

What are the most important code quality metrics to track first?

Start with test coverage and cyclomatic complexity, since coverage is among the strongest predictors of defect proneness identified in recent research, and high complexity compounds that risk. Add code churn next, since files that change often and have thin coverage are the most likely to hide unresolved defects.

Sources

Code Quality Scores: Why Metrics Need Verified PR Reviews · Veridical.dev