Veridical19 min readArticle

Developers: Build a Valid SARIF File, One Tool, One Run, One Result

Developers: Build a Valid SARIF File, One Tool, One Run, One Result

Isometric SARIF validation structure

SARIF is the standardized JSON format for static analysis outputs that lets any scanner feed a common viewer, dashboard, or CI gate. For tool producers, it means writing one output format instead of maintaining custom exporters for every downstream consumer. For consumers, it means one parser handles findings from ESLint, CodeQL, or a proprietary scanner alike, an approach documented in the OASIS SARIF specification and explained further on Sarif.


TL;DR:

  • Using stable rule identifiers and consistent partialFingerprints based on file path, rule ID, and line number is crucial for effective deduplication across scans.
  • Platform limits on file size and result counts demand proactive trimming, splitting, and prioritization of high-severity findings to ensure successful upload and ingestion.
  • Always validate SARIF files with schema checkers and visual viewers before upload to prevent silent validation failures and ensure accurate rendering on platforms like GitHub and GitLab.
  • Incorporate platform-specific properties like partialFingerprints and correct path resolution early in your workflow to avoid common ingestion issues and ensure reliable results display.
  • Start simple with a minimal valid SARIF example, then gradually add richness, focusing on stable identifiers and canonical URIs to future-proof your pipeline.

Veridical
Strengthen Your Repository Reviews
Veridical provides evidence-based pull request reviews with verified findings, detailed summaries, and inline evidence for GitHub repositories.

Table of Contents

SARIF file anatomy: sarifLog, runs, results, schema and version

Every SARIF file starts with the same skeleton, and once you recognize it, reading an unfamiliar tool’s output becomes mechanical rather than mysterious. The root object is called sarifLog, and it holds exactly two things that matter on first read: a $schema reference and an array called runs.

Each entry in runs represents one execution of one analysis tool. A single file can contain multiple runs, which is how you get combined output from, say, a linter and a security scanner in one artifact. Inside each run, you’ll find a tool object describing what produced the findings, and a results array, which is the payload everyone actually came for.

A result is the atomic unit of a SARIF file: one flagged issue, one location, one message. At minimum, a valid result needs a message (what’s wrong, in plain text) and typically a ruleId (which check flagged it) and at least one entry in locations (where it happened). Skip the location and most viewers will accept the file but render the finding as orphaned or unclickable, which defeats the purpose of structured output in the first place.

The $schema and version properties deserve attention because they’re the first thing any validator or viewer checks. version should read "2.1.0", matching the OASIS SARIF specification. For $schema, point to the canonical schema URI rather than a local or cached copy, since mismatched or missing schema references are a common source of silent validation failures downstream.

A few habits will save you debugging time later:

  • Always set version to "2.1.0" and match it with a $schema pointing to the correct published schema URI.
  • Treat results as the primary array to get right first; everything else in a run supports it.
  • Keep one runs entry per tool invocation rather than merging unrelated tools into a single run object.
  • Validate the skeleton (sarifLog to runs to results) before worrying about optional richness like code flows or related locations.

Getting this shape right on the first pass avoids most of the downstream ingestion problems covered later in this guide.

Key SARIF objects and properties you will actually use

Beyond the skeleton, a handful of objects do most of the work in any real-world SARIF file. Understanding them well is the difference between output that merely validates and output that’s genuinely useful to a human reviewing findings or a platform trying to deduplicate them.

tool.driver and the rules array. Every run’s tool object contains a driver describing the analysis engine: its name, version, and critically, a rules array. Each rule entry should carry an id that result objects reference via ruleId, plus a helpUri linking to documentation and a short properties bag for anything tool-specific. Including helpUri matters more than it looks: without it, a developer seeing an unfamiliar rule ID has no path to understanding why the finding exists or how to fix it. Omit the rules array entirely and your ruleId references become opaque strings with no explanatory context.

result fields that drive triage. A result’s ruleId ties it back to its rule definition, while level (typically error, warning, or note) and message.text tell a reviewer how urgent it is and what happened. locations points to where in the codebase the issue lives, and partialFingerprints gives platforms a stable identifier to track the same finding across multiple scans, even as line numbers shift with unrelated edits.

location objects and the path down to region. A location wraps a physicalLocation, which wraps an artifactLocation (the file, via uri and often uriBaseId) and a region (the specific line and column range). This nesting looks verbose at first glance, but each layer solves a real problem: artifactLocation separates the file reference from the base path so the same result stays valid whether it’s read on a developer’s laptop or in a CI container, and region narrows the finding to something a viewer can actually highlight. The SARIF 2.1.0 specification treats at least one physical location as effectively required for any result a viewer is expected to display meaningfully.

property bags for tool-specific metadata. Nearly every SARIF object supports an optional properties field, a free-form object for anything the standard doesn’t formally define. This is where tool authors attach extra context: confidence scores, internal categorization, links to supporting evidence. Because consumers that don’t recognize a property simply ignore it, property bags let you extend SARIF output without breaking any platform or viewer already parsing your files, a pattern the OASIS SARIF specification explicitly supports.

Pro Tip: Populate rules[].helpUri even for internal or proprietary checks; a placeholder documentation link is more useful to a confused developer than none at all.

Get these four object families right, tool, result, location, and property bags, and you’ll be able to read or write the overwhelming majority of SARIF files you encounter without consulting the spec line by line.

Four connected SARIF object families

A minimal valid SARIF example with field-by-field notes

The fastest way to internalize SARIF’s structure is to look at the smallest file that still qualifies as useful output. Here’s a minimal but realistic example: one tool, one run, one result.

{
  "$schema": "https://raw.githubusercontent.com/oasis-tcs/sarif-spec/main/Schemata/sarif-schema-2.1.0.json",
  "version": "2.1.0",
  "runs": [
    {
      "tool": {
        "driver": {
          "name": "ExampleScanner",
          "version": "1.0.0",
          "rules": [
            {
              "id": "EX001",
              "helpUri": "https://example.com/rules/EX001"
            }
          ]
        }
      },
      "results": [
        {
          "ruleId": "EX001",
          "level": "warning",
          "message": { "text": "Unsanitized input passed to query builder." },
          "locations": [
            {
              "physicalLocation": {
                "artifactLocation": { "uri": "src/db/query.js" },
                "region": { "startLine": 42 }
              }
            }
          ]
        }
      ]
    }
  ]
}

Every field here earns its place, and viewers rely on specific ones to render anything useful.

Field Purpose Required for display?
$schema Points validators to the schema definition Recommended, not strictly required
version Declares SARIF version, 2.1.0 Yes
tool.driver.name Identifies the scanner in the viewer UI Yes
rules[].helpUri Links a finding to documentation Optional, strongly recommended
message.text The human-readable description of the issue Yes
locations[].physicalLocation Where the finding appears in the file tree Yes for clickable results
region.startLine The specific line a viewer highlights Recommended

Extend this minimal shape with partialFingerprints once you’re deduplicating across runs, and with a fixes array if your tool can propose an automatic patch. Neither is required for a file to validate, but both are what separate a barely-compliant SARIF file from one that’s genuinely pleasant to consume in a pull request review or a security dashboard.

How to produce SARIF from your tools and CI pipeline

Most teams don’t write SARIF by hand. They either use an existing formatter built into a scanner or reach for a language-specific library when building a custom analyzer.

  1. Check for a built-in SARIF formatter first. ESLint, for instance, ships a SARIF formatter you invoke directly from the command line, and many commercial and open-source static analysis tools follow the same pattern.
  2. Use an official or community SDK for custom tooling. Libraries such as sarif-om on PyPI give you typed object models for constructing SARIF output in Python rather than hand-assembling JSON, and equivalent SDKs exist for other ecosystems.
  3. Generate one SARIF file per tool run in CI, rather than trying to merge multiple tools’ output into a single file at generation time. Merging is easier to do reliably as a separate post-processing step.
  4. Adopt a consistent naming convention for output artifacts, such as {tool-name}-{run-id}.sarif, so downstream steps and artifact retention policies can find and process files predictably.
  5. Retain SARIF artifacts alongside build logs for at least as long as your platform’s audit or compliance window requires, since regenerating historical scan results after the fact is rarely possible.
  6. Filter and batch results before upload when a scan produces a large number of findings, prioritizing by level so high-severity results survive any trimming needed to respect platform limits.

The batching step matters more than it sounds. A scanner that reports every stylistic nit alongside genuine defects will often exceed platform ingestion limits, discussed in detail further down, long before it exceeds any technical constraint in the SARIF format itself.

Pro Tip: If your CI pipeline runs the same scanner across multiple services in a monorepo, keep each service’s SARIF file separate rather than combining them; platforms handle per-component uploads more predictably than one oversized combined file.

Once generation is reliable, the next question is whether the output you’re producing is actually correct, which is where validation and viewing tools come in.

Validating and viewing SARIF before you trust it

A SARIF file that generates without errors isn’t the same as a SARIF file that’s correct. Validating early catches malformed output before it reaches a platform’s stricter and less forgiving ingestion checks.

  • Run a JSON schema validator against your file using the schema URI referenced in $schema; most general-purpose JSON schema tools handle SARIF’s schema without modification.
  • Use the VS Code SARIF Viewer extension for a fast visual check, since it renders locations, messages, and rule metadata the same way a platform’s UI eventually will.
  • Try the sarif.info explorer for a browser-based inspection when you don’t want to install anything locally.
  • For quick command-line spot checks, pipe the file through jq to confirm the shape of runs, results, and nested locations before reaching for a full viewer.
  • Add a CI step that fails the build on schema validation errors, rather than allowing invalid SARIF to reach an upload step that fails later with a less specific error.

The last point is worth building early rather than retrofitting. A validation failure caught in a dedicated CI step produces a clear, actionable error message. The same malformed file caught by a platform’s upload endpoint often produces a vague rejection with little indication of which field caused it.

Tests worth including in that CI step: confirming version and $schema are present and consistent, confirming every result has at least one location, and confirming ruleId values in results actually correspond to entries in the rules array. These three checks catch the overwhelming majority of malformed SARIF before it leaves your pipeline.

Platform ingestion and limits: GitHub and GitLab requirements

Passing schema validation is necessary but not sufficient. Both GitHub and GitLab implement a supported subset of the full SARIF specification, and understanding that subset is what separates a file that validates from one that actually displays correctly once uploaded.

GitHub parses SARIF 2.1.0 files and reads a defined subset of properties, using partialFingerprints specifically to deduplicate the same finding across repeated scans, according to GitHub’s SARIF support documentation. Without stable fingerprints, the same underlying issue can reappear as a “new” finding on every run, cluttering a repository’s alert history and eroding trust in the scan results. GitHub’s own guidance recommends constructing fingerprints from a stable hash of file path, region.startLine, and ruleId, deliberately avoiding volatile inputs like timestamps that would break identity across runs.

Platform limits require proactive handling. Both GitHub and GitLab impose practical soft and hard limits on file size and result counts, and files that exceed them produce upload errors documented in GitHub’s troubleshooting guide for oversized SARIF uploads. The standard itself permits arbitrarily large files, but production ingestion pipelines don’t, which means large scans routinely need pre-processing before upload.

  • Split a single oversized run into multiple smaller runs rather than attempting one monolithic upload.
  • Trim thread flow locations and other verbose optional fields when a file is close to a size limit.
  • Prioritize high-severity results for inclusion when a result count limit forces you to drop findings.
  • Set partialFingerprints deliberately using stable inputs, never timestamps or line offsets alone.

GitLab ingests SARIF 2.1.0 reports on a similar principle, mapping SARIF fields into its own internal vulnerability report types and documenting its own field mappings and limits separately from GitHub’s. The practical takeaway for producers targeting both platforms is the same: treat the SARIF spec as the floor, not the ceiling, and always check each platform’s own documentation for the specific properties and limits it enforces before assuming full compatibility.

One deduplication mechanism matters more than any other property in this section: stable partialFingerprints are what keep a finding recognized as “the same issue” across scans rather than resurrected as new noise every time your pipeline runs, a behavior confirmed directly in GitHub’s SARIF documentation.

Stable finding identity across scans

Common pitfalls and a troubleshooting checklist

Most SARIF problems teams encounter fall into a small number of recurring categories, and working through them in order resolves the majority of ingestion failures.

  1. Fix path resolution first. Relative artifactLocation.uri values need a matching uriBaseId to resolve correctly across a developer’s machine and a CI container. Emit originalUriBaseIds with both an absolute base and a relative one pointing at the repository root, a pattern GitHub’s documentation recommends specifically because it holds up across environments.
  2. Confirm every displayed result has a physicalLocation and region. A result without at least one resolvable location will validate against the schema but render as orphaned or invisible in most viewers, since the SARIF specification treats location data as what makes a result actionable rather than merely descriptive.
  3. Build stable partialFingerprints before worrying about anything else. Hash the file path, the rule ID, and the starting line together, and avoid any input that changes between otherwise-identical runs.
  4. Trim or split large files by severity, not by position. When a file exceeds a platform’s result count limit, drop or defer low-severity findings first rather than truncating arbitrarily by array order, which risks silently discarding the issues that matter most.

Pro Tip: When a finding mysteriously disappears from a platform’s UI after upload, check partialFingerprints before anything else; an unstable fingerprint is the most common reason a previously visible result stops appearing.

Working this checklist in order, path resolution, then locations, then fingerprints, then size, catches nearly every practical SARIF failure teams report, because each step depends on the one before it: a fingerprint built on a broken path is unstable by definition, and a location that doesn’t resolve can’t be trimmed intelligently by severity.

Best practices checklist for SARIF producers and CI integration

A small set of habits, applied consistently, separates SARIF output that’s merely valid from output that’s genuinely reliable across viewers, platforms, and time.

  • Always include both $schema and version at the top of every file, pointed at the canonical 2.1.0 schema.
  • Use canonical, uriBaseId-resolved URIs for every artifact location rather than machine-specific absolute paths.
  • Populate rules[] with id and helpUri for every rule your tool can trigger, even minimal internal ones.
  • Set partialFingerprints using stable inputs and keep the hashing logic unchanged across tool versions.
  • Limit large arrays proactively and split runs before platform limits force an upload rejection.
  • Document your own tool’s conversion or export steps so a future maintainer can reproduce or adjust the SARIF output without reverse-engineering it.
Practice Why it matters Where it’s enforced
Schema and version present Enables validation and correct parsing JSON schema validators
Canonical URIs with uriBaseId Prevents path mismatches across environments GitHub, GitLab ingestion
Stable partialFingerprints Enables deduplication across scan runs GitHub code scanning
Rules array with helpUri Gives reviewers context on unfamiliar findings Viewers, dashboards

Treat this list as a pre-upload checklist rather than a one-time setup task. Tool versions change, rule sets grow, and a fingerprinting scheme that worked for version 1.0 of your analyzer can quietly break when version 2.0 changes how it numbers internal checks.

How Veridical relates to SARIF workflows

SARIF standardizes the shape of a finding: what rule fired, where, and with what message. It says nothing about whether that finding is actually a real defect, which is a separate and harder problem. Veridical addresses that second problem directly, publishing findings only when tied to concrete evidence and reproducible checks, and closing every pull request review with a calibrated advisory score built on real-defect F1 metrics.

That distinction maps naturally onto SARIF’s own extension points. A property bag on a result object can carry a verification status, an evidence pointer, or an advisory score without breaking any consumer that ignores unrecognized properties, exactly the extensibility mechanism the OASIS specification describes.

Picture a merge gate that already consumes SARIF from a static analyzer for broad, rule-based coverage. Layering Veridical’s review alongside it adds evidence-backed findings on the specific change set, plus an advisory score that can feed the same gating logic. The static analyzer catches known patterns at scale; Veridical’s review adds verified findings with supporting evidence on the code actually changing in that pull request. Teams exploring this combination can review Veridical’s field notes on evidence-based code review for more on how verification metadata gets structured.

A pragmatic path for teams adopting SARIF

Teams new to SARIF tend to over-plan before writing a single file, and that’s backward. Start with the smallest valid example, get it through a schema validator, and only then add richness like fingerprints or fix suggestions.

Stable identifiers matter more than any other decision you’ll make early on. A ruleId scheme and a partialFingerprints hashing approach chosen on day one will outlive several versions of your tool, so get them right before optimizing anything else.

Platform support will keep shifting. GitHub and GitLab both evolve their supported property subsets and limits over time, which means a thin post-processing step between “SARIF my tool emits” and “SARIF the platform ingests” pays for itself. That layer is where you adapt to platform quirks without touching your core generation logic.

— Łukasz

Veridical as an adjacent option for teams running SARIF pipelines

If your team already runs static analyzers through a SARIF-based pipeline, the gap you’re likely feeling isn’t format compatibility, it’s confidence in which findings are worth a developer’s attention. Veridical reviews pull requests directly on GitHub, verifying findings against concrete evidence and reviewing repository context beyond the diff alone. Reviews conclude with a calibrated advisory score based on real-defect F1 metrics that can integrate with existing merge gating.

Veridical

Veridical doesn’t replace your static analyzers or your SARIF pipeline; it adds a second, evidence-based layer focused on the change actually under review. Open-source projects can use it through the free tier for open-source repositories, while teams needing higher review volume can compare the Launch tier, Standard, and Pro plans on the pricing page. Check the AI code review page for the full picture of how verification works before you merge.

Sources

For readers who want the primary sources rather than a paraphrase, start with the OASIS SARIF 2.1.0 specification plus Errata 01, the normative reference for every object and property discussed above. The SARIF project home at sarif.info offers tutorials and a browser-based explorer for newcomers. For platform-specific ingestion rules, GitHub’s SARIF support documentation and its upload troubleshooting guide cover supported properties and common errors, while GitLab maintains its own separate documentation on field mapping and limits.

FAQ

What is SARIF format?

SARIF is a JSON-based standard, published by OASIS, for representing static analysis tool output in a consistent shape. It lets scanners from different vendors feed the same viewers, dashboards, and CI gates, as defined in the SARIF 2.1.0 specification.

How do I open a SARIF file?

Open a .sarif file with the VS Code SARIF Viewer extension for a full visual rendering of rules, messages, and locations, or use the browser-based explorer at sarif.info if you don’t want to install anything. For a quick raw look, any text editor or jq on the command line will show you the underlying JSON structure.

Can you provide an example of a SARIF file format?

A minimal SARIF file contains a sarifLog root with $schema, version, and a runs array, where each run has a tool.driver and a results array. A single result typically needs a ruleId, a message, and at least one locations entry pointing to a file and line, following the structure defined in the SARIF specification.

What are SARIF reports?

A SARIF report is the JSON output a static analysis tool produces after scanning code, structured so any SARIF-compatible viewer or platform can read it. Platforms like GitHub and GitLab ingest these reports directly, mapping fields like results and partialFingerprints into their own code scanning or vulnerability views, per GitHub’s SARIF support documentation.

Developers: Build a Valid SARIF File, One Tool, One Run, One Result · Veridical.dev