← Back to blog

Profile First, Example Driven FHIR Genomics DocumentReference for Labs

October 3, 2026
Profile First, Example Driven FHIR Genomics DocumentReference for Labs

The Genomics DocumentReference (Genomic Data File) profile is a FHIR DocumentReference profile used to register and describe genomic data files such as VCF or BAM, not to encode variant observations. Use it as the metadata and access wrapper for the file itself, while keeping structured variant calls in Observation-based profiles. The canonical definition lives in the HL7 Genomics Reporting Implementation Guide, which your lab's FHIR implementation should treat as the source of truth.


TL;DR:

  • Use content URL for large files like BAM or CRAM, and reserve base64 embedding for small files or test data, to avoid overloading servers.
  • Link DocumentReference to GenomicStudy or Specimen using context.related to maintain provenance and prevent orphaned files, especially in workflows with multiple artifacts.
  • Populate the description with pipeline name, reference build, and sample data to facilitate troubleshooting and ensure clear documentation.
  • Apply strict security labels such as Restricted for sensitive genomic files, and always verify security metadata before sharing or caching files.
  • Validate against specific IG versions and keep profile references versioned to avoid inconsistencies, since the Genomics Reporting IG is still in trial-use status.

SignalPGx
Turn Genomic Data Into Reports
SignalPGx helps laboratories transform genotype and medication data into physician-reviewed, evidence-based pharmacogenomic reports.
Explore SignalPGx

Table of Contents

Where the genomic data file profile fits in the genomics reporting IG

The Genomic Data File profile is a constrained version of the base FHIR DocumentReference resource, purpose-built to describe large sequencing artifacts rather than clinical notes or scanned PDFs. It is published and maintained by the HL7 Clinical Genomics Working Group as part of the Genomics Reporting Implementation Guide artifacts, which also defines GenomicStudy, GenomicStudyAnalysis, and a family of Observation profiles for variants, haplotypes, and genotypes.

Two things matter before you write a line of integration code. First, confirm which build you are validating against: the IG's continuous build at build.fhir.org moves faster than published releases, so pin your validator to the specific profile version your trading partners expect. Second, understand the IG's standards status. Much of the Genomics Reporting IG, including the Genomic Data File profile, carries trial-use maturity rather than normative status, which means element bindings and cardinality can still shift between releases. That is not a reason to avoid it; it is a reason to version-pin your StructureDefinition references and retest after every IG update.

The practical question most teams ask is when to reach for DocumentReference versus the other genomic resources in the IG. A short way to decide:

  • Use DocumentReference (Genomic Data File profile) when you need to register, discover, or retrieve a raw or semi-processed file such as a VCF, BAM, CRAM, or BED.
  • Use GenomicStudy when you need to describe a sequencing or analysis event that may produce multiple files and multiple downstream observations.
  • Use Observation-based profiles (Variant, Haplotype, Genotype) when you need to encode a specific, queryable genetic finding that clinical decision support or a pharmacogenomics engine will act on.
  • Use DiagnosticReport when you need the clinician-facing narrative that ties findings, files, and interpretation together for a single order.

Treating these as a layered stack, file, study, observation, report, rather than interchangeable containers, is what keeps a genomics FHIR implementation queryable instead of becoming a pile of opaque attachments.

What goes in each JSON field, element by element

Once the profile choice is settled, the real work is getting the JSON right. The DocumentReference R4 specification defines the base elements, and the Genomics Reporting IG narrows them for genomic use. Here is the field-by-field breakdown implementers actually need.

  1. resourceType must be "DocumentReference". There is no genomics-specific resource type: the profile is a constraint, applied through meta.profile, not a new type.
  2. identifier should carry a stable, lab-assigned business identifier (an accession number or sample ID) so the file can be re-matched if the FHIR server's logical ID changes during a migration.
  3. status is required and almost always "current" for an active file; use "superseded" when a reprocessed file replaces an earlier one, and point the new resource back with relatesTo.
  4. subject is required and references the Patient (or a Group resource for trio or family analyses) the file describes.
  5. content.attachment.contentType is required and must accurately describe the file's MIME type, since downstream systems rely on it to choose a parser.
  6. content.attachment.url or content.attachment.data is required: exactly one of these must be present to either point to the file or embed it.
  7. description is recommended and should state the reference genome build, pipeline name, and pipeline version, information that is otherwise invisible to a consuming system.
  8. content.attachment.size, creation, and title are recommended for discoverability and for giving human reviewers a readable label without opening the file.
  9. content.format (a Coding) is recommended to name the file format precisely, separate from the raw MIME type, which helps when a single contentType like application/octet-stream covers multiple formats.
  10. meta.security and securityLabel are strongly recommended given the sensitivity of raw genomic data; most implementations apply a confidentiality coding such as Restricted.
  11. type should use a LOINC code when one exists for the document kind, supporting interoperability with systems that query by document type rather than by profile.
  12. relatesTo.code (values like replaces, transforms, appends) should be set whenever a file supersedes or derives from another, and docStatus should reflect whether the underlying file content itself is still preliminary or final, distinct from the resource's own status.

The DocumentReference R4 page shows worked examples of several of these elements together, including common securityLabel and content.attachment.contentType patterns you can adapt directly.

Pro Tip: Populate description with the pipeline name and reference build every time: it is the cheapest field to fill and the one that saves the most time during a troubleshooting call six months later.

Serving large genomic files without breaking your FHIR server

A VCF file for a single exome can run into tens of megabytes; a BAM file can reach into the gigabytes. That size reality drives most of the design decisions around content.attachment.

  • Use attachment.data (base64-encoded inline content) only for small files, informational metadata documents, or test fixtures. US Core and related implementation guides set pragmatic expectations that inline payloads stay in the low single-digit megabytes, with servers documenting their own limits for anything larger, as described in the US Core guidance on writing clinical notes.
  • Use attachment.url for anything beyond that threshold, pointing to a dedicated file server, object store, or genomics-aware API endpoint rather than embedding the file in the FHIR payload itself.
  • Set contentType precisely: text/vcf or application/octet-stream with a content.format coding for VCF, application/octet-stream for BAM and CRAM (since there is no dedicated registered MIME type for these binary formats), and a documented coding scheme for BED files so consuming systems do not have to guess.
  • Populate size and creation so a consumer can decide whether to retrieve the file at all before pulling gigabytes over a network connection.
  • Add integrity metadata, such as a SHA-256 checksum captured in the attachment description or a locally agreed extension, so a receiving system can verify the file was not corrupted or truncated in transit.

One practical figure to anchor expectations: FHIR's own DocumentReference specification is explicit that the resource's design is deliberately format-agnostic, meaning the burden of conveying size, format, and provenance correctly falls entirely on the metadata you populate, not on the resource structure itself.

Server-side obligations matter just as much as producer-side discipline. A server exposing genomic DocumentReference instances should document its own size limits, support content negotiation so a client can request a specific representation, and use signed URLs or token-based access rather than open, permanent links given how sensitive the underlying data is. Expiry policies on those signed URLs should be documented too, since a broken link six months after a report is signed creates a support problem for your lab and a trust problem for the ordering clinician.

Connecting files to the rest of the genomic and clinical graph

A DocumentReference that sits alone, unconnected to a specimen, a study, or a report, is close to useless for clinical workflows. The Genomics Reporting Implementation Guide's general reporting page is direct about this: file-level resources should be linked to GenomicReport, GenomicStudy, and Specimen to preserve provenance and keep the file discoverable from the clinical context that produced it.

  • Use context.related to reference the GenomicStudy, Specimen, or GenomicReport that the file belongs to, which is the primary mechanism for answering "what is this file for" without opening it.
  • Use derivedFrom when the file was produced by transforming another resource, such as a VCF derived from a BAM, to preserve the processing lineage.
  • Reserve hasMember and result style relationships for Observation and DiagnosticReport resources rather than DocumentReference itself; those relationships describe findings, not files, and mixing the two muddies both.
  • Map DocumentReference into broader EHR event and documentation models (Event pattern, Composition references) carefully, since many EHR integration layers expect a Composition or DiagnosticReport entry point rather than a bare DocumentReference, a pattern also discussed in general terms in practical guidance on EHR integration.

The most common pitfall implementers run into here is orphaning: publishing a file-level DocumentReference with no link back to the ordering ServiceRequest or the Specimen it came from. The IG's own guidance is to treat context.related and derivedFrom as mandatory discipline, not optional polish, specifically because orphaned files are what make genomic data warehouses unsearchable within a year of going live.

Minimal and fuller JSON examples for a Genomics DocumentReference

The cleanest way to internalize the field guidance above is to see it in a working instance. Both examples below follow the structure used in the Genomics Reporting Implementation Guide's example artifacts and the published DocumentReference-genomicFileGroupAsSubject example.

  1. Minimal single-patient VCF reference. This is the smallest instance that a validator will accept while still being operationally useful: it declares the resource type, a stable identifier, status, subject, and a file pointer with a correctly stated content type.
{
  "resourceType": "DocumentReference",
  "meta": {
    "profile": ["http://hl7.org/fhir/uv/genomics-reporting/StructureDefinition/genomic-file"]
  },
  "identifier": [
    {
      "system": "http://examplelab.org/accession",
      "value": "ACC-2026-00412"
    }
  ],
  "status": "current",
  "subject": {
    "reference": "Patient/pat-01"
  },
  "description": "Germline VCF, GRCh38, variant-calling pipeline v3.2",
  "content": [
    {
      "attachment": {
        "contentType": "text/vcf",
        "url": "https://files.examplelab.org/genomics/ACC-2026-00412.vcf.gz",
        "size": 18400213,
        "creation": "2026-02-11T09:22:00Z"
      }
    }
  ]
}

Each element here does a specific job: identifier lets the lab reconnect this resource to its own accessioning system if the FHIR logical ID changes; status tells a consumer the file is the current, authoritative version; description states the pipeline and reference build that a bare filename cannot convey; and the single content.attachment entry supplies the contentType and url required for retrieval, plus the size and creation timestamp that let a consumer decide whether and when to pull the file.

  1. Fuller trio-linked example with security labeling. Family or trio sequencing studies often use a Group as the subject, and the added sensitivity of multi-person genomic data typically calls for an explicit confidentiality label.
{
  "resourceType": "DocumentReference",
  "meta": {
    "profile": ["http://hl7.org/fhir/uv/genomics-reporting/StructureDefinition/genomic-file"],
    "security": [
      {
        "system": "http://terminology.hl7.org/CodeSystem/v3-Confidentiality",
        "code": "R",
        "display": "Restricted"
      }
    ]
  },
  "identifier": [
    {
      "system": "http://examplelab.org/accession",
      "value": "ACC-2026-00917-TRIO"
    }
  ],
  "status": "current",
  "type": {
    "coding": [
      {
        "system": "http://loinc.org",
        "code": "51969-4",
        "display": "Genetic analysis report"
      }
    ]
  },
  "subject": {
    "reference": "Group/trio-family-917"
  },
  "description": "Trio BAM alignment set, GRCh38, joint-calling pipeline v4.0",
  "content": [
    {
      "attachment": {
        "contentType": "application/octet-stream",
        "url": "https://files.examplelab.org/genomics/ACC-2026-00917-TRIO.bam",
        "size": 4821039552,
        "creation": "2026-03-04T14:05:00Z",
        "title": "Trio BAM, proband + two parents"
      },
      "format": {
        "system": "http://hl7.org/fhir/uv/genomics-reporting/CodeSystem/genomic-file-format",
        "code": "bam",
        "display": "BAM"
      }
    }
  ],
  "context": {
    "related": [
      {
        "reference": "GenomicStudy/study-917-trio"
      },
      {
        "reference": "Specimen/spec-917-proband"
      }
    ]
  },
  "derivedFrom": [
    {
      "reference": "DocumentReference/ACC-2026-00917-RAWREADS"
    }
  ]
}

The security label marks the resource Restricted, the type coding gives it a LOINC-recognized document category, context.related ties it to both the GenomicStudy and the Specimen it came from, and derivedFrom records that this BAM was produced from an earlier raw-reads file, preserving the processing lineage an auditor or downstream pipeline might need later.

When validating either example, point your FHIR validator or IG publisher at the specific profile URL under meta.profile rather than the generic DocumentReference base definition. Running validation against the base resource will pass structurally invalid genomics instances, since none of the profile-specific constraints, required security labeling, context.related expectations, get enforced unless the profile URL is explicitly declared and checked.

Practical checklist for producers and consumers

Shipping genomic DocumentReference instances reliably comes down to a short discipline on both sides of the exchange.

For producers: assign stable business identifiers that survive system migrations, compute and record a checksum for every file, declare contentType and content.format precisely rather than defaulting to application/octet-stream out of convenience, and always populate context.related back to the ordering Specimen or GenomicStudy.

For consumers: check securityLabel before caching or forwarding a file, follow derivedFrom and related-resource chains to locate the Observation or GenomicStudy that interprets the file rather than assuming the file itself carries findings, and respect whatever size and access policy the serving system publishes rather than hammering it with naive polling.

One operational judgment call worth making early: for any workflow producing multiple files from a single sequencing run, model the run as a GenomicStudy and attach each file as a DocumentReference under it, rather than publishing loose DocumentReference instances with no shared parent.

GenomicStudy linked to multiple file references

Pro Tip: Run a monthly query for DocumentReference instances with an empty context.related, since that single check catches most orphaned-file problems before a clinician ever notices a broken link.

The most consequential pitfall, worth repeating because it recurs in nearly every genomics FHIR integration review, is using DocumentReference to store the variant calls themselves instead of the file that contains them. The Genomics Reporting IG keeps files and structured observations deliberately separate so that clinical decision support systems can query variants directly without parsing a VCF at runtime; collapsing that separation defeats the purpose of adopting FHIR for genomics in the first place.

What I'd prioritize when implementing this in production

Three things earn disproportionate payoff in a genomics DocumentReference implementation: discoverability, integrity, and privacy, in roughly that order of daily impact. Discoverability means every file resolves back to a Specimen or GenomicStudy; integrity means every file carries a checksum a consumer can actually verify; privacy means security labeling is set before the resource is ever published, not patched in afterward.

On the DocumentReference versus GenomicStudy question, I lean toward reaching for GenomicStudy earlier than most teams expect, specifically once a workflow produces more than one artifact per order. A single VCF from a single assay rarely needs a GenomicStudy wrapper. A pipeline producing raw reads, an aligned BAM, and a final VCF almost always does.

If your team is weighing how these file-level patterns connect to pharmacogenomic reporting and EHR delivery, that is exactly the integration layer SignalPGx works in daily.

— Tarek

How SignalPGx turns FHIR genomic files into physician-ready reports

Getting a Genomics DocumentReference right is only half the integration problem: the other half is turning what it points to into something a clinician can act on without opening a VCF file. SignalPGx is built for that second half. The platform ingests genotype and medication data, including file-level artifacts exposed through HL7/FHIR, and produces physician-reviewed, evidence-graded pharmacogenomic reports without requiring your lab to build and maintain that interpretation layer in house.

SignalPGx

Because the integration runs on standard FHIR resources rather than a proprietary pipeline, the DocumentReference, GenomicStudy, and Observation patterns described above map directly into how SignalPGx connects to your lab's systems and the ordering provider's EHR. Living reanalysis then keeps those reports current as CPIC and FDA guidance evolves, so a report issued today does not go stale the next time a guideline changes. Labs deploying white-label PGx reporting often get a branded reporting service running more quickly than building this integration from scratch.

If your lab is scoping a FHIR-based pharmacogenomics rollout, visit the White-Label PGx Reporting page to see the Pilot, Standard, Scale, and Enterprise plans and start a conversation about your integration timeline.

This article is general information, not a substitute for advice from a qualified doctor. Consult a qualified healthcare professional about your own circumstances before acting on anything here.

FAQ

What is the difference between DocumentReference and GenomicStudy in FHIR genomics?

DocumentReference registers and describes a single genomic file, such as a VCF or BAM, including its location, format, and security metadata. GenomicStudy describes the broader sequencing or analysis event, often spanning multiple files and multiple downstream observations, as outlined in the Genomics Reporting Implementation Guide.

Can DocumentReference be used to store variant data directly?

No. DocumentReference is a file-level metadata wrapper and should never encode structured variant calls, haplotypes, or genotypes; those belong in dedicated Observation-based profiles within the Genomics Reporting Implementation Guide. Keeping files and structured findings separate is what allows clinical decision support systems to query variants without parsing raw files.

What file formats does the Genomics DocumentReference profile support?

The profile is format-agnostic by design, so it supports any genomic file type as long as content.attachment.contentType and content.format correctly describe it, as the DocumentReference R4 specification notes. Common examples include VCF, BAM, CRAM, and BED files, with application/octet-stream typically used for binary formats that lack a registered MIME type.

How should large genomic files be referenced instead of embedded?

Files beyond a small size threshold should be referenced through content.attachment.url pointing to a dedicated file server or object store, rather than embedded as base64 data in attachment.data. The US Core implementation guide notes that inline attachments are only practical for small payloads, with servers expected to document their own size limits for anything larger.

Use context.related to reference the GenomicStudy, Specimen, or GenomicReport associated with the file, and use derivedFrom when the file was produced by transforming another file. The Genomics Reporting Implementation Guide recommends this linkage specifically to preserve provenance and prevent orphaned files that cannot be traced back to an order.

Sources

Key HL7 resources for validation and deeper reading

For machine validation, point your validator at the profile pages directly: the Genomics Reporting Implementation Guide artifacts list the current StructureDefinitions, and the DocumentReference R4 page documents the base resource constraints the profile builds on. For conceptual grounding on why FHIR genomics separates files from observations, the general reporting guidance and a scholarly discussion of FHIR genomics integration are the most useful starting points, alongside HL7's own genomics overview.