FHIR Interoperability Testing: Certified, Not Connected
The following scenario is hypothetical, built from patterns that recur across health IT integration work, not a real patient or a real vendor incident.
A patient is discharged from a hospital after a cardiology admission. Her discharge summary and a reconciled medication list are sent electronically to her primary care practice, which runs a different, independently developed EHR. Both systems hold current ONC certification. Both passed their FHIR conformance test suites during certification testing, including the Inferno-based test kit for the Standardized API criterion. Three weeks later, a nurse at the primary care practice pulls up the patient's chart before a follow-up visit and finds a medication list with two entries where the hospital's discharge record had eleven. Nothing errored. No integration engine logged a failure. No certification body was notified, because nothing about either system's certified conformance had changed.
The eventual root-cause trace (again, hypothetical, but representative of a documented and well-understood class of problem) lands on a single design decision made independently by two engineering teams years apart. The hospital's EHR marks most of the fields on a MedicationStatement resource as optional in its own implementation guide, consistent with the base FHIR specification's philosophy that very few elements should carry a hard cardinality requirement. When a medication was entered through a legacy order-entry workflow rather than the newer structured medication reconciliation tool, the resource that left the hospital's system was missing a data-absent-reason code that the specification permits but does not require. The primary care EHR's ingestion logic, built by a different team against a different but equally valid reading of the same implementation guide, treated any MedicationStatement without a specific status extension it expected as not-current and silently excluded it from the active medication list the nurse actually saw. Neither system violated the specification. Neither system violated its own certification. The data still arrived. It was still readable. It was still wrong, in a way that a certification test could never have caught, because certification testing was never designed to catch it.
This is the argument this article makes at length: passing FHIR conformance testing is proof that a system built a resource with the correct shape. It is not proof, and was never designed to be proof, that two systems can exchange information that means the same thing on both ends. Treating the first as evidence of the second is an expensive and increasingly regulated mistake.
The Certification Illusion
Health IT vendors, buyers, and even some engineering teams inside vendor organizations routinely conflate two different claims. The first claim is "our system conforms to the FHIR specification and the relevant implementation guide." The second is "our system is interoperable with other certified systems." The first claim is testable in a lab, and the health IT industry has built genuinely good tooling to test it. The second claim can only be verified by actually connecting two independently built systems and watching what happens to real or realistic clinical data as it crosses that boundary. Almost nobody does this systematically before go-live.
The confusion is understandable, because the industry's own regulatory and testing infrastructure uses the word "interoperability" to describe the certification programs themselves. The ONC Health IT Certification Program's Standardized API criterion, § 170.315(g)(10), exists specifically to promote interoperability, and Inferno (the ONC-funded, open-source tool used to validate conformance against that criterion) describes itself as a FHIR testing tool built to support exactly this certification goal. None of that is inaccurate. But certification against a specification, however well the specification is designed, is a structural test performed against a reference behavior, not a live test performed against another organization's production system. A system can pass every assertion Inferno runs and still have never once exchanged a resource with the specific EHR, payer platform, or health information exchange it will actually connect to in production.
This distinction matters more in FHIR than it would in a tightly constrained API standard, because FHIR was deliberately built to be loose at its base and tightened only through implementation guides layered on top. The specification says so directly: in most FHIR resources, "very few elements have a minimum cardinality of 1," because resources are meant to function across a wide range of contexts, some of which will have incomplete information available. That flexibility is a genuine strength: it is why FHIR could become a single base standard adopted across radically different care settings. It is also the exact mechanism that allows two conformant, certified implementations to disagree about what a well-formed, clinically complete resource actually contains.
What FHIR Conformance Testing Actually Validates
Before mapping where interoperability breaks, it is worth being precise about what the standard testing tools genuinely check, because the tools themselves are well built and the industry benefits from taking their actual scope seriously rather than either overtrusting or dismissing them.
Inferno runs test kits scoped to specific implementation guides. Its ONC Certification (g)(10) test kit validates conformance to the Standardized API for Patient and Population Services criterion. Its US Core test kit validates server and client behavior against a specific version of the US Core Implementation Guide. Separate kits exist for Da Vinci Prior Authorization Support, Documentation Templates and Rules, Coverage Requirements Discovery, Bulk Data Access, SMART App Launch, and CARIN Blue Button, among others. Each of these is a genuinely useful, automated, repeatable way to check that a server returns resources with the right structure, the right required elements, and the right authorization behavior for a defined use case.
Touchstone, operated by AEGIS.net, plays a similar role for a broader set of FHIR-based implementation guides and is widely used both for ONC certification support and for HL7 Connectathon testing, where implementers bring early-stage FHIR servers together to test against published test scripts in a shared, monitored environment. Both tools are, by design, specification-conformance engines: they compare what a server returns to what a specification or implementation guide says a server should return, under test conditions that the tool itself constructs.
Here is what that actually covers, and what it does not.
| What conformance testing validates | What it does not validate |
|---|---|
| Resource structure matches the profile (correct elements, correct cardinalities, correct data types) | Whether the data inside those elements means the same clinical thing to the system receiving it |
Required (1..1 or 1..*) elements are present |
Whether optional or "must support" elements that carry real clinical weight are actually populated in practice |
| Authorization flows (SMART App Launch, OAuth scopes) behave correctly against the test harness | Whether a specific hospital's identity provider, consent model, or scope-granting workflow behaves the same way in production |
| Terminology bindings match the value set the profile specifies, when the binding is checked | Whether two systems chose the same code from a permissible value set, or the same coding system, for the same clinical concept |
| The server responds correctly to the exact search parameters and interactions the test kit exercises | Whether the server behaves the same way for search parameter combinations, pagination volumes, or data population patterns the test kit never exercises |
| Conformance to a single implementation guide version, in isolation | Behavior when connected to another organization's real system, real patient population, and real data-entry habits, over the full lifecycle of an exchange |
This is not a criticism of Inferno, Touchstone, or the ONC certification program. A specification-conformance test that also tried to guarantee cross-vendor semantic agreement would need to somehow simulate every other certified vendor's implementation choices, which is not a solvable problem for a general-purpose test kit. The industry built the right tool for the job it assigned itself. The failure is organizational, not technical: treating a passed conformance suite as if it answered a question it was never built to answer.
There is also a practical reason conformance tools cannot close this gap even if a future version tried to. Inferno's test kits, and Touchstone's test plans, work by scripting a specific, known test patient or a small fixed set of test fixtures against your server, then checking the response against the profile. That is exactly the right design for a repeatable, automatable pass/fail check, and it is also, by construction, a test against a handful of scripted scenarios rather than against the messy distribution of real clinical documentation a production system actually produces. A server can return a perfect result for Inferno's specific test patient, built once, carefully, to exercise the must-support elements the test kit checks, while the same server's real production data, generated by hundreds of clinicians using dozens of different order sets and legacy workflows, populates those same must-support elements far less consistently. The test kit is not wrong about what it measured. It measured one server's best-case behavior against a small number of engineered scenarios, and reported that behavior accurately. Nothing in that result describes the server's average-case behavior against its own live population, let alone another organization's.
Five Places Certified, Compliant Systems Still Diverge
The specification's own flexibility, not carelessness on any single vendor's part, is what produces the recurring failure modes below. Each one is legal under the specification, each one is compatible with certification, and each one has caused real production incidents across the industry.
Must-support elements and the absence-of-data escape hatch
US Core and most other US-market implementation guides use a mustSupport flag to mark elements that a conformant server is obligated to populate when the underlying system holds the data and to handle correctly when it does not. This sounds like it should close the optionality gap. In practice, the specification is explicit that must-support obligations describe how a system behaves, not that data will always be present. The US Core Implementation Guide's own guidance states that when information for a data element is not present and the reason is unknown, a US Core server "SHALL NOT" include that element in the resource. It also separately obligates client systems to interpret a missing must-support element simply as "data not present," gracefully, without erroring.
Both halves of that rule are reasonable engineering guidance on their own. Combined, they create a specification-sanctioned path for a clinically important field (an allergy severity, a medication status, an encounter reason) to be silently absent from a resource that is nonetheless fully conformant, fully must-support-compliant, and fully certifiable. The receiving system is required to treat that absence as "no data," not as "an error," which is exactly correct from a conformance standpoint and exactly the mechanism by which real information can vanish between two certified systems without either one doing anything the specification prohibits.
Consider two simplified, illustrative MedicationStatement resources built from the same clinical fact (a patient currently taking a maintenance medication) but produced by two different order-entry paths inside the same hospital's own EHR. Both are structurally valid US Core resources:
{
"resourceType": "MedicationStatement",
"status": "active",
"medicationCodeableConcept": { "text": "Example maintenance medication" },
"subject": { "reference": "Patient/example" },
"effectiveDateTime": "2026-01-14"
}
{
"resourceType": "MedicationStatement",
"status": "unknown",
"medicationCodeableConcept": { "text": "Example maintenance medication" },
"subject": { "reference": "Patient/example" }
}
The first resource, produced through the hospital's structured reconciliation tool, carries a clear status and an effective date. The second, produced through a legacy free-text order path, carries a status of unknown and no effectiveDateTime at all: both legal values under the specification, since status is required but unknown is a permitted code, and effectiveDateTime is not mandatory. A receiving system's ingestion logic, written to treat status: unknown as insufficiently reliable to include in an "active medications" view, will quietly drop the second resource from the list a clinician sees, while a different receiving system, written with a more permissive rule, would show it. Neither ingestion rule is a specification violation. The difference in what the clinician actually sees is entirely a function of implementation choices that certification testing never compares against each other, because certification testing only checks each side against the specification, never against the other side's interpretation of it.
Extension proliferation for the same clinical concept
FHIR's extension mechanism exists because, as the specification itself puts it, "specific implementations have valid requirements that are not part of these agreed common requirements." Extensions let an implementer capture something the base resource and profile do not model, without waiting for a future specification version. That is a deliberate and valuable design choice. It also means two organizations can each build a perfectly legitimate extension for the same underlying clinical idea (a care-preference flag, a locally significant risk score, a state-specific consent indicator), using different URLs, different value sets, and sometimes different data types, with neither one aware the other exists.
The specification's guidance for handling unrecognized extensions is that applications "should not reject resources merely because they contain extensions" they don't understand, which is the right default for keeping data flowing. It also means a receiving system can correctly, conformantly, and silently ignore an extension that the sending system considered clinically important, because from the receiving system's point of view, an unrecognized extension is indistinguishable from a harmless local annotation. Nothing errors. Nothing gets flagged. The information is present in the payload and absent from the workflow.
An illustrative version of this pattern is easy to construct. Two organizations both want to flag that a patient has declined a specific category of screening, a real clinical workflow need the base Observation resource does not model directly. Organization A defines its own extension for this:
"extension": [{
"url": "http://example-hospital-a.org/fhir/StructureDefinition/screening-declined-reason",
"valueCodeableConcept": { "text": "Patient declined, documented verbally" }
}]
Organization B, solving the identical clinical problem independently, defines a different extension:
"extension": [{
"url": "http://example-hospital-b.org/fhir/StructureDefinition/patient-refusal-flag",
"valueBoolean": true
}]
Both are legitimate, specification-compliant extensions. If Organization A sends a resource carrying its extension to Organization B, Organization B's system has no reason to recognize the URL, no obligation to reject the resource because of it, and, per the specification's own guidance, every reason to simply pass the resource through unchanged while ignoring the extension entirely. The clinical fact the sending system considered important enough to encode is present in the transmitted data and invisible to anyone using the receiving system, and no conformance test run against either organization in isolation would ever surface this, because each organization's server behaves exactly as its own profile says it should.
Terminology and value-set binding mismatches
FHIR profiles bind coded elements to value sets with varying binding strength (required, extensible, preferred, or example), and even a "required" binding only constrains which code system and which permissible codes are valid, not which of several clinically equivalent codes a given system will choose. A condition, an allergy, or a lab result can be coded correctly, using a code from the correct value set, by two different systems that nonetheless picked different codes for what a clinician would consider the same concept: a local lab's internal analyte code mapped to one LOINC code by one system's terminology service and to a closely related but distinct LOINC code by another. Both codes may be individually valid. Downstream decision support, allergy-interaction checking, or population-health queries built against one code will not match against the other. This is not a bug in either terminology service. It is the ordinary, expected behavior of independently maintained code-mapping tables applied to a genuinely ambiguous real-world source value, and no conformance test checks whether two systems' mapping tables agree with each other, because that would require the test tool to know what "agree" means for every possible clinical concept in scope.
An illustrative example, using clearly fictional code values rather than real terminology identifiers: a regional laboratory reports a potassium result under an internal analyte code specific to its instrument vendor. Hospital A's terminology mapping service, maintained by its own informatics team, translates that internal code to one code from an appropriate standard code system. Hospital B, receiving the same lab feed through a different integration engine with its own independently maintained mapping table, translates the identical internal code to a different, closely related code from the same code system, perhaps one that distinguishes serum from plasma specimen type where the source system's own documentation was ambiguous about which applied. Both codes are individually valid members of the value set the relevant profile specifies. A clinical decision support rule at Hospital A that checks for elevated potassium using its expected code will not fire on a resource carrying Hospital B's code, and vice versa, even though both resources describe the same lab result for the same patient. Conformance testing validated that each code belongs to a permitted value set. Nothing validated that the two mapping choices agree with each other, because no general-purpose conformance tool can encode "these two codes describe the same real-world lab test" for every ambiguous source value in a health system's terminology.
Date, time, and time zone handling inconsistencies
FHIR's dateTime and instant data types support multiple levels of precision (year only, year and month, full date, full date and time with a time zone offset), and a profile can legitimately accept any of these depending on what the source system actually captured. A conformance test typically validates that a returned value is a syntactically valid FHIR date or instant; it does not typically validate that a system consistently applies a time zone offset, correctly interprets a partial date supplied by another system, or correctly orders events across resources that mix precision levels. In a lab-result exchange between a facility on Eastern time and a receiving system that defaults to UTC when no explicit offset is present, or between a system that stores an admission date with only day-level precision and one expecting full date-time precision for chronological sorting, resource timestamps can be individually valid and collectively produce an incorrect clinical timeline once assembled on the receiving end: a lab value that appears to have resulted before the specimen was drawn, or a medication administration that appears to precede the order that authorized it.
An illustrative pair of values makes the mechanism concrete. A specimen collection event recorded at a facility in the Eastern time zone might carry "effectiveDateTime": "2026-03-14T22:10:00-04:00". A batch interface at a different facility, feeding results into the same longitudinal record overnight, might submit a related result carrying only "effectiveDateTime": "2026-03-15", a legal, lower-precision value under the specification, because the source system never captured a time component for that particular interface. A receiving system's sort logic, if it treats a date-only value as midnight UTC by default rather than flagging it as lower precision, can order that second result before the first even though the underlying events happened in the reverse sequence. Both values are individually valid FHIR dateTime instances. Neither system violated its profile. The assembled timeline a clinician or a downstream analytics query sees is simply wrong, and no single-resource conformance check was ever positioned to catch an ordering error that only exists once multiple resources are combined.
Patient-matching failures across systems
None of the preceding failure modes matter if the two systems cannot first agree that they are talking about the same patient, and this is arguably the most consequential and best-documented interoperability gap in the industry, sitting entirely outside what any FHIR conformance test evaluates. The U.S. Government Accountability Office's review of health IT interoperability found that inaccurate, incomplete, or inconsistently formatted demographic data, such as a hyphenated surname recorded differently by two intake systems, a Social Security number present in one record and absent in another, or a preferred name used instead of a legal name, routinely undermines automated patient matching between systems, and that no single fix, including any specific software product, resolves the problem on its own (GAO-19-197). ONC's own explanation of patient matching makes a related point directly relevant to conformance testing: there is no industry-accepted minimum performance baseline or standardized testing approach for patient-matching algorithms across systems, which means two certified EHRs can each run a technically sound matching algorithm and still fail to link the same patient's records when the demographic inputs differ even slightly between the systems that captured them. A MedicationStatement resource with a perfectly complete must-support profile, correctly coded terminology, and correctly formatted timestamps is clinically useless if it gets attached to the wrong patient record, or never gets matched to any record at all and sits stranded as an orphaned result.
An illustrative near-miss shows how little variance it takes to break automated matching. Two systems each hold a record for what is, in fact, the same person, captured at different intake points:
| Field | System A record | System B record |
|---|---|---|
| Legal name | Katherine J. Alvarez-Reyes | Katherine Alvarez Reyes |
| Date of birth | 1987-03-04 | 1987-03-04 |
| Sex | F | F |
| Address on file | 118 Birchwood Ln, Apt 4 | 118 Birchwood Lane |
| Identifier present | State driver's license number | None captured |
A deterministic matching algorithm requiring an exact string match on name and address, common in older integration engines, will not link these two records: the hyphenation and the apartment-number formatting differ enough to fail an exact comparison, even though every human reviewer would immediately recognize them as the same patient. A probabilistic matching algorithm, weighting partial similarity across multiple fields, might link them correctly, or might not, depending on its configured threshold and which fields it weights most heavily. Neither system's FHIR resources were malformed. Neither system's matching algorithm was defective by any absolute standard; each performed within the range ONC's own research describes as normal for current patient-matching technology, where, as ONC has noted publicly, there is no industry-accepted minimum performance baseline or standardized testing approach that would let a buyer compare one algorithm's real-world accuracy against another's before deployment. The practical consequence is identical either way: a clinician working from System B's record has no way of knowing System A holds additional history for the same person, because from System B's point of view, no additional history exists to request.
The Risk Map
The table below consolidates the failure modes above into a single reference for engineering and QA leaders assessing where their own integration testing coverage actually stops.
| Failure mode | What conformance testing checks | What it misses | How to actually test for it |
|---|---|---|---|
| Optional / must-support field variance | That must-support elements are structurally correct when present, and that absence doesn't break the schema | Whether the field is populated often enough, and consistently enough, to be clinically usable in practice | Run a statistically meaningful sample of real or realistic resources from the actual trading partner's sandbox and measure population rates per must-support field, not just presence/absence on a handful of test-kit fixtures |
| Extension mismatch for the same concept | That any extensions present use valid URLs and valid data types | Whether the receiving system's business logic actually reads and acts on a given extension, versus silently ignoring it | Trace a specific extension from the sending system's payload through the receiving system's downstream workflow (clinical decision support, chart display, analytics) and confirm it surfaces somewhere a human or a rule engine can act on it |
| Terminology / value-set mismatch | That codes belong to the value set the profile specifies | Whether two systems' code-mapping tables produce the same or clinically equivalent codes for the same real-world source value | Build a terminology reconciliation test set of source values known to be ambiguous in your domain (local lab codes, legacy problem-list text) and compare both systems' mapped output side by side |
| Date / time zone inconsistency | That timestamps are syntactically valid FHIR date/time values | Whether chronological ordering survives once resources with mixed precision and mixed time zone handling are assembled together | Construct a synthetic timeline test that deliberately mixes precision levels and time zones and verify the receiving system reconstructs the correct clinical sequence, not just valid individual timestamps |
| Patient-matching failure | Nothing. Patient matching sits outside the FHIR resource-conformance layer entirely | Whether the two systems will reliably resolve the same real person to the same patient record given realistic, imperfect demographic data | Run a dedicated matching test set with deliberately imperfect but realistic demographic variants (nicknames, hyphenation, missing identifiers) against both systems and measure match and mismatch rates directly |
When an Interoperability Gap Becomes a Regulatory Problem
Until 2016, a vendor whose systems technically conformed but practically failed to exchange usable data had, at most, a reputational and contractual problem. That is no longer the whole picture. The 21st Century Cures Act and ONC's Cures Act Final Rule created a federal information-blocking framework, codified at 45 CFR Part 171, that prohibits practices likely to interfere with the access, exchange, or use of electronic health information, subject to a defined set of exceptions. The rule applies to three categories of actors: health care providers, health IT developers of certified health IT, and health information exchanges or health information networks. The knowledge standard differs by actor: ONC's own summary frames the health-IT-developer standard as whether the actor "knows, or should know" that a practice interferes with EHI access, exchange, or use, a materially different and generally stricter bar than the "knows... is unreasonable" standard applied to providers.
What follows is professional analysis of how this framework intersects with the interoperability gaps described above, not a legal conclusion, and any organization facing an actual information-blocking inquiry should get its own regulatory counsel rather than rely on an engineering blog post. With that caveat stated plainly: the information-blocking framework was built around practices, not around isolated bugs, and a single instance of a must-support field going unpopulated because of a legacy workflow is not, on its own, the kind of thing the rule was designed to capture. But the "know, or should know" standard for health IT developers creates a meaningfully different exposure once a specific interoperability gap has been identified, reported by a trading partner, and left unaddressed. A vendor that becomes aware, through a support ticket, a partner escalation, or its own QA process, that its handling of a specific must-support element, extension, or terminology mapping is systematically suppressing data for a defined class of resources, and continues shipping that behavior without a published, defensible technical or practical justification, has moved from "conformant implementation choice" toward the territory the knowledge standard was built to address. The distance between "our system passed certification" and "we knew about this gap and left it in place" is exactly the distance the regulation is designed to close, and it is a distance that a bilateral testing program, described below, is specifically built to shrink before it becomes a live question.
There is a second, more immediate practical stake that does not require any information-blocking analysis at all: a hospital or payer that experiences repeated interoperability failures with a certified vendor's product has every incentive, contractual and reputational, to escalate that failure publicly, to a regulator, or to a certifying body, regardless of whether the underlying behavior technically qualifies as information blocking. The CMS Interoperability and Patient Access Final Rule (CMS-9115-F) compounds this by requiring Medicare Advantage, Medicaid, CHIP, and federally facilitated exchange qualified health plans to expose FHIR-based Patient Access and provider directory APIs built on the same content and vocabulary standards ONC finalized under the Cures Act rule, which means payer-side FHIR APIs sit under direct CMS oversight in addition to whatever certification status the underlying health IT carries. A vendor's interoperability gap does not need to be adjudicated as information blocking to become a business problem; it only needs to become visible to a partner with regulatory leverage.
It is worth being precise about where the regulatory overlap actually sits, because it is easy to overstate. The information-blocking framework governs practices that interfere with EHI access, exchange, or use: it was built primarily around deliberate or negligent restriction, not around good-faith implementation variance between two independently engineered systems each trying to follow the same specification. A must-support field that goes unpopulated because of a legacy order-entry path is, in isolation, closer to an engineering defect than to the kind of conduct the rule was written to target. The relevant shift happens only once an organization has specific knowledge of a specific, recurring gap and takes no action, which is exactly why the testing framework in this article matters as much for the paper trail it produces as for the defects it catches. A documented bilateral test program that identified and remediated a gap before go-live is close to the opposite of information blocking; an identical gap discovered only through a partner complaint, with no record of prior testing, sits in a meaningfully different position under the same facts.
TEFCA and the Direction of Travel
The regulatory trend line points toward more exchange, at greater scale, between more heterogeneous systems, not less. The Trusted Exchange Framework and Common Agreement (TEFCA), operated under ONC's authority, establishes a "network-of-networks" model in which Qualified Health Information Networks (QHINs) connect to each other under a shared Common Agreement, so that a participant connected to one QHIN can, in principle, reach data held by any organization connected to any other QHIN. The first QHINs were designated in December 2023, and data has continued to flow and expand since, including work extending TEFCA's use to purposes such as government benefits determination.
TEFCA's entire value proposition is that a healthcare organization does not need a bespoke bilateral connection with every other organization it might need to exchange data with: it connects once, to a QHIN, and inherits reach across the network. That is a genuine reduction in integration burden. It is also, from an interoperability-testing standpoint, a multiplier on exactly the risk this article describes. A bilateral connection between two organizations that have tested against each other directly at least has the option of catching must-support gaps, extension mismatches, and terminology divergence specific to that pair. A query answered through a QHIN network-of-networks model may originate from an organization the responding system's engineering team has never heard of, running an EHR they have never tested against, with implementation choices they have no visibility into until the data arrives. The more the industry succeeds at TEFCA's stated goal of making exchange the default rather than the exception, the more often certified systems will be exchanging data with certified systems they have never validated against directly, which makes the layered testing approach below more relevant with each new QHIN connection added to the network, not less.
Beyond Conformance: A Four-Layer Testing Framework
Conformance testing is a necessary floor, not a finished testing program. The framework below is built specifically for teams that already run Inferno or Touchstone as part of certification or pre-certification work and need to know what to add, in what order, before trusting a production connection.
Layer 1 — Specification conformance testing. Run the appropriate Inferno test kit or Touchstone test plan against your own server and, wherever possible, against your client's outbound requests. This layer catches structural defects (wrong cardinality, wrong data type, missing required elements, broken authorization flows) cheaply and early, and there is no reason to skip it or to treat the layers below as a replacement for it. Its output is a pass/fail against a specification, run in a controlled environment your own team built or configured.
Layer 2 — Bilateral testing with the actual trading partner. Before a production connection with a specific hospital system, payer, or health information network goes live, run a dedicated test cycle against that partner's own sandbox or test environment, using resources built from real workflows on both sides rather than the conformance tool's generic fixtures. The goal of this layer is narrow and specific: identify where this partner's implementation choices (their must-support population patterns, their extensions, their terminology mappings, their time zone handling) differ from your own, before either side is exchanging live patient data. This is the layer most organizations skip entirely, on the reasonable-sounding but incorrect assumption that two certified systems talking to each other is equivalent to two systems that have actually talked to each other.
Layer 3 — Synthetic-patient lifecycle testing across the full exchange, not a single request-response pair. A single successful GET against a Patient or Observation endpoint proves almost nothing about how an exchange behaves over time. Build synthetic patients (clearly fabricated, non-real identities constructed specifically for testing) and run them through a full realistic lifecycle: initial registration and identity resolution, an encounter, a referral to a specialist on the other system, a diagnostic result returned, a medication reconciliation, and a discharge or care-transition summary. Test what happens when the same synthetic patient's data accumulates on both sides over multiple encounters, not just what a fresh single query returns. Many of the failure modes described earlier, such as timestamp ordering across multiple encounters, must-support fields that are populated on a first encounter but dropped on a subsequent update, and extensions that get attached once and never referenced again, only become visible across a sequence of exchanges, not a single transaction.
Layer 4 — Testing against more than one certified sandbox, not only a single reference server. A test program that validates only against Inferno's own reference server, or only against one partner's sandbox, is really validating against one specific set of implementation choices, however conformant those choices are. The major certified EHR platforms each expose developer or partner sandbox environments with materially different data population patterns, different default behavior for must-support fields, and different extension conventions from each other and from Inferno's reference implementation. A vendor building a product meant to connect to multiple hospital systems or payers should run the same synthetic-patient lifecycle test from Layer 3 against at least two independently built certified sandboxes and compare the results directly. Differences that show up in that comparison are, by definition, exactly the class of interoperability gap that a single-sandbox test program would never surface.
Healthcare Interoperability Testing Checklist
The checklist below is built specifically around the four layers above, meant as a pre-go-live gate rather than a general FHIR testing reminder list.
- Run the applicable Inferno or Touchstone conformance test kit against your own server and confirm a clean pass on the current implementation guide version you intend to support in production.
- Obtain sandbox or test-environment access from the specific trading partner you are about to connect to, not a generic reference server, and confirm it is populated with data patterns representative of that partner's real production workflows.
- Build a must-support population audit: for every element your profile marks must-support, query a realistic sample of resources from the partner's sandbox and record the actual population rate, not just whether the schema allows the field to be absent.
- Catalog every extension either system uses for the resource types in scope, and confirm each one is read and acted on somewhere in the receiving system's workflow, not merely accepted without erroring.
- Build a terminology reconciliation test set from your domain's known-ambiguous source values and compare the codes both systems produce for the same underlying concept.
- Construct a mixed-precision, mixed-time-zone timestamp test and verify chronological ordering survives once resources from both systems are assembled into a single patient timeline.
- Run a dedicated patient-matching test set with realistic imperfect demographic variants and measure both false-non-match and false-match rates before trusting automated linkage in production.
- Build synthetic patients and run them through a full exchange lifecycle (registration, encounter, referral, result, medication update, care transition) rather than testing single request-response pairs in isolation.
- Repeat the synthetic-patient lifecycle test against a second, independently built certified sandbox if your product will connect to more than one EHR or payer platform, and compare results directly.
- Document every implementation-guide-permitted variance discovered during bilateral testing, including the business or clinical impact if left unaddressed, so the organization has a defensible record of active remediation rather than a known, unaddressed gap.
Two Further Examples
A hypothetical payer-side scenario. A regional health plan builds its CMS-9115-F Patient Access API on a FHIR server that passed its (g)(10)-equivalent conformance testing without issue. A digital health app used by the plan's members to aggregate their own records pulls claims-derived ExplanationOfBenefit resources correctly, but the app's engineering team discovers, months after launch, that a meaningful share of prior-authorization-linked claims are missing a CodeableConcept extension the payer's system uses internally to flag denial reason categories, an extension the app never knew existed and had no reason to expect, because nothing in the base profile required it and the conformance test never exercised that specific data population pattern. Members see claims with no visible denial reason, support tickets rise, and the root cause is not a conformance failure on either side but an extension that existed, was structurally valid, and was never tested end to end because nobody owned the responsibility of confirming the receiving app actually surfaced it.
A hypothetical multi-EHR scenario. A specialty telehealth SaaS platform integrates with two different hospital systems' EHRs to pull recent lab results before a virtual visit. Against the first hospital's sandbox, everything works: lab values arrive with full-precision timestamps, and the platform's chronological result view displays correctly. Against the second hospital's production environment, six months after go-live, a support escalation reveals that a subset of lab results, specifically those entered through an overnight batch interface rather than the hospital's real-time order-entry system, arrive with date-only precision instead of full date-time. The platform's display logic, tested only against the first hospital's full-precision pattern, silently defaults missing time components to midnight, which occasionally reorders same-day results incorrectly on the clinician-facing timeline. Both EHRs are certified. Both resources are valid FHIR. The platform had only ever run its synthetic-patient lifecycle test against one sandbox, which is exactly the Layer 4 gap the framework above is built to close.
A Closer Look at the Medication List That Shrank
It is worth returning to the opening scenario with more technical detail, because the general shape of that failure (again, a hypothetical composite, not a real incident) recurs across enough real production interoperability work to be worth walking through step by step, using the structure that makes any interoperability failure analyzable: the initial situation, the assumption nobody stated out loud, the specific technical cause, the consequence, the decision the organization actually faces, and the better approach.
The initial situation. A hospital's certified EHR sends a discharge medication list to a primary care practice's certified EHR through a standard FHIR-based document exchange, built on MedicationStatement resources referencing the same patient and encounter.
The hidden assumption. Both engineering teams assumed that "certified" meant "tested against how the other side's system actually behaves." Neither team had, in fact, ever run a synthetic-patient lifecycle test against the other's real sandbox. Each had only run the standard conformance test kit against its own server, in isolation, during its own certification cycle, potentially years apart, against potentially different implementation guide versions.
The technical cause. The hospital's structured medication reconciliation workflow reliably populates status, effectiveDateTime, and a specific dosage extension on every MedicationStatement it creates. A smaller share of medications, entered through an older order-entry path retained for a specific specialty workflow, produces resources with status: unknown and no effectiveDateTime, exactly as shown in the earlier code example, both legal under the specification. The primary care practice's ingestion logic, built independently and tested only against the hospital's structured-workflow output during an earlier, narrower pilot, treats any resource with status: unknown as unreliable and excludes it from the active medication list rendered in the chart, without erroring or logging a rejection anyone would notice.
The consequence. Nine of eleven medications from the structured workflow display correctly. Two, entered through the older order path, are silently absent from what the receiving clinician sees, with no indication that anything was dropped, no discrepancy count, and no alert distinguishing "patient takes eleven fewer medications" from "two resources were filtered out by a business rule neither team documented to the other."
The decision the organization actually faces. Once discovered, typically through a clinical near-miss or a sharp-eyed nurse rather than through any automated signal, the organization has to decide whether this is an isolated bug to patch quietly or a signal that the entire integration was never tested at the level of rigor its certification status implied. Patching only the specific rule that dropped status: unknown resources fixes this one instance and leaves the underlying gap, meaning no bilateral testing program, no synthetic-patient lifecycle testing, and no must-support population audit, completely intact for the next divergence neither team has discovered yet.
The better approach. A Layer 2 bilateral test cycle, run before go-live, using synthetic patients built to include both the structured and legacy order-entry paths on the sending side, would have surfaced this specific gap in a test environment rather than a live chart. A must-support population audit, run against a realistic sample of the hospital's actual production-pattern data rather than a single conformance-tool fixture, would have shown that a meaningful share of MedicationStatement resources carry status: unknown, prompting the receiving team to ask, before go-live rather than after, exactly how their own ingestion logic handles that value.
Comparing the Testing Layers
| Layer | Primary question it answers | Typical tooling | What a clean result actually proves |
|---|---|---|---|
| 1. Conformance testing | Does this server produce structurally valid resources against the specification and implementation guide? | Inferno, Touchstone | The system can pass certification; nothing about a specific real-world exchange |
| 2. Bilateral partner testing | Does this specific trading partner's implementation agree with ours on optional fields, extensions, and terminology? | Partner sandbox, shared test data | This particular pair of systems has reconciled its known implementation differences |
| 3. Synthetic-patient lifecycle testing | Does data stay complete, correctly ordered, and correctly linked across a full multi-step exchange, not just one request? | Synthetic patient generators, custom test harnesses | The exchange behaves correctly over time and across resource types, not only on a first query |
| 4. Multi-sandbox testing | Does our product behave correctly across more than one certified implementation, or only the one we happened to test against? | Two or more independent certified sandboxes | The product generalizes across real-world implementation variance rather than one vendor's specific choices |
Metrics That Actually Tell You Something
Engineering and QA leaders reporting integration health upward tend to reach for the metrics certification tooling already produces, because those numbers are easy to get and easy to put in a slide. Most of them are close to useless for judging real interoperability risk, and a few less obvious ones carry most of the actual signal.
| Metric | Why it is misleading, or why it is useful |
|---|---|
| Percentage of Inferno or Touchstone test assertions passed | Misleading as a production-readiness signal. A 100 percent pass rate describes structural conformance against a fixed set of scripted fixtures, not behavior against a partner's real data distribution. Treat it as a certification gate, not a go-live gate. |
| Number of certified implementation guides supported | Misleading on its own. Supporting more implementation guides expands surface area for divergence without saying anything about how well any single connection has been validated against a real partner. |
| Must-support field population rate, measured against a specific partner's real or sandbox data | Useful. Directly measures whether a field the specification calls clinically important is actually showing up often enough to be relied on for that specific connection, rather than only in test fixtures. |
| Extension recognition rate: the share of extensions present in inbound data that the receiving system's business logic actually reads and acts on | Useful and rarely tracked at all. A low number here is a direct, quantifiable measure of the extension-proliferation gap described earlier. |
| Terminology mapping concordance rate between two systems for a defined set of ambiguous source values | Useful. Requires deliberately building a reconciliation test set, but produces a concrete number (the share of source values both systems map to the same or clinically equivalent target code) that a generic conformance score can never produce. |
| Patient match rate and mismatch rate against a realistic, imperfect demographic test set | Useful, and arguably the single highest-leverage metric on this list, since every other metric is worthless if the underlying patient link is wrong. |
| Count of production support tickets attributable to missing or misinterpreted data after go-live | Useful as a lagging indicator, and worth tracking specifically as its own category rather than folding it into generic "integration issues," since it is the clearest downstream signal that Layer 2 or Layer 3 testing was skipped or insufficient. |
The pattern across this table is not that certification metrics are worthless, since a clean conformance pass is still a real prerequisite, but that every metric worth reporting to a CTO or product leader as evidence of production readiness has to be measured against a specific trading partner's actual data, not against a conformance tool's own fixtures.
A Sample Cost Comparison: Bilateral Testing Versus Discovering the Gap in Production
The figures below are an illustrative estimate built for this article, not a benchmark from any published study or real QAtronic engagement, and any organization applying this kind of model should build its own estimate from its own engineering rates and incident history rather than treat these numbers as universal.
Assume a mid-sized digital health company is about to connect to a new hospital trading partner and is deciding whether to fund a scoped Layer 2 and Layer 3 test cycle (bilateral sandbox testing plus a synthetic-patient lifecycle test) before go-live, at an estimated two to three engineer-weeks of effort, versus skipping it and relying on certification alone.
| Cost category | Proactive bilateral and lifecycle testing (illustrative estimate) | Discovering the gap after go-live (illustrative estimate) |
|---|---|---|
| Direct engineering time | 2–3 engineer-weeks to build synthetic patients, run the partner sandbox cycle, and document findings | 1–2 engineer-weeks of root-cause investigation once a discrepancy is reported, often across two organizations' teams rather than one |
| Clinical or informatics review | A few hours of clinical informatics review to confirm test scenarios reflect real workflows | A formal clinical safety review once a real discrepancy affecting patient data is confirmed, typically involving more senior staff time and more documentation overhead |
| Partner relationship cost | Minimal. Most partners expect and welcome a pre-go-live test cycle | A support escalation that damages trust with the trading partner and can trigger a broader review of the entire integration, not just the specific defect found |
| Remediation cost | Fixes are made and re-tested before any real patient data is affected | Fixes must be made under time pressure, potentially requiring a data correction or backfill process for records already affected in production |
| Regulatory exposure | Essentially none. A documented pre-go-live test program is itself evidence of good-faith diligence | Non-zero once a partner or provider raises the issue as a possible information-blocking concern, even if it is ultimately resolved without a formal finding |
The illustrative comparison is not meant to produce a precise return-on-investment figure, since real organizations will have very different engineering costs, partner relationships, and risk tolerance. The directional point is durable regardless of the specific numbers: the cost of finding this class of gap before go-live is bounded and mostly falls on your own engineering calendar, while the cost of finding it after go-live is open-ended and increasingly falls on people who were not in the room when the original implementation decision was made.
Ownership: Who Actually Tests for This
A common reason bilateral and lifecycle testing gets skipped is not that anyone decided against it: it is that no single role clearly owns it. Certification testing usually has an obvious owner, because certification is a discrete, deadline-driven project with a named regulatory outcome. Bilateral partner testing and full-lifecycle synthetic testing tend to fall into a gap between teams that each have a partial, reasonable claim to a piece of it.
| Responsibility | Most natural owner | Common failure pattern when ownership is unclear |
|---|---|---|
| Running and maintaining conformance test kits (Inferno, Touchstone) | Integration engineering or a dedicated interoperability engineering function | Treated as a one-time certification project rather than a regression suite re-run on every implementation guide version bump |
| Negotiating sandbox access with a specific trading partner | Product or partnership management, with engineering support | Access gets requested only after a production issue is already reported, rather than before go-live |
| Defining must-support population expectations and terminology mappings | Clinical informatics, in collaboration with engineering | Engineering assumes a field will be populated because the specification allows it to be, without confirming actual clinical workflow behavior |
| Building and maintaining synthetic-patient lifecycle test scenarios | QA or a dedicated integration-testing function | Built once for the first partner and never adapted or re-run for subsequent partners with different data patterns |
| Monitoring production support tickets for interoperability-specific patterns | Support or customer success, escalated to engineering | Tickets get resolved case by case without anyone aggregating them into a pattern that points back to a specific untested gap |
| Documenting known, accepted implementation variance for regulatory defensibility | Compliance or legal, informed by engineering's own findings | No one documents that a gap was known and evaluated, which is precisely the condition that increases exposure under the information-blocking knowledge standard described earlier |
Assigning each row above to a named function before the next integration begins is a cheaper fix than any of the testing layers themselves, because it is the reason the testing layers get funded and staffed in the first place.
Warning Signs an Organization Is Trusting Certification Too Much
None of these signs individually proves a problem exists. Together, they describe an organization that has quietly substituted a certification event for an ongoing testing discipline.
- The only FHIR test suite anyone can name is the one used during ONC certification, and it has not been re-run since.
- New trading partner connections go live with a project plan that includes a security review and a data-use agreement but no dedicated interoperability test cycle against that specific partner's sandbox.
- Support tickets describing missing or incorrect clinical data get triaged as one-off data-quality issues rather than checked against a known list of must-support fields, extensions, and terminology mappings that have caused problems before.
- No one on the engineering or QA team can say, without checking, which specific implementation guide version and which specific partner sandbox the last synthetic-patient lifecycle test was run against.
- The organization's only FHIR test environment is a single reference server, often the certification test kit's own reference implementation, rather than more than one independently built certified sandbox.
- Conversations about interoperability risk default immediately to "we're certified," as though certification status answers the question being asked rather than a narrower, related one.
Questions Executives Should Ask Before the Next Integration Goes Live
A CTO, VP of Engineering, or product leader who wants a real answer rather than a reassuring one can ask a small number of specific questions that certification status alone cannot answer:
- Which specific trading partner's sandbox, not a generic reference server, was this integration tested against before go-live, and when?
- What is the current must-support field population rate for the resource types this integration depends on, measured against that partner's real or sandbox data, not against our own test fixtures?
- Has anyone catalogued every extension either side uses for these resource types, and confirmed the receiving system actually acts on each one rather than silently accepting and ignoring it?
- What happens, specifically, when this integration receives a partial or lower-precision timestamp, and has that behavior been tested rather than assumed?
- What is our current patient match and mismatch rate against a realistic, imperfect demographic test set, and who owns improving it if the number is not acceptable?
- If this specific interoperability gap were reported by a partner tomorrow, do we have a documented record showing when we became aware of it and what we did, or would that record start today?
What This Means for Different Organizations
A digital health startup building its first FHIR integration against a single hospital partner can reasonably treat Layers 1 through 3 as close to mandatory before go-live and defer Layer 4 until it has a second partner in its pipeline, since the marginal cost of a second sandbox is not justified until there is a second real connection to protect. A health IT vendor selling a product meant to connect to many different certified EHRs across many customers has a much weaker excuse for skipping Layer 4: every new customer is effectively a new, untested implementation pairing, and the vendor is the only party positioned to catch cross-implementation variance before it reaches a customer's production environment. A payer building CMS-9115-F-mandated APIs sits in a slightly different position again, because its API's consumers are largely third-party apps it does not control and may never test against directly, which makes Layer 3's full-lifecycle synthetic testing, run against the payer's own claims and eligibility data patterns, the highest-leverage layer available, since bilateral testing with every possible consuming app is not realistic at that scale.
Enterprise health systems integrating many inbound feeds from many affiliated practices and referral partners face the inverse problem from the vendor case: the volume of trading partners makes true bilateral testing with every partner impractical, which is exactly the scenario TEFCA-style network exchange is meant to address, and exactly why Layer 3's lifecycle testing needs to be automated and repeatable enough to run against new partners quickly rather than treated as a one-time manual exercise reserved for the largest connections.
A useful way to think about the budget trade-off across these organization types is in terms of what each is actually protecting. A startup protecting a single early hospital relationship is protecting a reference customer whose trust, once damaged by a visible data gap, is disproportionately expensive to rebuild relative to the company's size, which argues for front-loading bilateral testing even before the product has the volume to justify a dedicated integration-testing function. A vendor protecting dozens of customer relationships across many EHR platforms is protecting a portfolio, where the marginal cost of automating Layer 3's lifecycle test once and re-running it against each new sandbox is small relative to the cumulative cost of re-discovering the same class of gap independently at each customer. A payer or large health system protecting a regulatory relationship with CMS or ONC is protecting something closer to a license to operate, where the cost of a documented, repeatable testing program is better understood as a fixed compliance cost than as a discretionary engineering investment, because the alternative, an undocumented gap discovered externally, is precisely the condition the information-blocking knowledge standard treats least favorably.
The Distinction This Article Is Not Making
None of the above should be read as an argument that certification testing is a waste of effort, that Inferno or Touchstone are poorly designed, or that ONC's certification program has failed at its stated goal. The program was built to establish a common structural floor across a health IT market that, before certification requirements existed, had no such floor at all, and that floor is a genuine prerequisite for the layered testing described here: none of it is possible without a baseline of structurally conformant, certified systems to test against in the first place. The argument is narrower and, for an engineering or QA leader making a resourcing decision, more actionable: conformance testing answers a real and necessary question, it does not answer the question "will this integration work," and budgeting zero engineering time for the four layers above because certification already happened is a decision, not a default.
Where Leaders Should Draw the Line
The organizations that get this right tend to share one trait: they treat "certified" as a procurement and legal prerequisite, and treat "tested against this specific partner, across a full exchange lifecycle, against more than one real implementation" as the actual engineering bar for production readiness. The gap between those two bars is not a rounding error. It is the difference between a medication list with eleven entries and one with two, discovered by a nurse rather than a test harness.
The question worth taking back to an engineering team is not whether the last integration passed certification. It almost certainly did, or it would not have shipped. The question is which of the four layers above that integration actually went through before its first real patient's data crossed the wire, and whether anyone can name, specifically, which trading partner's sandbox it was tested against.
A useful distinction to hold onto, separate from any specific testing tactic: certification answers "can this system be trusted to build a conformant resource," and interoperability testing answers "will this specific pair of systems, in this specific configuration, agree on what that resource means." An organization that has only ever answered the first question has not yet asked the second one, regardless of how many certification badges sit on its compliance page.
Where QAtronic Fits
The gap this article describes is fundamentally an integration-testing problem wearing healthcare-specific clothing: two independently built systems, each individually correct against its own specification, producing an outcome that is only visible once you test the connection itself rather than either endpoint in isolation. That is the kind of problem functional and integration testing work is built to catch, and it is where automated testing earns its keep for exactly the layered, repeatable, multi-partner test cycles described above. Running the same synthetic-patient lifecycle scenario against a second sandbox is exactly the kind of test worth automating once, not re-running by hand for every new partner. QAtronic works with engineering teams on API and integration testing programs of this shape across regulated and unregulated industries alike; for a team already carrying FHIR conformance testing as a certification requirement, the practical next step is usually a scoped bilateral and lifecycle test plan built around the trading partners that matter most, not a wholesale rebuild of an existing test suite.
Frequently Asked Questions
Does passing ONC certification mean our FHIR API is interoperable with other certified systems? No. Certification, including the Inferno-based (g)(10) Standardized API test kit, confirms that your system produces structurally conformant resources against a specification and implementation guide. It does not test your system against any specific trading partner's real implementation, and it cannot, because a general-purpose conformance tool has no way to simulate every other certified vendor's independent implementation choices.
What is the fastest way to find out if we have an undiscovered interoperability gap with an existing partner? Run a must-support population audit against a realistic sample of resources from that partner's actual production or sandbox data, not synthetic conformance-tool fixtures, and compare population rates for every field your profile marks must-support. Gaps that show up as a field being technically present but rarely populated are usually the first sign of a workflow-driven divergence rather than a specification violation.
Is FHIR conformance testing still worth doing if it doesn't catch these problems? Yes, without qualification. Conformance testing is the floor every subsequent layer of testing depends on: there is no meaningful way to run bilateral or lifecycle testing against a system that has not first established basic structural conformance. Skipping conformance testing to jump straight to bilateral testing would mean debugging specification-level structural defects and real-world implementation variance at the same time, which is harder to isolate, not easier. The point is that conformance testing is a necessary floor, not a finish line, and treating a passed test kit as the end of the testing program rather than its starting point is the specific mistake this article is arguing against.
How is this different from testing HIPAA compliance? HIPAA compliance testing addresses who is authorized to access protected health information and whether access controls, audit logging, and transmission security hold up technically and procedurally. FHIR interoperability testing addresses a completely different question: whether data that both systems are fully authorized to exchange arrives complete and gets interpreted the same way on both ends. A system can have flawless access controls and still lose a medication list to an extension neither side tested end to end.
Does participating in TEFCA reduce the need for this kind of testing? It changes where the risk sits rather than removing it. TEFCA reduces the integration burden of connecting to many partners individually, but it also means your system will increasingly exchange data with organizations you have never tested against directly through a QHIN network. That makes automated, repeatable lifecycle testing, Layer 3 in the framework above, more valuable, not less, since bilateral testing with every possible network participant is not realistic at TEFCA's intended scale.
What is the single highest-value test to add if we can only add one? For most organizations, a synthetic-patient full-lifecycle test run against a second certified sandbox, not just the one used during initial development, surfaces the largest share of real-world interoperability gaps for the smallest testing investment, because it exposes cross-implementation variance directly rather than requiring you to guess where it might occur.