Random sampling can pass a claims population that still contains a meaningful, concentrated error. This is not necessarily a sign of poor execution.
A properly randomized sample of a few hundred claims has a limited statistical chance of capturing an error that affects only a narrow slice of total volume, such as a specific coder, provider group, procedure code, or other defined subgroup, even when the sampling process runs exactly as designed.
The audit result and the underlying claims population are not the same, and the gap between them is the focus of this piece. Errors that fall outside a random sample are not necessarily evidence that the sampling process failed. They can be a predictable consequence of what a sample is designed to detect.
Distinguishing the Concepts Involved
Sample size, representativeness, and detection rate describe different aspects of an audit's reliability, and improving one does not automatically improve the others.
Sample size is simply how many claims a review examines. Representativeness is whether the sample reflects the actual composition of the broader population, including any meaningful subgroups within it. Detection rate is the share of actual errors the review identifies.
A larger sample size can improve the chance of detecting evenly distributed errors while still missing a clustered error entirely if the sampling process does not structure the sample to represent the specific subgroup where that error is concentrated.
Stratification, dividing a population into meaningful subgroups before sampling from each, addresses representativeness directly and can improve the ability to detect errors concentrated within those subgroups when the sampling design and sample sizes are appropriate.
Why This May Reflect a Structural Limitation, Not an Execution Gap
A published academic study compared random-sample auditing against a full-claims-review methodology, using actual paid claims data from two Fortune 100 companies. The research found that random sampling caught fewer than 10 percent of the errors the full-review method identified, missing between $200,000 and $750,000 in errors across the datasets examined.
The study does not establish that random sampling generally catches only 10 percent of errors across all healthcare claims environments. It demonstrates that, under the specific conditions modeled, correctly executed random sampling can miss the substantial majority of errors that a full-claims-review approach identifies.
The study's authors also had a commercial interest in the 100-percent auditing approach, which is relevant context when interpreting the findings. The paper was nevertheless published in a peer-reviewed academic journal.
The study does highlight an important statistical premise: correctly executed random sampling can miss a substantial number of errors when those errors are concentrated in parts of the population that are not adequately represented in the sample. That is a limitation of what the sample can reveal, not necessarily a flaw in the sampling process itself.
Organizations should treat this as a documented risk to be evaluated against their own claims population, not as a universal detection rate. Audit sampling limitations of this kind are easier to plan around once an organization understands which risks its current sampling design can detect and which may require additional review methods.
What Can Cause Random Sampling to Under- Detect Errors
Rare or Clustered Errors Have Limited Odds of Selection
An error affecting a small percentage of total claims has a correspondingly small statistical chance of appearing frequently enough within a modest random sample to register as a pattern rather than an isolated instance.
When an error concentrates within a narrow subgroup, a specific coder, provider, or procedure code, the odds of an unstratified sample surfacing it as a pattern decline further, since the sample would need to draw disproportionately from that subgroup to reveal what is happening within it.
Heterogeneous Populations Can Distort Unstratified Sampling
A claims population often includes meaningfully different subgroups by location, service line, provider, or billing period, each with distinct documentation practices and risk profiles.
A sample drawn from a varied population without first dividing it into subgroups may provide limited information about a narrow subgroup, since a sample sized appropriately for the full population may be too small to yield statistically reliable estimates for a narrow slice within it.
Partial and Line-Item Errors Complicate Standard Error Assumptions
Many sampling approaches assume an error affects an entire claim or none of it, an assumption that holds reasonably well for some audit objectives and poorly for others.
A claim with multiple line-item charges may contain errors in only a portion of those charges, and a sampling method built around whole-claim assumptions can understate the actual financial exposure a partially incorrect claim represents.
An Executive Diagnostic for Assessing Current Exposure
Applying three questions to an existing sampling process can help indicate the organization's likely exposure to this specific risk.
First, identify whether the sampling methodology divides the claims population into meaningful subgroups before drawing a sample, or draws uniformly across the entire population.
Second, for any known concentrated risk areas, such as a specific provider group, claim type, or coding pattern, estimate whether the current sample size and structure would realistically capture enough instances of an error in that subgroup to identify a meaningful a pattern.
Third, review whether any mechanism exists for flagging a concentrated pattern outside the sampled claims, or whether the sampled result is being treated as representative of the entire population without additional risk-based review.
An organization that can answer all three with reasonable confidence has a clearer understanding of the exposure described in this piece. One unable to answer them has specific, addressable areas to examine rather than open-ended concern.
Why Expanding Coverage Does Not Resolve the Underlying Issue on Its Own
Expanding claims review beyond a random sample can improve detection, but it introduces operational requirements that often go unaddressed.
A team that moves from a small routine sample to reviewing a substantially larger share of total volume will likely surface more findings, and a review, rebuttal, and resolution process not scaled to match that increase can produce a growing backlog rather than more errors being resolved. Expanding coverage without a corresponding increase in root-cause analysis capacity can leave a team identifying considerably more problems than it has the capacity to explain or prevent from recurring.
Expanding coverage unevenly can create another blind spot. If more scrutiny is applied to straightforward claim types while complex claim types receive less attention, the review may still underrepresent the areas where errors are concentrated.
None of this suggests expanding coverage is unwise. It suggests the operational capacity to act on additional findings deserves the same planning attention as the statistical decision to look for them.
A Framework for Evaluating a Sampling Approach
Three criteria help indicate whether an existing sampling process is likely to catch the errors that matter the most, financially:
Stratification: Does the sampling process draw from defined subgroups within the population, or treat the population as uniform?
Sample size relative to concentration risk: Does the sample size account for how concentrated an error would need to be for the process to realistically detect it, or does it only calibrate for evenly distributed problems?
Visibility outside the sample: Does any mechanism exist for surfacing a concentrated pattern in claims the sample did not include, or does the sampled result stand in for the full population by default?
A program that can answer all three affirmatively has likely reduced much of the exposure described in this piece. A program that cannot have a defined starting point, rather than a general sense that its audit process could improve.
Where Claims Quality Management Practices Appear to Be Shifting
Expanded and continuous claims review has become more technically and financially feasible than it was when random sampling became the standard approach, largely because the cost of reviewing a larger share of claims has declined, not because the underlying statistical case for sampling has changed.
Some organizations are accordingly evaluating QA sampling healthcare claims processes against that changing cost equation, treating the choice between sampling and broader coverage as an operational decision rather than the only realistic option it once was.
This shift is not yet universal, and organizations considering it should evaluate their own claims population and error patterns directly, rather than assume a general industry trend applies uniformly to their specific circumstances.
What a Passed Random Audit Does, and Does Not, Establish
A passed random audit does not establish that a claims population is free of significant errors. It establishes that, if any errors exist, they did not fall within a specific, randomly drawn slice of that population.
Healthcare quality assurance audits designed around that distinction, treating a clean sample result as a limited signal rather than a conclusive finding, are asking a more precise question than sampling alone can answer: not whether the sample looked acceptable, but whether the broader population it came from actually is.
Frequently Asked Questions
Why might random sampling miss errors even when the audit process runs correctly?
Random sampling gives every claim an equal chance of selection, which works reasonably well for errors distributed evenly across a population but performs poorly against errors that cluster within a narrow claim type, provider, or coding pattern. A modest sample size has limited statistical power to capture enough instances of a concentrated error to reveal it as a pattern rather than an isolated case.
What does the available research actually indicate about random sampling's error detection rate?
One peer-reviewed study comparing random-sample auditing to full-claims-review auditing found random sampling caught fewer than 10 percent of the errors the full-review method identified, using real paid claims data from two large companies. An audit firm with a commercial interest in the outcome funded the study, and it represents a single dataset under specific conditions rather than a universally applicable detection rate.
How does population heterogeneity affect sampling accuracy specifically?
When a claims population includes meaningfully different subgroups, by provider, service line, or billing period, a sample drawn without first stratifying by those subgroups can misrepresent any one of them, since a sample sized for the overall population is often too small to support reliable conclusions about a narrow slice within it.
Does moving to a broader claim review entirely resolve the limitations of sampling?
Not entirely on its own. Broader coverage tends to surface more findings, and without a rebuttal, resolution, and root-cause process capable of handling that increase, the practical result can be a larger backlog rather than proportionally more errors resolved. The operational side of the transition warrants the same attention as the statistical rationale.
What should a QA program examine first regarding its current sampling approach?
Whether the process stratifies the claims population into meaningful subgroups before sampling is often the most informative starting point, since an unstratified sample applied to a genuinely varied population is one of the more common ways a technically correct sampling process still under detects concentrated, high-impact errors.
Katrina Huynh is a healthcare strategy and operations leader with more than 15 years of experience, including over a decade working within Blue plans. Her experience spans health plan operations, enterprise strategy, client relationships, transformation, and strategic partnerships. As Director, Strategic Partnerships and Growth at MDI NetworX, she drives strategic partnerships, identifies growth opportunities, and translates organizational priorities into execution. With firsthand health plan experience, Katrina brings a broad perspective on how people, technology, data, and operations come together to create meaningful results. Known for bringing clarity to complexity and connecting people and capabilities, she is focused on advancing practical, high-impact solutions across healthcare operations.