You Don't Have a Retrieval Problem. You Have an Abstraction Problem.
InsightsHEDIS & Stars

You Don't Have a Retrieval Problem. You Have an Abstraction Problem.

Why the last mile of data collection is where Stars performance actually gets made or lost.

Peter Saah, DBA, MBA, CPHQCEO & Co-Founder, Podero HealthMay 22, 20268 min read

Most Stars misses are not missing care.

The member got the care. The physician documented it. The chart was retrieved. The evidence existed in the record.

And the measure still stayed open anyway.

When I ask health plan operators why a gap remained open after a successful chart retrieval, the default answer is almost always the same: the evidence was insufficient. So the member gets added to next year's outreach list, another retrieval gets ordered, another provider gets contacted.

But in a significant percentage of those cases, the evidence was not insufficient. It was misinterpreted. The care happened. The documentation was there. Abstraction got it wrong.

That distinction matters more than most plans have recognized — because the intervention for a retrieval gap and the intervention for an abstraction failure are completely different. And spending money on retrieval to solve an abstraction problem compounds year after year.

The Measurement Most Plans Do Not Have

Before getting into the specific failure modes, consider this question.

What is your abstraction accuracy rate — by coder, by measure, by data source, and by season?

Not your retrieval completion rate. Not your overall HEDIS score. Not your submission rate. The percentage of abstraction decisions that were correct against the current year's specification on the first pass.

Most health plan quality programs cannot answer that question. Not because the question is unreasonable — it is the most reasonable quality control question in the entire pipeline — but because abstraction accuracy is almost never tracked as a standalone metric. It gets absorbed into the final submission output. If the measure closes, abstraction is assumed to have worked correctly.

That assumption is where Stars performance leaks.

What About Vendor QA and AI-Assisted Abstraction?

Two objections come up immediately when abstraction accuracy is raised as an unmanaged problem, and both deserve a direct answer.

The first is vendor quality assurance. Many abstraction vendors report sub-one-percent internal error rates and conduct one-hundred-percent over-read processes. Those are real and meaningful quality controls. What they measure is consistency — whether the primary coder and the reviewing coder reached the same conclusion. What they do not measure is whether the shared interpretation logic applied by both the coder and the reviewer is correct against the measure specification. Two reviewers can agree on the wrong interpretation. When an entire abstraction team systematically misapplies a lookback window, misreads qualifying language, or applies prior-year specifications to a current-year record, that error passes internal QA checks consistently. The over-read validates the error. And it still reaches submission.

The second objection is AI-assisted abstraction. Modern clinical NLP is genuinely capable. Well-trained HEDIS abstraction systems handle negation, chronology, and clinical nuance. They surface compliant evidence from unstructured text at a scale no manual team can match at volume. Vendors including Reveleer report accuracy rates above ninety-five percent on AI-assisted evidence identification, and these are real gains.

But the standard for AI-assisted HEDIS abstraction is hybrid: the AI identifies and surfaces relevant evidence, and the human abstractor makes the final adjudication. That human decision — whether the AI-surfaced evidence actually satisfies the measure, in the full context of the clinical documentation, against this year's specification — is still a human call. And that call is almost never measured separately from the AI's evidence identification accuracy. Those are two different metrics. Most plans track one of them.

The failure modes that follow all live in the judgment layer that sits between what the documentation says and what gets entered into the abstraction system — whether that judgment is preceded by manual chart review or AI-assisted evidence surfacing.

The Seven Ways Abstraction Fails

Abstraction failures in HEDIS production environments fall into seven categories.

One: Clinical Interpretation Errors

This is the most common and most consequential failure mode. The chart contains valid evidence. The abstractor draws the wrong conclusion.

The most frequent variants: misidentifying procedure type — a diagnostic colonoscopy counted when the specification requires a screening colonoscopy; misreading referral language as completed care; missing qualifying language that changes what the documentation supports. A physician note reading "colonoscopy recommended — patient declined" is a referral. It is not evidence of completed screening. Under production pressure, with high chart volume and time constraints, interpretation variance of this kind happens regularly.

Clinicians document for clinical care. Abstractors interpret for compliance logic. Those are not the same activity, and the documentation was not written with HEDIS abstraction in mind.

Two: Measure Specification Misapplication

The abstractor has the data. They apply the wrong rule.

This failure mode is particularly prevalent with lookback windows, which vary by procedure type within a single measure and change annually. Under current COL-E specifications, the lookback windows are: nine years for colonoscopy, four years for CT colonography, four years for flexible sigmoidoscopy, two years for stool DNA with FIT, and the current measurement year only for standard FOBT.

Misclassifying CT colonography as colonoscopy and applying a nine-year lookback to a procedure with a four-year window is not a cosmetic coding difference. Depending on when the procedure occurred, it produces either a false closure or a false miss. Either outcome is a material abstraction error that passes every internal QA check if both the coder and the reviewer share the same misunderstanding.

NCQA updates its technical specifications annually. An abstractor running on prior-year training is a systematic specification error risk for the entire measurement season.

Three: Data Element Mapping Errors

The clinical fact is correct. The structured output is wrong. A screening colonoscopy gets logged as a diagnostic procedure. A lab result gets mis-tagged as imaging. The date of service gets confused with the result date. A CPT or HCPCS code gets mapped incorrectly relative to the current value set.

The outcome: evidence exists in the record. It is structurally unusable for submission. It passed every upstream check — retrieval, feed ingest, reconciliation — and failed at the moment of translation from clinical documentation to abstraction record.

Four: Source Prioritization Errors

Multiple sources exist for the same member. The abstractor selects the wrong one. An outdated problem list overrides a recent specialist note. Claims data overrides direct encounter documentation. Duplicate records are not reconciled and the abstractor works from the incorrect version.

Without a consistently enforced source hierarchy — one that is specific enough to produce the same decision regardless of which coder handles the record — every abstractor makes individual judgment calls on source priority. Across a production team and vendor partners, the result is not one standardized process. It is multiple individual processes that appear standardized from the outside.

Five: Timing and Measurement Window Errors

Correct event. Wrong time context. A procedure falls outside the applicable lookback period. A result date is confused with a service date. A preventive service date is read as a diagnostic follow-up date. Measurement year boundaries are misapplied to a specific procedure type.

These are temporal errors, and they produce either false closures — a measure marked as satisfied when it legally is not — or false exclusions — valid care disqualified by a date misread.

Six: Cognitive Load Omission

Where most programs break down

This is where most quality programs break down

Execution across retrieval, abstraction, and outreach is rarely connected. That's where gaps stall.

See how Podero connects the workflow

The evidence is in the chart. The abstractor does not see it. Not because they made a wrong decision — because they did not encounter it under production conditions.

High chart volume, workflow fatigue, ambiguous documentation that requires deliberate interpretation, year-end pressure that compresses review time — these are the conditions under which cognitive load omission occurs. This is the silent failure. The chart shows as reviewed in the system. The evidence simply does not get captured. And the member returns to the gap list the following year for reasons no system report will accurately describe.

Seven: Incentive-Driven Abstraction Drift

Abstractors working in production environments optimize, consciously or not, for what gets measured and rewarded in that environment. When throughput is the primary metric, speed wins over accuracy. When closure rate targets are emphasized, marginal evidence becomes sufficient that a more rigorous reviewer would not accept. When audit avoidance is the cultural priority, abstractors code conservatively and miss legitimate closures to avoid challenge exposure.

Unlike random error — which averages out across a large member population — systematic directional bias produced by incentive misalignment does not average out. It compounds across every measure it touches, in the same direction, across every production season.

Three Root Causes

Seven failure modes share three underlying causes. The root cause determines the fix, so naming them precisely matters.

Root Cause One: Interpretation Variance

Human judgment is inconsistent. Different abstractors reviewing the same chart documentation do not always reach the same conclusion — not because one is wrong and one is right, but because clinical documentation was written for clinical purposes, and the translation into a HEDIS compliance determination requires judgment that varies by training, experience, and the cognitive conditions at the moment of review.

Interpretation variance is measurable. Most plans do not measure it. Inter-rater reliability audits measure whether coders agree with each other — not whether their shared interpretation is correct against the specification.

Root Cause Two: Specification Complexity

NCQA's measure specifications are technically demanding. Annually versioned. Measure-specific. With procedure-specific lookback windows, age stratifications, continuous enrollment requirements, exclusion logic, and value set dependencies that change from year to year.

The complexity is unavoidable — the measures are complex because the clinical realities they capture are complex. The problem is that most organizations train specifications at season kickoff and do not enforce them systematically at the point of abstraction decision. The gap between seasonal training and point-of-decision enforcement is where specification drift accumulates invisibly.

Root Cause Three: Human Adjudication Accuracy Is Unmeasured

This is the most operationally significant root cause. AI-assisted abstraction has improved the evidence identification step. It has not eliminated the human adjudication step. The final decision — whether AI-surfaced evidence, in the full context of the documentation, against this year's specification, actually satisfies the numerator — remains a human call.

What almost no quality program tracks is the outcome accuracy of that human call separately from the AI's evidence identification accuracy. Vendors report extraction accuracy. Plans report submission rates. Neither metric tells you whether the human abstractor's final determination was correct. That is the measurement gap where Stars performance is lost in a way that no downstream report will accurately attribute.

The Financial Calculation Most Plans Have Never Run

Abstraction error rates in real production HEDIS environments — accounting for the full range of failure modes, not just internal QA disagreements — vary based on measure complexity, chart volume, and seasonal pressure. Independent research on clinical NLP in healthcare contexts suggests meaningful residual error rates even in AI-assisted environments, particularly for complex multi-source clinical reasoning scenarios.

For illustration, consider a plan with ten thousand members on a hybrid measure and a five-percent abstraction error rate. Five percent is five hundred members where care occurred, evidence existed, the chart was retrieved — and the measure still stayed open because of an abstraction error.

The financial significance depends on measure weighting. A five-percent abstraction error rate on a heavily weighted measure — Diabetes Care, Breast Cancer Screening, Controlling Blood Pressure — has materially different Stars and bonus threshold implications than the same rate on a lower-weighted measure. For plans where two or three measures are doing the heavy lifting on the Part C summary rating, abstraction accuracy on those specific measures is not a quality operations metric. It is a financial variable.

Plans that have not run this calculation have not connected abstraction accuracy to Stars performance at the measure level. They are treating abstraction as a process function and making resource allocation decisions accordingly. The result: continued investment in retrieval and outreach to solve a problem that originates in the abstraction layer.

What Good Abstraction Operations Actually Look Like

Four operational practices separate the plans managing abstraction accuracy from the plans assuming it works.

First: They Track Accuracy, Not Just Completion

A coder-level accuracy metric exists — not throughput, not chart completion count. The specific percentage of decisions that were correct against the current year's specification on secondary review, tracked by coder, by measure, by source, and over time. That metric drives training decisions between production seasons, not just at season kickoff.

Second: They Embed Enforcement Into the Workflow

Abstraction decisions are validated at the point of entry, not discovered in post-production audit. If an abstractor enters a colonoscopy date outside the nine-year lookback window, the system flags it before the record closes. If a CT colonography is coded using a nine-year window instead of four, it gets caught at entry. Enforcement built into the workflow prevents errors from propagating to submission. Post-production audit catches some of them afterward. The operational and financial difference between the two is significant.

Third: They Operationalize Specification Updates Before Production

When NCQA publishes annual technical specification updates, these organizations run historical charts through the updated logic before season kickoff. They identify precisely where updated rules produce different results than prior-year rules and retrain on those specific decision points before the first production chart is reviewed. Not a general communication about this year's changes — targeted retraining on exactly where prior-year logic is now incorrect.

Fourth: They Treat Q4 Accuracy Degradation as Predictable Risk

Q4 abstraction accuracy is systematically lower than Q2 accuracy. Higher chart volume, more complex charts, compressed review cycles, year-end production pressure. This is not surprising — it is predictable. Organizations that manage abstraction well staff and build QA capacity around Q4 accordingly. The downstream cost of Q4 abstraction errors is identical to the downstream cost of late-year retrieval failure: the measure does not close, the gap stays open, and the member returns to the gap list the following year.

The Question to Take Back

You know your retrieval rate. You know your ingest success rate.

Do you know your abstraction accuracy rate — by coder, by measure, by source, and by season?

If the answer is no — or if the answer is "we conduct QA reviews but I could not give you a specific number" — the biggest data quality problem in the pipeline may not be retrieval. It may be the unmeasured interpretation layer at the very end of it.

The most expensive chart in healthcare is the chart you already retrieved correctly — and still failed to close.

Go Deeper — This article is based on Episode 6 of The Execution Gap Podcast: "You Don't Have a Retrieval Problem. You Have an Abstraction Problem." 16 minutes on the seven ways abstraction fails in real HEDIS production environments, why vendor sub-1% error rates measure consistency and not correctness, why AI-assisted abstraction shifts the judgment layer without eliminating it, and the four operational practices that separate plans managing abstraction from plans assuming it works. Listen to Episode 6 →

How Podero Health Approaches This

Podero Health is built to give health plans visibility and control at the data layer — from retrieval and digital feeds through reconciliation and abstraction — on a single execution pipeline.

On the abstraction accuracy problem specifically, Podero pilots establish four things: where interpretation variance exists across the abstraction workflow, where specification drift is occurring between training and point-of-decision enforcement, where source hierarchy is inconsistently applied, and where Q4 volume pressure is producing measurable accuracy degradation.

That baseline — abstraction accuracy by coder, by measure, by source, by season — is the number most plans have never had. It is also the number that tells you where the real performance opportunity is, before more resources get allocated to retrieval or outreach targeting the wrong layer of the problem.

If you want to see what this looks like against your actual data, that is what a pilot conversation is designed to establish.

P

Peter Saah, DBA, MBA, CPHQ

CEO & Co-Founder, Podero Health

See how quality execution actually runs end-to-end

From chart retrieval to abstraction, outreach, and evidence submission — without fragmented workflows.