What AI Changes for the Member Waiting on a Coverage Review, and the Health Plan Performing It

Any health plan running coverage reviews is working against two requirements at once.

Two rules that look like a conflict

Decide faster.

The Centers for Medicare & Medicaid Services requires prior authorization decisions within seven calendar days, or 72 hours when expedited2. Under 2025 industry commitments, participating plans pledged to answer at least 80% of electronic approvals in real time by 20273.

Don't let software decide.

California's SB 1120 bars artificial intelligence from denying, delaying, or modifying care on its own, and requires that medical necessity be determined by a competent clinician4. Texas prohibits AI as the sole basis for a denial, and thirty-seven states have enacted or introduced laws governing AI in health care5.

Together they look like a trap: move faster, without the tool that would let you. They aren't — and the way out is not working harder at either. It's putting each part of a coverage review where it belongs.

Where the work goes

A coverage review asks several unrelated questions at once, and only one of them — whether this is the right amount of care for this person — needs clinical judgment. The rest is settled in two steps — the criteria check, and the routing that follows it — and in both, work can end up with the wrong party.

The criteria check.

A nurse reviewer is checking documentation against rules, working across disparate sources by hand. That is rule-checking work done by a clinician, and it costs twice: it caps throughput, and it produces inconsistent answers. A sample of Medicare Advantage denials found 13% met Medicare's own coverage rules — meaning they likely should have been approved1. One cause recurred: the plan denied for insufficient information that was already in the file. Each of those is weeks of waiting at best — an appeal, a resubmission — and at worst, care the member was entitled to and didn't get.

The routing step.

Plans do separate these cases somewhere: reason codes distinguish them, and pend workflows ask for what's absent. But a distinction recorded as an outcome isn't the same as one used to route. Where it isn't, a failed check sends the request to a physician under a single label — criteria not met — that records nothing about why. An absent document, a document that doesn't say quite the right thing, and a genuinely ambiguous case all arrive on the same desk looking identical. Two of those three need a phone call or a careful read, not a physician. The member's wait is set by the slowest available path rather than by what the request actually needed.

Figure 1: a request through coverage review, with the two failure modes marked. Request submitted goes to the criteria check (failure mode one: throughput ceiling). Requests that meet criteria get a determination; requests that fail go to a routing step (failure mode two: cause discarded), where a document absent, a document insufficient and a genuine clinical question all travel as 'criteria not met' to physician review, arriving looking identical.
Figure 1. A request through coverage review, with the two failure modes marked. Blue: the process as designed. Orange: the two steps where the member's wait is actually set.

Four changes that put the work in the right place

1

Compile the criteria into rules

Some criteria already are executable — purchased sets ship with software. A plan's own medical policy and benefit grids often live as documents, and the grids are what settle whether a service is covered, whether authorization is required, and which limits and exclusions apply. Where those pieces aren't joined, the reviewer becomes the integration layer, which is how a clinician ends up doing clerical work.

Compiling the rest into logic is a problem AI can solve. Extracting the rules means reading disparate sources the way a reviewer would, then rendering what's found as explicit logic a clinician can inspect, correct and approve. Not someone else's criteria replacing the plan's — the plan's own policy, made executable alongside them.

Note where the AI sits: it builds the rules, and the rules make the decisions. Once approved, the logic returns the same approve, pend or deny for the same inputs every time, with the governing clause cited alongside the result. Reviewers stop synthesizing and start confirming — a different and much faster task — and each decision carries its own defense, which matters when the plan justifies it on appeal or publishes its overturn rate.

Approving rules one at a time isn't the same as trusting the set — run the compiled logic against the plan's own historical determinations before go-live, and keep those scenarios as a regression suite.

And the result belongs to the plan. The rules, the citations, and the test scenarios are a versioned asset the plan's own staff maintain — not a dependency on whoever built it.

2

Govern the AI — and use automation to watch it

Start with where the model is allowed to act. The language model reads, classifies, extracts and drafts — it does not decide medical necessity. In a compiled rule set the model isn't running when a determination is served — it works during extraction and review, not at the moment of decision. That separation is what distinguishes a compiled rule set from the black-box denial tools now in litigation, and what makes the result explainable in procurement, regulator review, and appeal.

Federal policy draws the same line. In an AI-assisted authorization model now running for traditional Medicare, software may approve a request; anything not approved goes to a human clinician with relevant expertise6. California and Texas set the same limit in statute.

Software may approve. A clinician must deny.

Then the part that only works because the process is automated. Human review can be audited by sample; nobody reads every decision a reviewer made last quarter. A model can be watched in full — every request, continuously, with anomalies flagged before they spread. That matters because a reviewer's error affects one member, while a model with a drifted threshold affects every request it touches until somebody notices. The same automation that creates the exposure is what makes complete oversight possible.

One decision to make explicitly: what rate of incorrect approvals is acceptable, and who signed off on it.

3

Use the interoperability requirements to deliver the evidence

Plans have to stand up standard software connections by January 20272. Three published specifications sit under the prior authorization requirement: one lets a physician's ordering system ask whether authorization is needed and what documentation is required, one pulls the relevant records from the chart and attaches them, and one submits the request and returns the decision.

These can be built to satisfy the interoperability requirement and nothing more. But automatically pulling records from the chart removes the most common reason a request stalls: the evidence exists and nobody can reach it. It also feeds the compiled rules better inputs than a reviewer can assemble by hand.

4

Let AI read what the rules can't

Some requests will still fail the check, even with good rules and clean inputs, and nothing in the process says why. Separating an absent document from an insufficient one from a real clinical question means reading the file — the progress note, the narrative, the attachment. A rules engine can check whether a field is populated; it cannot judge whether a note establishes a failed trial. Language models can, which makes three things possible.

  • Retrieve. Find evidence already in the chart and attach it, rather than pending the request and waiting.
  • Classify. Label the failure by cause, and route each type to the fastest path that can resolve it.
  • Check sufficiency. Name what's missing, so the next submission is the complete one rather than the third.

What reaches a physician is then labeled for what it is: the evidence is present, and someone needs to judge whether it's enough.

Where to start

In the order above. One and two wait on nothing.

This is a scoped program, not a transformation. It moves in phases — discovery and governance framework, extraction and review, then deployment and handoff — each producing something specific. The first is a working rules set for one service line, which is also the proof that the rest is worth doing.

First, together 12

Compiling works with the inputs a plan already has, in whatever form they exist, and fixes consistency and throughput without waiting for cleaner data. Governance belongs in at the same time rather than after — retrofitting an audit trail onto a system already issuing decisions is harder than building one that emits it.

Clinical judgment is what makes the rule set defensible. Every rule carries a clinician's approval before it goes live, and that judgment is exercised once per rule rather than once per request — it then governs every request that rule touches. Criteria change, so the review continues, and approval capacity becomes the pacing item. That is a better constraint to have than reviewer throughput, and worth planning for rather than discovering.

In parallel 3

Three runs in parallel. A plan doesn't have to finish the connections to improve the decision: the January 2027 program continues on its own schedule, and coverage review gets faster before it lands.

Last 4

Four comes easier for having waited. Once the rules are explicit and evidence is arriving in structured form, a model classifying why a request failed has something definite to classify against, and a governance framework already built to hold it. Attempted first, it's a harder problem with less to show.

The opportunity sizes itself from numbers a plan already has: annual authorization volume, times the fully loaded cost of a review, times the share that stalled on a document rather than a clinical question. If that third term is already measured, the case makes itself. If it isn't, it's a short piece of work — and the one that decides whether the rest is worth doing.

The medicine belongs to clinicians — criteria content, judgment calls, what counts as adequate evidence that a treatment failed. Compiling that policy into logic, delivering the chart, classifying what fails, and governing the model is engineering work. Put each where it belongs and the review gets faster and more accurate at once, which is why the member's wait and the plan's cost move together rather than against each other.

We've taken this from concept through detailed solution design with a health plan administrator — the rules architecture, the governance model, the interoperability endpoints, and the delivery plan. If you're weighing where to start, we can show you what that looks like.

Request a walkthrough

Sources

  1. Improper denials and information present in the case file — HHS Office of Inspector General, “Some Medicare Advantage Organization Denials of Prior Authorization Requests Raise Concerns About Beneficiary Access to Medically Necessary Care,” OEI-09-18-00260, April 2022. Stratified random sample of 250 prior authorization denials from June 2019 across 15 large Medicare Advantage organizations. oig.hhs.gov
  2. Decision timeframes; interoperability and prior authorization requirements — CMS Interoperability and Prior Authorization Final Rule (CMS-0057-F). cms.gov
  3. 2025 industry commitments — AHIP, “Health Plans Take Action to Simplify Prior Authorization,” June 2025. ahip.org
  4. Clinician-review requirement in state law — California SB 1120 (2024). leginfo.legislature.ca.gov
  5. State AI legislation survey — KFF, “Regulation of AI in Prior Authorization and Claims Review,” May 2026. kff.org
  6. Approval permitted, denial reserved to clinicians — CMS, WISeR Model Request for Applications and Model Fact Sheet. cms.gov
Share

About the Author

Digineer

Related Posts


Format: Article ·

Change is your Friend

25 years ago, Digineer was founded in the midst of change, Y2K, the dot.com boom. In those 25 years, we have seen a little company called Apple go...

Read More
Capability: Organizational Effectiveness ·

Engaging your IT culture

In case you missed it, here is our most recent webinar with Shannon Rose Farrell-Jackson, a leader in organizational change management and culture...

Read More