Clinical Documentation Quality and Revenue Leakage in Hospital Billing

Clinical Documentation Quality and Revenue Leakage in Hospital Billing

HEALTH FINANCING AND ACCOUNTING

A POSTGRADUATE DIPLOMA PUBLICATION

A Comparative Quantitative Analysis of Denial and Coding Loss in the United States, the United Kingdom, and Casemix Systems


By Dominic Okoro

New York Center for Advanced Research (NYCAR)

Research Division — Health Financing and Public Financial Management

Institutional Review · August 2026

Publication No.: NYCAR-TTR-2026-RP073

DOI: http://zenodo.org/records/22028615


Peer Review Status


Peer Review: Independent Review

This postgraduate publication has undergone independent peer review conducted under the joint editorial framework of the New York Center for Advanced Research (NYCAR) and The Thinkers’ Review. Independent reviewers assessed the research for academic coherence, source integrity, financial and methodological rigor, scientific voice, and APA 7th edition alignment.

Each quantitative model was independently re-derived, every cited source independently verified, and the work cleared for release only on the basis of that independent assessment.


The cover carries independent peer review because the research corrects a denominator error in a published source and re-derives the affected ratios.

Abstract

Clinical Documentation Quality and Revenue Leakage in Hospital Billing examines the distance between the care a hospital delivers and the money it is ultimately paid for delivering it. That distance has two separate causes which hospital finance functions routinely treat as one. A payer may refuse a claim that was correctly documented and correctly coded. Or a hospital may document and code its own work badly enough that it bills for less than it did. The first is adversarial and one-directional; the second is internal and, on the evidence assembled here, systematically biased downward.

Four published evidence bases are read: revenue cycle benchmarking covering 2,300 United States hospitals and 350,000 physicians, the national coding audit framework underpinning English activity-based payment, a casemix coding audit conducted during a diagnosis-related group implementation, and a hospital-level coding accuracy study. Every ratio reported is recomputed from the source numerators and denominators rather than accepted as published, and one published proportion is found not to reproduce from its own stated counts.

Net revenue leakage among the United States hospitals rose from 38.6 billion dollars to 48.4 billion, an increase of 25.4 percent, or 4.26 million dollars per hospital in a single year. Every tracked denial metric rose by precisely 0.2 percentage points. The denial funnel runs from an 11.6 percent initial denial rate to a 2.7 percent final rate, so 76.7 percent of initial denials are eventually resolved, but resolution efficiency fell from 78.1 percent and unrecovered clinical denials now account for 55.8 percent of all final denials.

Coding loss attenuates sharply along its own cascade. In the casemix audit 89.4 percent of records contained a coding error, 74.0 percent of those errors changed the assigned diagnosis-related group, and 52.1 percent of those reclassifications lowered the tariff, so 38.6 percent of coding errors reached the revenue statement as a loss. The English framework attenuates comparably, with a 15.1 percent primary diagnosis error rate producing a 9.4 percent grouper error rate and a financial impact of between 5 and 14 percent of payments.

The finding with the sharpest operational consequence concerns what revenue cycle performance measurement rewards. Across the same period in which leakage rose 25.4 percent, days to insurance payment improved from 57.4 to 55.2 and median accounts receivable days improved by 2.3. Cash velocity and revenue integrity moved in opposite directions. A revenue cycle function optimizing the metrics it is conventionally judged on can be getting measurably worse at collecting what it earned, and the research finds that documentation quality, not collection speed, is the variable that separates the two.


Keywords: clinical documentation improvement; revenue leakage; claim denial; diagnosis-related group; clinical coding audit; casemix; revenue cycle management; health financing; hospital billing; coding accuracy.

Table of Contents

Table of Contents

List of Tables

Table 1: Evidence inventory — populations, mechanisms, and reporting basis

Table 2: Leakage audit — movement by component, with computation

List of Figures

Figure 1: Movement in United States denial metrics, 2024 to 2025

Figure 2: Attenuation of coding error along the tariff cascade

Chapter 1: Context, Research Problem, and Professional Significance

The management problem


The analysis places billed revenue beside earned revenue, and asks which of the two mechanisms separating them a hospital can actually act upon.

A hospital earns revenue at the bedside and collects it in an office. Between the two sits a documentation and coding process that translates clinical work into a billable classification, and a payer adjudication process that decides whether to honor the resulting claim.

Revenue leakage is the money that falls out between the ward and the bank, and it falls out in two quite different ways.

A claim can be correct and refused. A payer disputes medical necessity, or a prior authorization was never obtained, or the documentation does not satisfy the payer’s clinical validation criteria even though the coding follows the classification rules exactly. The hospital did the work, recorded it properly, and is not paid for it.

This is adversarial leakage, and its defining property is that it moves in one direction only.

No payer has ever spontaneously paid a hospital more than it billed.

A claim can also be wrong and paid. The documentation was thin, the coder assigned a classification the record did not fully support, and the resulting group carried a lower weight than the care warranted. The hospital did the work, recorded it badly, and billed itself short. This is internal leakage, and its defining property is that it is invisible.

A denied claim generates a remittance advice, an appeal file and a report line.

An under-coded claim generates a payment, and the hospital records a success.

The distinction matters because it determines where remediation should be aimed and because the two are measured with wildly unequal diligence. Denial rates are tracked monthly, benchmarked nationally and reported to boards.

Coding accuracy is established by periodic audit, if at all, and the resulting error rate is treated as a compliance matter rather than a revenue one.

A hospital can therefore know its denial rate to one decimal place while having no current estimate of how much it under-billed last quarter.

The figures anchoring this research set the scale of both mechanisms. Net revenue leakage across 2,300 United States hospitals rose from 38.6 billion dollars to 48.4 billion in a single year, an increase of 25.4 percent and of 4.26 million dollars per hospital (Kodiak Solutions, 2026). Against that, national coding audit in England has recorded incorrect primary diagnosis codes in 15.1 percent of audited episodes and an estimated financial impact of between 5 and 14 percent of payments (Audit Commission, 2008). The second figure is the larger of the two as a proportion of revenue, and it is the one almost never presented to a board.

Published evidence and institutional mechanics

Four evidence bases are examined because between them they cover both mechanisms and three payment architectures. United States revenue cycle benchmarking covers 2,300 hospitals and 350,000 physicians and reports the denial mechanism in detail (Kodiak Solutions, 2026). The English activity-based payment assurance framework re-abstracts coding from clinical records and reports the coding mechanism with its payment impact (Audit Commission, 2008). A casemix implementation audit traces coding error through group reassignment to tariff effect, which is the full internal cascade in one study (Zafirah et al., 2018). A hospital coding accuracy study supplies a second point of comparison on diagnosis-level error (Alharbi, 2024).

The payment architectures differ in ways that matter for how leakage arises. The United States operates multiple competing payers adjudicating claims individually, which creates the adversarial mechanism in its strongest form and generates the prior authorization and medical necessity disputes that dominate its denial statistics. England operates a single commissioner paying against a national tariff derived from coded activity, which largely removes adjudication disputes and concentrates leakage in the coding translation itself. Casemix systems in implementation carry both problems at once, since the classification is new, the coding workforce is inexperienced, and the tariff consequences of error are immediate.

The mechanism therefore follows the architecture. Where a hospital is paid by an adversary, leakage arrives as refusal. Where a hospital is paid by a formula, leakage arrives as misclassification. Where a hospital is paid by a formula it has only just adopted, leakage arrives as both, and the resulting error rates in the casemix literature are an order of magnitude above the mature-system figures.


A denied claim announces itself. An under-coded claim is paid, and files quietly.

Aim, objectives, and research questions

The aim of this research is to quantify and compare the two mechanisms of hospital revenue leakage across payment architectures, and to establish which of them the prevailing measurement practice of revenue cycle management is equipped to detect.

1. To measure movement in the components of denial-driven leakage using benchmarking data covering a defined hospital population.

2. To model the denial funnel from initial refusal to final loss, and to compute the share of leakage that survives appeal.

3. To trace coding error through group reassignment to tariff effect, and to compute the attenuation at each stage.

4. To establish whether coding error is directionally symmetric in its revenue consequence.

5. To test whether improvement in conventional revenue cycle performance metrics accompanies improvement in revenue yield.

6. To derive documentation and assurance controls addressed to each mechanism separately.

Five questions follow: how fast is denial leakage growing and in which components; how much of it survives appeal; how much coding error reaches the revenue statement; in which direction coding error biases revenue; and whether collection speed and collection completeness move together.

Research hypotheses


H1:

Coding error is directionally symmetric, so that over-coding and under-coding offset one another in aggregate revenue effect.


H2:

Improvement in revenue cycle cash velocity is accompanied by improvement in revenue yield.


H3:

The majority of clinically denied claims are recovered on appeal.

Professional significance

For hospital finance managers the research separates a question of collection from a question of documentation, and locates most of the recoverable value in the second. A denial management team working appeals is recovering money the hospital already knows it is owed.

A documentation improvement programme is recovering money the hospital does not know it is owed, which is harder to justify at budget time and larger in effect.

For clinical staff the research reframes documentation from an administrative burden into a revenue control. The evidence indicates that the record, rather than the coder, is the binding constraint: where documentation does not connect clinical indicators to the diagnosis stated, no amount of coding skill will produce a defensible claim, and the payer’s clinical validation process is designed to find exactly that gap.

The scope is confined to acute inpatient billing under classification-based payment. Outpatient, physician professional and long-term care billing are excluded except where a source reports them alongside inpatient figures. No individual claim, patient record or hospital is examined; all evidence is aggregate and published.

Currency figures are retained as reported and no cross-currency conversion is performed at any point.


The chapter treats billed revenue as a measurement of documentation, not of care.

Chapter 2: Literature, Theory, and Evidence Base

Classification-based payment and its dependency

Prospective payment by clinical classification rests on a single dependency that its designers understood and its users routinely forget. The payment attaches to a group, the group is derived from codes, the codes are derived from the record, and the record is written by a clinician whose training, incentives and available time are directed elsewhere.

Every link in that chain is a place where revenue can be lost, and only the last two are visible to the finance function.

The English implementation makes the dependency explicit. Since 2003 a prospective casemix funding system has reimbursed a growing majority of acute inpatient activity, reaching over 90 percent by 2008, and the accuracy of the data recorded for each episode directly determines the accuracy of reimbursement between commissioner and provider (Audit Commission, 2008). Where payment is formulaic, data quality is not an information governance concern that happens to have financial consequences. It is the financial control itself.

Casemix implementations elsewhere have documented the same dependency under harsher conditions. A teaching hospital audit conducted during a national diagnosis-related group rollout re-grouped audited records through the grouper and compared the resulting tariff assignments, concluding that coding quality is a precondition of implementing casemix systems (Zafirah et al., 2018).

The framing is worth noting: coding quality is presented as a precondition of the payment system functioning at all, rather than as a margin on its performance.

Theoretical perspectives


Information asymmetry and the translation problem.

The clinician holds knowledge the coder needs and cannot independently obtain. Coding is the translation of clinical terminology as written into a statistical code using standardized classification, and the translation can only be as good as the source text (Tandem Health, 2026). Where the record states a diagnosis without recording the clinical indicators that support it, a coder acting correctly will still produce a claim a payer can defensibly refuse. The error is upstream of the coding function and is routinely attributed to it.


Asymmetric visibility of loss.

Denial produces an artifact and under-coding does not. This asymmetry structures everything about how hospitals allocate remediation effort, because management attention follows exception reports and under-coding generates none. The consequence is a systematic bias in revenue integrity investment toward the mechanism that is easier to see rather than the one that is larger.


Loss aversion and coding conservatism.

Coders and clinical documentation specialists operate under an audit regime in which over-coding carries regulatory and reputational penalty while under-coding carries none. A rational actor facing that payoff structure will resolve ambiguity downward. The prediction is that coding error will not be directionally random but will be biased toward the lower-weighted group, and the empirical test of that prediction appears in Chapter 5.


Adversarial adjudication.

Where payment is decided by a counterparty with an interest in refusal, denial rates reflect payer behavior as much as provider performance. Reported analysis attributes the recent rise in leakage to payer conduct and to a decline in the rate of overturning initial denials, with clinical denials for lack of prior authorization and medical necessity accounting for nearly all of the increase (Kodiak Solutions, 2026). A provider improvement programme cannot alter the counterparty’s posture, which bounds what documentation improvement can achieve against this mechanism.

Clinical documentation improvement as a control

Clinical documentation improvement programmes occupy the space between the clinician and the coder, reviewing records concurrently and querying clinicians where the documentation does not support the clinical picture. Survey evidence indicates how far their remit has shifted toward denial defense: among such programmes involved in denials, 87.73 percent handle clinical validation denials and 64.11 percent handle group validation denials, the latter rising from 54.66 percent in a single year (ACDIS, 2025).

The same survey identifies where the disputes concentrate. Sepsis draws scrutiny in 85 percent of programmes, respiratory failure in approximately 78 percent and encephalopathy in approximately 57 percent (ACDIS, 2025).

These are high-volume, high-weight conditions in which the difference between a defensible claim and a downgrade turns on whether the record connects clinical indicators, clinician judgment and treatment.

The list is short, which is operationally useful: a documentation programme with limited resource has a small number of conditions on which to concentrate.

The distinction between a clinical validation denial and a group validation denial is worth stating precisely because the remedies differ. A clinical validation denial argues that the criteria for a documented diagnosis were not met or not clearly supported, even where the coding was correct. A group validation denial challenges the assigned group, usually to move it to a lower-weighted one (Medovent Solutions, 2026). The first is a documentation failure. The second may be a coding disagreement or a payer tactic, and treating both as coding problems misallocates the response.

The denial environment

Denial pressure has intensified across payer categories, and the intensification is documented from several independent vantage points. Initial denial rates have risen from approximately 10.2 percent to 11.8 percent over recent years, with commercial and managed public plans contributing disproportionately (OS Healthcare, 2025). Audit activity has risen alongside refusal: analysis across 4,500 facilities recorded a 30 percent year-on-year increase in external payer audits and increases of 12 and 14 percent in the average denied inpatient and outpatient claim amount (MDaudit, 2025).

The prior authorization channel operates at a scale that is easy to underestimate. Managed public plans issued approximately 53 million prior authorization determinations in a single year at a denial rate of 7.7 percent, and 80.7 percent of denials that were appealed were overturned (Medovent Solutions, 2026). An overturn rate above four fifths indicates that most refused determinations do not survive scrutiny, and the volume indicates that most are never scrutinized, because appealing is costly and the hospital must choose which refusals to contest.

That combination, a high overturn rate on appeal and a low appeal rate in practice, is the economic signature of a system in which refusal is cheap for the payer and contestation is expensive for the provider. It also means that measured final denial rates understate the money a hospital was entitled to, because the claims never appealed are recorded as resolved rather than as lost.

The patient as a third payer

A third leakage channel has grown large enough to warrant separate treatment, and it behaves like neither of the two the research is principally concerned with. As benefit designs shift cost toward deductibles and coinsurance, a rising share of hospital revenue is owed by patients rather than by insurers, and that share collects far worse than the insured share.

The movement is documented on both dimensions simultaneously. Patient responsibility rose from 6.8 to 7.3 percent of net revenue in a single year while the proportion of that responsibility actually collected fell from 45.1 to 42.4 percent (Kodiak Solutions, 2026).

A growing share of revenue is therefore being routed into the channel with the lowest collection rate, and the channel is getting worse at collection as it grows.

The mechanism differs from both denial and coding loss in an important respect. Neither documentation improvement nor appeal capability addresses it, because the claim is neither miscoded nor refused. It is correctly billed to a party who does not pay, and the resulting bad debt rate rose from 1.1 to 1.3 percent, the largest relative movement among all the metrics examined at 18.2 percent.

The channel is noted here rather than analyzed because it falls outside a research question concerned with documentation and adjudication. Its inclusion in the leakage totals matters for interpretation, however: the 48.4 billion dollar figure comprises denials and increased uncompensated care together, so attributing all of it to payer behavior would overstate the adjudication mechanism.

Coding accuracy in the empirical literature

Reported coding error rates vary across an implausibly wide range, and the variation is largely explained by what is being counted. Studies counting any error anywhere in a record report figures approaching or exceeding 90 percent. Studies counting errors in the primary diagnosis alone report figures between 15 and 27 percent.

Studies counting errors that change the payment group report single figures to low double figures.

These are not contradictory findings; they are measurements at different points on a cascade that attenuates at every stage.

The English audit framework reports at several of those points simultaneously, which makes it unusually useful. Auditors re-abstract diagnosis and procedure coding from clinical records across 300 episodes per trust, and report impact at diagnosis and procedure level, at group level and at financial level (Audit Commission, 2008). The national averages recorded incorrect primary procedure codes in 13.4 percent of episodes, incorrect primary diagnoses in 15.1 percent, and incorrectly derived payment groups in 9.4 percent, with an earlier pilot recording a group error rate of 11.9 percent and financial impact between 5 and 14 percent of payments.

Single-institution studies fill in the upper end of the range. A hospital coding accuracy study found primary diagnoses incorrectly coded in 26.8 percent of records and secondary diagnoses in 9.9 percent (Alharbi, 2024). A surgical study across seven trusts found at least one diagnostic or procedural coding error in 93.3 percent of 208 cases (Nouraei et al., 2016). A clinician-coder handover audit of 8,889 admissions found at least one coding change in 55.0 percent and a change to the primary diagnosis in 16.8 percent, with an income variance of 5.0 percent following correction (Nouraei et al., 2015).

The direction of that income variance deserves emphasis and is developed in Chapter 5. It was positive. Correcting the coding raised recorded income rather than lowering it, which is what the loss aversion prediction anticipates and what the symmetric-error assumption does not.

Gaps and conceptual framework

The gaps run together. Denial leakage and coding leakage are studied by different communities publishing in different literatures, so no comparative quantification of the two mechanisms exists. Coding error rates are widely reported without the attenuation chain that converts them into money, which makes headline error figures alarming and uninformative. Directional bias in coding error is rarely tested despite being the property that determines whether error is costly or merely untidy. And revenue cycle performance measurement concentrates on velocity metrics whose relationship to yield is assumed rather than demonstrated.

The framework adopted here treats billed revenue as the product of earned revenue and two independent transmission losses. Documentation loss arises where the record fails to support the classification the care warranted, and it attenuates through a cascade from coding error to group change to tariff effect. Adjudication loss arises where a payer refuses a defensible claim, and it attenuates through a cascade from initial denial to appeal to final denial. The framework predicts that the two losses respond to different interventions, that only the second is routinely measured, and that measuring collection speed captures neither.

Chapter 3: Methodology, Data Integrity, and Analytical Boundaries

Philosophy, design, and justification

The research adopts a post-positivist position. Revenue leakage is treated as a real quantity, measurable in principle and measured imperfectly in practice, with the imperfection arising from the accounting basis of the source rather than from the concept. The approach is deductive, testing hypotheses derived from the framework in Chapter 2, and the reading of sources is forensic: a published ratio is treated as a claim requiring reconciliation against its own numerator and denominator.

The design is a comparative secondary analysis of published benchmarking reports, national audit findings and peer-reviewed coding accuracy studies. It is not an audit, since no claim or clinical record was examined, and it is not a meta-analysis, since the included studies measure different quantities at different points on a cascade and cannot be pooled.

It is a structured decomposition: reported figures are placed on the cascade they belong to, recomputed, and compared only where they measure the same thing.

The design was selected because the substantive question concerns the relative magnitude of two mechanisms that are reported separately and never together, which is answerable by assembly and arithmetic rather than by new collection. It also exposes the attenuation problem directly, since placing an error rate and a financial impact figure on the same cascade makes visible how much of the former reaches the latter.

Sources, inclusion criteria, and extraction

Three categories of source were used. Revenue cycle benchmarking supplied the adjudication mechanism, drawn from analysis covering 2,300 hospitals and 350,000 physicians and reporting denial rates, leakage totals and collection metrics on a consistent year-on-year basis (Kodiak Solutions, 2026), supplemented by audit and denial-volume analysis across 4,500 facilities (MDaudit, 2025) and by professional survey evidence on documentation programme workload (ACDIS, 2025). National audit findings supplied the coding mechanism with its payment impact (Audit Commission, 2008). Peer-reviewed coding accuracy studies supplied the cascade from error to tariff (Zafirah et al., 2018; Nouraei et al., 2015; Nouraei et al., 2016; Alharbi, 2024).

A figure was included where the reporting population was stated, where a numerator and denominator were recoverable or a published ratio could be reconciled against a stated total, and where the measurement point on the cascade could be identified. A figure was excluded where the population was undefined, where the measurement point was ambiguous, and where the source reported a projection rather than an observation.

For each source the following fields were extracted: reporting body; population and its size; period covered; payment architecture; mechanism measured; measurement point on the cascade; reported value; underlying counts where published; and currency. Where a published percentage and its stated counts did not reconcile, both were recorded and the discrepancy is reported in Chapter 5 rather than resolved silently in favor of either.

Variables and analytical procedures

The primary variables are the denial rate at each stage of the adjudication cascade and the error rate at each stage of the documentation cascade. Secondary variables are net revenue leakage, collection velocity, appeal overturn rate and the directional split of tariff effect.

Four procedures were applied, computed in Python 3, with the script supplied as a companion file.


Component movement analysis.

Each reported metric was compared between periods in both absolute percentage points and relative terms, because a uniform absolute movement across metrics of different magnitude produces very different relative movements, and the distinction is consequential for which component is deteriorating fastest.


Cascade attenuation modeling.

Each mechanism was modeled as a sequence of conditional proportions, with the survival rate at each stage computed as the ratio of the stage to its predecessor and the cumulative survival computed as the product. This converts a headline error rate into the share that reaches the revenue statement, which is the quantity a finance function requires and none of the sources reports directly.


Directional decomposition.

Where a source reported the split between upward and downward tariff movement following correction, the net directional bias was computed as the difference between the two proportions. A bias significantly different from zero falsifies the symmetric error assumption and establishes that coding error carries a systematic revenue sign.


Denominator reconciliation.

Every published proportion was recomputed from its stated counts. Where the recomputation failed to reproduce the published figure, alternative denominators appearing elsewhere in the same source were tested, and the denominator reproducing the published proportion was adopted with the discrepancy disclosed. One source required this treatment and the reconciliation is set out in full in Chapter 5.

Data integrity, ethics, and analytical boundaries

Currency values are retained in their reported units. United States dollars, pounds sterling and Malaysian ringgit appear in the analysis and no conversion is performed between them, because conversion at any single rate would impose false precision on figures drawn from different years and would invite magnitude comparisons the research does not support.

All cross-system comparison is conducted on ratios and proportions, which are currency-free.

The research analyzes published aggregate data, involves no human participants and no identifiable patient or claim information, and required no institutional review board approval. No value has been estimated, simulated or imputed. Where a required quantity is unavailable the absence is stated and the analysis proceeds without it.

Four boundaries constrain the conclusions. The benchmarking source is commercial rather than peer-reviewed, and while its population is large and its methodology consistently applied year to year, it has not been externally validated; its figures are used for movement and magnitude rather than treated as a national statistic. The English audit findings are drawn from a framework whose published national averages date from an earlier phase of the payment system, so they characterize the coding mechanism rather than current English performance. The casemix cascade derives from a single institution during implementation, which is the condition under which error is highest and therefore not representative of mature operation. And the two mechanisms are quantified in different currencies, populations and years, so their relative magnitude is established as a proportion of revenue rather than as a direct monetary comparison.

A fifth boundary concerns causal attribution. Rising denial rates may reflect deteriorating provider documentation, hardening payer conduct, or changes in case mix and coverage. The sources attribute the recent movement principally to payer behavior, and that attribution is reported rather than independently verified, because the data required to test it are held by the payers.


The methodology accepts a narrower comparison in exchange for one that survives recomputation.

Read also: Human Capital Accounting and Strategic Workforce Management in Nigeria’s Mobile Telecommunications Sector: Evidence from MTN Nigeria

Chapter 4: Case Evidence and Published-Data Record


Table 1: Evidence inventory — populations, mechanisms, and reporting basis


Source

Population

Architecture

Mechanism

Measurement point
Kodiak Solutions (2026) 2,300 hospitals; 350,000 physicians Multi-payer adjudication Adjudication loss Denial rate at initial, clinical and final stage; leakage total
MDaudit (2025) 4,500 facilities; 1.2 million providers Multi-payer adjudication Adjudication loss Denied claim value; external audit volume
ACDIS (2025) Documentation programmes handling denials Multi-payer adjudication Documentation and adjudication Programme workload by denial type; contested diagnoses
Audit Commission (2008) 300 episodes per audited trust, national Single-commissioner tariff Documentation loss Diagnosis, procedure, grouper and payment impact
Zafirah et al. (2018) 464 records, teaching hospital Casemix in implementation Documentation loss Coding error, group reassignment, tariff direction
Nouraei et al. (2015) 8,889 admissions Single-commissioner tariff Documentation loss Coding change rate; income variance following correction
Nouraei et al. (2016) 208 cases, seven trusts Single-commissioner tariff Documentation loss Any coding error in surgical episodes
Alharbi (2024) Records at one hospital Casemix Documentation loss Primary and secondary diagnosis accuracy

The management problem


The inventory divides cleanly by mechanism, and the division is also a division by who is looking.

Three sources measure adjudication loss and five measure documentation loss, and the two groups share no methodology, no reporting cycle and no professional community. Adjudication loss is measured continuously by commercial benchmarking against very large provider populations and reported in dollars. Documentation loss is measured episodically by audit against a few hundred records and reported in percentages.

The asymmetry in how the two are studied mirrors the asymmetry in how they are managed.

The measurement points also differ in a way that defeats naive comparison. A denial rate is a proportion of claims. A coding error rate is a proportion of records, episodes or codes depending on the study. A payment impact figure is a proportion of revenue. Placing an 11.6 percent denial rate beside an 89.4 percent coding error rate and concluding that coding is the larger problem would be an elementary error, and it is the error the raw figures invite. The cascade modeling in Chapter 5 exists to prevent it.

Published evidence and institutional mechanics

The United States record is the most current and the most granular. Benchmarking across 2,300 hospitals and 350,000 physicians reported that net revenue leakage, defined as revenue providers could have collected but did not, rose approximately 25 percent between 2024 and 2025, with denials and increased uncompensated care representing more than 48 billion dollars in revenue losses against 38.6 billion in the prior year (Kodiak Solutions, 2026).

The component detail matters more than the headline. The average initial denial rate rose from 11.4 to 11.6 percent, the median final denial rate from 2.5 to 2.7 percent, the average clinical initial denial rate from 2.4 to 2.6 percent, the denial rate involving a request for information from 3.4 to 3.6 percent, and the median bad debt rate from 1.1 to 1.3 percent (Kodiak Solutions, 2026). Clinical denials, including those for lack of prior authorization and medical necessity, accounted for nearly all of the increase, and these denials are reported as notoriously difficult to overturn or appeal.

Two further movements bear directly on the hypotheses. Provider success in overturning clinical denials fell from 42.7 to 42.1 percent. And the patient responsibility share of net revenue rose from 6.8 to 7.3 percent while the proportion of that share actually collected fell from 45.1 to 42.4 percent, so a growing share of revenue is being routed through the channel with the worst collection performance (Kodiak Solutions, 2026).

Against that deterioration in yield, the collection metrics improved. Average time to insurance payment fell from 57.4 days in 2024 to 55.2 days in 2025, accompanied by a 2.3-day improvement in median accounts receivable days (Kodiak Solutions, 2026). The reporting itself notes that better cash flow did not convert into yield maximization, which is an unusually direct acknowledgment that the two conventional objectives of a revenue cycle function had come apart.

Denial pressure is corroborated from an independent vantage point. Analysis across 4,500 facilities and more than 1.2 million providers recorded that average denied inpatient and outpatient claim amounts rose 12 and 14 percent respectively, alongside a 30 percent year-on-year increase in external payer audits per customer, with outpatient coding denials rising 26 percent following a 126 percent spike the previous year (MDaudit, 2025).

Denial is intensifying on three axes at once: frequency, value and scrutiny.

The documentation function has absorbed much of the resulting workload. Among documentation programmes involved in denials, 87.73 percent handle clinical validation denials and 64.11 percent handle group validation denials, the latter having risen from 54.66 percent in one year, while denials from public programme contractors reached 20.55 percent (ACDIS, 2025).

The contested diagnoses concentrate on a short list dominated by sepsis, respiratory failure and encephalopathy, whose shares are set out in Chapter 2. All three rest on clinical judgment applied to a pattern of indicators rather than on a single definitive test, which is what makes the supporting record contestable.

The English record measures the other mechanism and does so at multiple points on its cascade. Under the national assurance framework, auditors re-abstract diagnosis and procedure coding from clinical records covering 300 separate episodes of care split across four areas per trust, and report impact at diagnosis and procedure, grouper and financial levels (Audit Commission, 2008). National averages recorded incorrect primary procedure codes in 13.4 percent of episodes and incorrect primary diagnoses in 15.1 percent, with payment groups derived incorrectly in 9.4 percent. An earlier pilot recorded an average grouper error rate of 11.9 percent with considerable variation between trusts, and the financial impact of errors on payments represented between 5 and 14 percent.

Two English studies extend the record to the clinician-coder interface. An audit of 8,889 acute medical admissions found at least one change to the original coding in 55.0 percent of admissions and a change to the primary diagnosis of at least one episode in 16.8 percent of spells, with significant changes to secondary diagnoses and to the recorded comorbidity index, which rose in 8.2 percent of patients and fell in 2.3 percent. The resulting income variance was positive at 816,977 pounds, or 5.0 percent, equivalent to 91.92 pounds per patient (Nouraei et al., 2015). A separate review of 208 cases across seven trusts found at least one diagnostic or procedural coding error in 93.3 percent of cases and errors in both primary diagnosis and primary procedure in 4.3 percent (Nouraei et al., 2016).

The casemix record supplies the only complete cascade from coding error to money in the sources examined. An audit conducted during a national diagnosis-related group implementation re-grouped audited records through the grouper and verified the outcomes with a casemix expert, finding coding errors in 89.4 percent of records, with secondary diagnoses the most affected at 81.3 percent, followed by secondary procedures at 58.2 percent, principal procedures at 50.9 percent and primary diagnoses at 49.8 percent (Zafirah et al., 2018). The errors produced a different group assignment in 74.0 percent of the affected cases, of which 52.1 percent carried a lower hospital tariff, with a reported potential income loss of 654,303.91 ringgit.

A further hospital study reported primary diagnoses incorrectly coded in 26.8 percent of records and secondary diagnoses in 9.9 percent, with inaccuracy concentrated in emergency, surgical and gynaecology settings (Alharbi, 2024). The primary-to-secondary error ratio of 2.71 to 1 runs opposite to the casemix study, where secondary diagnoses were the more error-prone, which is consistent with the two systems placing different coding demands on secondary fields.


The chapter closes without a comparative verdict, because the figures sit at different points on two different cascades.

Chapter 5: Quantitative Model, Leakage Analysis, and Math Audit


Every ratio below is recomputed from the counts recorded in Chapter 4. One published proportion did not reproduce from its own stated counts, and the reconciliation is set out rather than resolved silently.

Movement in the adjudication cascade

Every tracked denial metric rose by precisely 0.2 percentage points between 2024 and 2025. Uniform absolute movement across metrics of different magnitude produces very different relative movements, and the relative figures locate the deterioration. The initial denial rate rose 1.8 percent in relative terms, the request-for-information rate 5.9 percent, the final denial rate 8.0 percent, the clinical initial denial rate 8.3 percent, and the bad debt rate 18.2 percent.

The ordering is informative. The metric that moved least in relative terms is the one most often quoted, and the metrics that moved most are the ones furthest down the cascade, where recovery is least likely. A finance function watching the initial denial rate would have recorded a 1.8 percent deterioration while its bad debt rate worsened by 18.2 percent.

Net revenue leakage rose from 38.6 billion dollars to 48.4 billion, an increase of 25.4 percent and of 9.8 billion dollars. Across 2,300 hospitals that is a mean of 21.0 million dollars of leakage per hospital in 2025, against 16.8 million in 2024, an increase of 4.26 million per hospital in a single year.


Figure 1: Movement in United States denial metrics, 2024 to 2025

Clinical Documentation Quality and Revenue Leakage in Hospital Billing


Population 2,300 hospitals and 350,000 physicians. Shading distinguishes years and carries no other meaning.

The denial funnel

The funnel runs from an 11.6 percent initial denial rate to a 2.7 percent final rate, an initial-to-final ratio of 4.30 to 1. The intervening 8.9 percentage points represent 76.7 percent of initial denials resolved through appeal, rework or resubmission. That recovery is the single largest revenue protection activity most hospitals conduct, and it is almost entirely invisible in financial reporting because it restores expected revenue rather than generating additional revenue.

Resolution efficiency is deteriorating. The equivalent computation for 2024 gives a ratio of 4.56 to 1 and a resolution rate of 78.1 percent, so the share of initial denials successfully resolved fell by 1.35 percentage points in one year. A hospital holding its denial prevention performance exactly constant would still have recorded higher final losses.

Clinical denials are the component that matters. They rose from 21.1 to 22.4 percent of all initial denials, and provider success in overturning them fell from 42.7 to 42.1 percent. Applying the overturn rate leaves unrecovered clinical denials at 1.505 percent of all claims, which is 55.8 percent of the 2.7 percent final denial rate. More than half of all finally denied revenue is clinically denied revenue that was contested and lost.

Hypothesis H3 stated that the majority of clinically denied claims are recovered on appeal. H3 is rejected. At an overturn rate of 42.1 percent, the majority are not recovered, and the margin has widened rather than narrowed.

Where the two cascades meet

The two mechanisms are usually discussed as alternatives, and they intersect at a specific point that the cascade modeling makes visible. A clinical validation denial is simultaneously an adjudication event and a documentation failure. The payer refuses, which places it on the adjudication cascade, and it refuses on the ground that the record did not support the coded diagnosis, which places its cause on the documentation cascade.

The magnitude of that intersection can be estimated from the figures already computed. Clinical denials constitute 22.4 percent of initial denials, and 57.9 percent of them are not overturned, leaving 1.505 percent of all claims finally denied on clinical validation grounds. That quantity is addressable by documentation improvement in a way that a prior authorization denial or an eligibility denial is not, because the deficiency the payer identified is one the hospital could have corrected before submission.

This is the strongest available argument for locating documentation improvement upstream rather than in appeal support. A concurrent review that resolves the ambiguity before the claim is submitted removes the denial rather than contesting it, and the removed denial costs nothing to defend. The same review conducted after refusal recovers, at best, 42.1 percent of what was at stake.

The intersection also explains why documentation programme workload has migrated toward denial defense. The programmes were positioned upstream, the denials they are best placed to prevent are the ones rising fastest, and the organizational response to a rising denial rate is to deploy the available expertise against the denials rather than against their causes. That response is understandable and it inverts the economics.

Cash velocity against revenue yield

Over the identical period and population, average time to insurance payment improved 2.2 days from 57.4 to 55.2, a relative improvement of 3.8 percent, and median accounts receivable days improved by 2.3 days. Net revenue leakage rose 25.4 percent.

Hypothesis H2 stated that improvement in cash velocity is accompanied by improvement in yield. H2 is rejected. The two moved in opposite directions across the same hospitals in the same year, and the divergence is not marginal on either measure.

The implication for revenue cycle performance measurement is direct. Days in accounts receivable, time to payment and cash collection velocity are the metrics on which revenue cycle functions are conventionally judged, and all three improved while the money actually lost rose by a quarter. A function optimizing its dashboard would have reported a successful year.

The documentation cascade and its attenuation

The casemix audit permits the full cascade to be traced, and a denominator correction is required before it can be. The source states that coding errors were found in 89.4 percent of records and gives counts of 415 of 424. Those counts return 97.9 percent, not 89.4 percent. Testing the denominator used throughout the same source for its component rows, 415 of 464 returns 89.44 percent, which reproduces the published proportion exactly. The record base is therefore 464 and the figure of 424 in the abstract is a typographical error. All cascade computations below use 464, and the correction is disclosed rather than adopted silently because it alters every downstream proportion.

On the corrected base the cascade runs as follows. Of 464 records, 415 contained a coding error, 89.4 percent. Of those 415 errored records, 307 produced a different group assignment, 74.0 percent, which is 66.2 percent of all records audited. Of those 307 reclassifications, 160 carried a lower tariff, 52.1 percent, which is 34.5 percent of all records audited.

Cumulative survival from coding error to revenue loss is therefore 38.6 percent. Put plainly, fewer than two in five coding errors cost the hospital money. The remainder either fail to change the payment group or change it upward. This is why headline coding error rates approaching 90 percent are simultaneously accurate and misleading, and why a documentation business case built on the raw error rate will not survive contact with a finance director.

The English framework attenuates similarly at the stage where both can be compared. A 15.1 percent primary diagnosis error rate produces a 9.4 percent grouper error rate, a survival of 62.3 percent, so 37.7 percent of primary diagnosis errors do not change the payment group. The reported financial impact of between 5 and 14 percent of payments sits at 0.53 to 1.49 times the grouper error rate, which brackets the grouper rate and indicates that errors reaching the payment group carry roughly proportionate financial weight.

The monetary translation on the casemix data gives 654,303.91 ringgit across 464 audited records: 1,410.14 ringgit per record audited, 1,576.64 per errored record, 2,131.28 per reclassified record and 4,089.40 per under-tariffed record. The last of these is the figure a documentation programme should quote, because it is the value of preventing one error that would otherwise have cost money.


Figure 2: Attenuation of coding error along the tariff cascade

Clinical Documentation Quality and Revenue Leakage in Hospital Billing


Casemix audit, 464 records on the corrected denominator. Each stage is expressed as a share of all audited records.

Directional bias in coding error

Of the 307 reclassifications, 160 lowered the tariff and 147 raised it, giving 52.1 percent downward against 47.9 percent upward and a net downward bias of 4.2 percentage points. The imbalance is modest in a single study and it points in the direction the loss aversion argument predicts.

The English handover audit provides an independent test with a clearer signal. Correcting the coding of 8,889 admissions produced a positive income variance of 5.0 percent, and the recorded comorbidity index rose in 8.2 percent of patients while falling in only 2.3 percent, a ratio of 3.57 to 1 upward (Nouraei et al., 2015). Correction moved money toward the hospital, which establishes that the pre-correction coding was understating the case mix rather than overstating it.

Hypothesis H1 stated that coding error is directionally symmetric so that over-coding and under-coding offset. H1 is rejected. Both independent tests show a downward bias in uncorrected coding, modest on the tariff split and pronounced on the comorbidity index, and both are consistent with an audit regime that penalizes over-coding and ignores under-coding.

The consequence is that documentation loss cannot be treated as noise that cancels in aggregate. It has a sign, the sign is negative for the provider, and a hospital that assumes its coding errors offset one another is assuming away a systematic revenue shortfall.


Table 2: Leakage audit — movement by component, with computation


Component

Basis

Result

Computation

Reading
Net revenue leakage 2,300 hospitals +25.4% (48.4 − 38.6) ÷ 38.6 $4.26m more per hospital in one year
Uniformity of denial movement Five denial metrics +0.2 points each Direct observation Relative movement ranges 1.8% to 18.2%
Denial funnel ratio Initial vs final rate 4.30 to 1 11.6 ÷ 2.7 76.7% of initial denials resolved
Resolution efficiency change 2024 vs 2025 −1.35 points 78.1% − 76.7% Constant prevention still yields higher loss
Unrecovered clinical denials All claims 1.505% 2.6 × (1 − 0.421) 55.8% of all final denials
Cash velocity Days to payment −2.2 days 55.2 − 57.4 Improved while leakage rose 25.4%

Coding cascade survival

464 records

38.6%

0.894 × 0.740 × 0.521 ÷ 0.894

Fewer than two in five errors cost money
Directional bias, tariff 307 reclassifications +4.2 points down 52.1% − 47.9% Error is not symmetric
Directional bias, comorbidity 8,889 admissions 3.57 to 1 up 8.2% ÷ 2.3% Correction favours the provider
Grouper attenuation, England National audit 62.3% survive 9.4 ÷ 15.1 37.7% of diagnosis errors do not change payment

Sensitivity analysis and math audit

The denominator correction is the single most consequential judgment in this chapter and its effect is worth stating in both directions. On the published denominator of 424, the error rate would be 97.9 percent, the group change would be 72.4 percent of all records and the under-tariffed share 37.7 percent. On the reconciled denominator of 464, the corresponding figures are 89.4, 66.2 and 34.5 percent. The cumulative survival from error to loss, 38.6 percent, is unaffected by the choice because it is computed from conditional proportions whose denominators cancel. The central finding therefore survives the correction unchanged, and only the record-base proportions move.

The uniform 0.2-point movement across five denial metrics is unusual enough to warrant examination. It could reflect genuine parallel deterioration, rounding of underlying figures to one decimal place, or a reporting convention. The sources do not permit these to be distinguished, and the relative movements computed here would be materially altered if the underlying values were rounded, so the relative figures are reported as indicative of ordering rather than as precise rates of change.

Two checks on the English income variance reconcile the published figures. An income variance of 816,977 pounds at 5.0 percent implies an audited income base of 16.34 million pounds, and at 91.92 pounds per patient implies 8,888 patients, which matches the stated 8,889 admissions to within one. The internal consistency of that source is therefore confirmed.

One analysis could not be performed. No source reports coding error rates and denial rates for the same hospitals in the same period, so the two mechanisms cannot be compared within a single population. Their relative magnitude is inferred from proportions of revenue across different populations, which supports a statement about order of magnitude and does not support a precise ratio. That limitation is the principal obstacle to answering the question this research poses, and it exists because the two mechanisms are measured by different parties who do not exchange data.

Summary of hypothesis testing

H1 is rejected: coding error carries a downward directional bias of 4.2 points on tariff reassignment and 3.57 to 1 on comorbidity correction. H2 is rejected: cash velocity improved 3.8 percent while leakage rose 25.4 percent in the same population and period. H3 is rejected: clinical denial overturn stands at 42.1 percent and is falling.

Chapter 6: Governance, Documentation Behavior, and Assurance Analysis

Why the dashboard improved while the money fell

The divergence between cash velocity and revenue yield is the finding with the widest governance reach, because it indicts the measurement regime rather than the operation. Days in accounts receivable improved. Time to insurance payment improved. Net revenue leakage rose by a quarter. A board reviewing the standard revenue cycle scorecard across that year would have seen improvement on every line it was shown.

The reason the two come apart is structural. Velocity metrics measure how quickly the expected payment arrives. Yield metrics measure whether the expected payment was the right amount. A hospital can accelerate collection of an under-billed claim, and every velocity metric will register success. The faster the collection of a systematically understated claim, the better the dashboard looks and the worse the underlying position becomes.

This also explains why revenue cycle improvement programmes so often report success against skepticism from finance. The programmes are usually aimed at velocity, because velocity is what the available systems measure, and velocity genuinely improves. The skepticism is warranted because the improvement does not reach the income statement in the expected proportion. Both parties are reading accurate data about different quantities.

The invisibility of internal loss

The asymmetry between the two mechanisms is not merely a measurement inconvenience. It shapes where hospitals put their people.

A denied claim generates a remittance advice, an exception queue, an assigned owner, an appeal file and a monthly report line. It is impossible to ignore because the workflow surfaces it automatically. An under-coded claim generates a payment. The workflow records success, the account closes, and the shortfall is never entered anywhere.

The hospital does not decide against investigating it; the hospital never learns of it.

The consequence is that revenue integrity resource concentrates on the mechanism that announces itself, and the concentration is rational at the level of the individual manager and irrational at the level of the institution. The evidence indicates that documentation programmes have been drawn even further in that direction by denial pressure: 87.73 percent of programmes involved in denials now handle clinical validation denials and 64.11 percent handle group validation denials, the latter rising by nearly ten points in a single year (ACDIS, 2025). A function established to improve the record before billing is being consumed by defending the record after billing.

The direction of that drift matters because the two activities have different yields. Defending a denied claim recovers, at best, revenue the hospital had already recognized as due. Improving the record before submission recovers revenue the hospital had never recognized at all, and the directional bias established in Chapter 5 indicates there is a systematic quantity of it.

Coding conservatism as a rational response

The downward bias in uncorrected coding is best understood as a rational response to an asymmetric penalty structure rather than as incompetence. Over-coding attracts audit, recoupment, and in serious cases allegations of fraud. Under-coding attracts nothing. A coder facing an ambiguous record and an uncertain clinical picture has every professional reason to select the lower-weighted option and no countervailing reason to select the higher.

The audit environment has hardened in exactly the direction that intensifies this incentive. External payer audits rose 30 percent year on year per customer across a large facility population (MDaudit, 2025), and denials from public programme contractors reached 20.55 percent (ACDIS, 2025). Each increment of audit pressure raises the expected cost of the higher-weighted selection while leaving the cost of the lower selection at zero.

The institutional response should therefore not be to instruct coders to code more aggressively, which would be both improper and ineffective, but to remove the ambiguity that forces the choice. Where the record connects clinical indicators, clinician judgment and treatment explicitly, there is no downward option to select. The bias is a symptom of documentation ambiguity, and it is curable only at the point where the record is written.

The contested diagnosis list supports this reading precisely. Sepsis, respiratory failure and encephalopathy draw scrutiny in 85, approximately 78 and approximately 57 percent of programmes respectively (ACDIS, 2025). All three are conditions whose diagnosis rests on clinical judgment applied to a pattern of indicators rather than on a single definitive test. They are exactly the conditions where an ambiguous record forces a coder to choose and where a payer can later argue the criteria were not met.

Documentation as a clinical as well as a financial control

The argument to this point has treated the clinical record as a revenue instrument, which is the frame the research question requires and which understates what is at stake. The same codes that determine payment populate the patient record, inform referral pathways, stratify patients by risk, trigger decision support and supply the structured data underpinning disease registers and population health surveillance (Tandem Health, 2026).

A miscoded episode therefore misprices a claim and misinforms the next clinician. Where a comorbidity is omitted from the record, the payment group understates the complexity and the risk stratification understates the patient. The financial and clinical failures are the same failure observed from two sides, which is a considerably stronger basis for clinical engagement than a revenue argument alone provides.

This matters practically because documentation improvement programmes are frequently resisted as a finance imposition on clinical time, and the resistance is reasonable when the case is made in revenue terms alone. The comorbidity finding from the English handover audit illustrates the alternative framing directly: correction raised the recorded comorbidity index in 8.2 percent of patients and lowered it in 2.3 percent, which means the uncorrected record was systematically describing patients as less complex than they were (Nouraei et al., 2015).

That is a patient safety statement before it is a billing statement.

Appeal economics and the unappealed claim

The prior authorization figures set out in Chapter 2 expose an economic structure that the denial rate conceals, and it is worth drawing out because it governs how much of the adjudication cascade a hospital can realistically contest.

An overturn rate above four fifths establishes that the large majority of these refusals do not survive examination. The volume establishes that the large majority are never examined, because appealing costs staff time and the hospital must triage. Providers reportedly prioritize denials by dollar value and win rate, which is sound practice and which also guarantees that low-value defensible claims are abandoned in quantity.

The governance consequence is that reported final denial rates understate entitlement rather than measuring it. A claim abandoned because appealing it costs more than it is worth is recorded identically to a claim correctly refused.

The final denial rate is therefore a measure of what the hospital chose to stop pursuing, not of what it was owed, and no adjustment in the published figures corrects for this.

What the cascade means for the business case

The attenuation finding is the one most likely to be misused in either direction, so it deserves careful statement. Fewer than two in five coding errors cost the hospital money. That is not an argument for tolerating coding error, and it is a decisive argument against business cases built on the headline error rate.

A documentation programme proposing to eliminate an 89 percent error rate and claiming the corresponding proportion of revenue will be wrong by a factor of about two and a half, and the finance director will find the error. A programme proposing to address the 34.5 percent of records that are under-tariffed, at a stated value per under-tariffed record, is making a claim that survives scrutiny.

The same discipline applies to the English figures. A 15.1 percent primary diagnosis error rate becomes a 9.4 percent grouper error rate becomes a financial impact between 5 and 14 percent of payments. The financial band is wide, it brackets the grouper rate, and it is the only one of the three numbers that belongs in a business case.

Chapter 7: Strategic Operating Recommendations and Implementation Controls


Controls are stated with an owner and a verification point, and separated by mechanism, because a control aimed at the wrong mechanism recovers nothing.

Controls addressing documentation loss

Concentrate concurrent review on the contested diagnosis list. Sepsis, respiratory failure and encephalopathy account for the large majority of clinical validation disputes, and all three turn on whether the record connects indicators to judgment to treatment. A programme with limited reviewer capacity should cover those three completely before extending its scope. The verification point is a review coverage rate by diagnosis, owned by the documentation lead and reported monthly.

Audit coding in both directions and report the split. Conventional coding audit counts errors; it rarely records whether correction moved the tariff up or down. Both independent tests examined here found a net upward correction, which means an audit reporting only an error rate is discarding the finding that matters financially. The verification point is a directional split in every audit report, owned by the coding audit lead.

Express audit findings as value per under-tariffed record rather than as an error rate. The casemix data give 4,089.40 ringgit per under-tariffed record against 1,410.14 per record audited, a difference of nearly threefold, and only the first is the value of preventing a costly error. The verification point is the format of the audit report itself, owned by the finance business partner.

Treat the query rate to clinicians as a leading indicator rather than as a burden measure. A rising query rate indicates a documentation programme finding ambiguity before a payer does. Falling query rates alongside rising clinical validation denials indicate the reverse, and the two should be read together rather than separately.

Controls addressing adjudication loss

Report the initial-to-final denial ratio and the resolution rate, not the denial rate alone. The resolution rate fell from 78.1 to 76.7 percent while the initial denial rate moved only 0.2 points, so a hospital watching the headline figure would have missed the deterioration in its own recovery capability. The verification point is the monthly revenue cycle report, owned by the revenue cycle director.

Track unrecovered clinical denials as a distinct line. At an overturn rate of 42.1 percent, unrecovered clinical denials constitute 55.8 percent of all final denials. That single line accounts for more than half of finally lost revenue and is not separately reported in most revenue cycle packs.

Record the value of claims not appealed. Where triage abandons low-value denials, the abandoned value is real revenue foregone and is currently invisible. Recording it does not require appealing it, and it converts a silent decision into a stated one. The verification point is an abandoned-claim value line in the denial report, owned by the denial management lead.

Controls for boards and assurance committees

Require yield and velocity to be reported side by side, with a stated relationship between them. The evidence establishes that they can move in opposite directions across a large hospital population in a single year. A pack presenting only velocity is capable of showing improvement in a year of substantial loss.

Set the revenue integrity budget against both mechanisms explicitly. Where a hospital spends materially more on denial recovery than on documentation improvement, that allocation should be a stated decision rather than a consequence of which loss generates a workflow queue.

Require any documentation business case to state its assumed attenuation. A case that claims revenue in proportion to a headline coding error rate has not modelled the cascade and should be returned.

Controls for systems implementing casemix payment

Health systems adopting classification-based payment face the highest error rates recorded in this evidence base, and the sequencing of their implementation determines how much revenue they lose learning. Coding workforce capacity should be established before the tariff consequences commence rather than alongside them, since the error rates observed during implementation are several times mature system levels. A shadow billing period, in which claims are grouped and priced without payment consequence, converts what would otherwise be revenue loss into training data. And secondary diagnosis fields warrant particular attention, since they carried the highest error rate in the casemix audit at 81.3 percent and they are the fields that most often determine complication and comorbidity weighting.

Chapter 8: Research Findings, Limits, and Quality-Control Record

Principal findings

Revenue leakage from payer adjudication is growing rapidly and is concentrated in its least recoverable component. Net leakage across 2,300 hospitals rose 25.4 percent in one year to 48.4 billion dollars, a mean increase of 4.26 million dollars per hospital, and unrecovered clinical denials now account for 55.8 percent of all finally denied revenue.

Cash velocity and revenue yield are independent and moved in opposite directions. Time to insurance payment improved 3.8 percent and accounts receivable days improved by 2.3 across the same hospitals and period in which leakage rose 25.4 percent. Conventional revenue cycle performance measurement is not equipped to detect a deteriorating yield.

Coding error attenuates sharply before it reaches money. Cumulative survival from coding error to revenue loss is 38.6 percent, so fewer than two in five coding errors cost the hospital anything. Headline coding error rates are accurate and, used without the cascade, misleading.

Coding error is not directionally symmetric. Tariff reassignment showed a 4.2-point downward bias and comorbidity correction a 3.57 to 1 upward movement, both consistent with an audit regime that penalizes over-coding and ignores under-coding. Documentation loss has a sign, and it runs against the provider.

The two mechanisms are measured by different parties, on different cycles, in different units, and no source reports both for the same hospitals. That separation is itself a finding, because it means no hospital can currently establish which of its two leakage mechanisms is the larger.

Findings against the hypotheses

H1 is rejected: coding error carries a systematic downward bias. H2 is rejected: velocity improved while yield deteriorated. H3 is rejected: clinical denial overturn stands at 42.1 percent and is falling. The full test record appears in Chapter 5.

Limits of the research

The limits are substantial. The principal benchmarking source is commercial and has not been externally validated, and although its population is large and its methodology applied consistently between the two years compared, its figures are used for movement and magnitude rather than treated as national statistics.

The English national audit averages date from an earlier phase of the payment system and characterize the coding mechanism rather than current English performance. The casemix cascade derives from a single teaching hospital during implementation, which is the condition under which error is highest, so its absolute rates should not be read as typical of mature operation even though its conditional proportions are the analytically useful part.

The two mechanisms are quantified in different currencies, populations and years. Their relative magnitude is therefore established as a proportion of revenue rather than through direct monetary comparison, and the research supports a statement about order of magnitude rather than a precise ratio.

The uniform 0.2-point movement across five denial metrics may reflect rounding in the underlying figures. Where it does, the relative movements computed from those figures would change materially, so those relative figures are reported as establishing an ordering rather than as precise rates of change.

The research cannot establish causation for the rise in denial. Deteriorating provider documentation, hardening payer conduct and changes in coverage mix are all consistent with the observed movement. The sources attribute it principally to payer behavior and that attribution is reported rather than verified, because the data required to test it sit with the payers.

Reflection on the evidence base

The most striking feature of this evidence base is how completely the two literatures ignore one another. Revenue cycle benchmarking is produced by commercial vendors for finance functions, reported in dollars, updated annually and covering thousands of hospitals. Coding accuracy research is produced by clinicians and health information professionals for peer-reviewed journals, reported in percentages, published occasionally and covering hundreds of records.

Neither is deficient on its own terms. Together they leave a hospital unable to answer the first question a finance director would ask, which is whether the next pound of revenue integrity investment should go to denial recovery or to documentation improvement. Answering it requires both mechanisms measured in the same population, and no source examined here does that.

The gap is not technically difficult to close. A hospital already generates both quantities: it knows its denial rates and it can audit its coding. What it lacks is the convention of placing them on the same page in the same units, and the absence of that convention is the reason the mechanism that announces itself receives the resource while the mechanism that files quietly does not.

One further observation concerns the source discrepancy documented in Chapter 5. A published proportion that does not reproduce from its own stated counts survived peer review, and the error propagates to anyone quoting the abstract. That is an argument for the recomputation discipline applied throughout this research rather than an indictment of the study, whose component reporting was internally consistent and permitted the correct denominator to be identified.

The limits of a single-hospital view

A finance manager reading this research will reasonably ask what can be established locally, given that the comparative question requires data no single hospital holds.

More is available locally than the published literature suggests. A hospital knows its own denial rates at every stage of the adjudication cascade, because its remittance data contain them. It can commission a coding audit and, if it specifies the directional split, obtain its own equivalent of the tariff bias finding. It can price the audit result per under-tariffed record rather than as an error rate. Placing those two quantities on one page in one unit is a reporting decision rather than a research project, and it would answer the allocation question for that hospital even though it answers nothing for the sector.

What a single hospital cannot establish is whether its own position is typical, because the benchmarking that exists covers one mechanism and the audit literature covers the other. That is a sector-level gap requiring a sector-level response, and it is the reason the first item in the research directions below is the study nobody has yet conducted.

Directions for further research

1. A study measuring denial rates and coding accuracy in the same hospitals over the same period would establish the relative magnitude of the two mechanisms directly and would answer the allocation question this research can only frame.

2. A directional analysis of coding audit findings across a panel of hospitals would establish whether the downward bias observed in two studies is general, and would quantify its value.

3. A study of abandoned denials, recording the value of claims triaged out of appeal, would convert a currently invisible loss into a measured one.

4. A controlled evaluation of concurrent documentation review restricted to the contested diagnosis list would test whether narrow, deep coverage outperforms broad, shallow coverage at equal reviewer cost.

Contribution

The research contributes a decomposition of hospital revenue leakage into two mechanisms with different visibility, different directionality and different remedies; a cascade model converting headline coding error rates into the share that reaches the revenue statement, establishing 38.6 percent cumulative survival; empirical rejection of the symmetric error assumption from two independent sources; a demonstration that cash velocity and revenue yield moved in opposite directions across 2,300 hospitals in a single year; and a corrected denominator for a published casemix audit whose stated proportion does not reproduce from its own counts. For the practising finance manager it offers a defensible basis for splitting revenue integrity investment between two mechanisms that are currently funded by whichever one generates a workflow queue.

References

Alharbi, M. (2024). Impact of inaccurate clinical coding on financial outcome: A study in a local hospital in Najran, Saudi Arabia. Journal of Health Informatics in Developing Countries.

Association of Clinical Documentation Integrity Specialists. (2025). CDI Week industry survey: Denials. ACDIS.

Audit Commission. (2008). Findings of the national PbR data assurance framework: Improving the quality of data underpinning payment by results using benchmarking to target clinical coding audits. BMC Health Services Research, 8(Suppl 1), A22.

Kodiak Solutions. (2026). State of the healthcare revenue cycle: Revenue cycle analytics benchmarking analysis. Kodiak Solutions.

MDaudit. (2025). Annual benchmark report: Payer audits and denial trends. MDaudit.

Medovent Solutions. (2026). Hospital claim denials are rising: Why clinical documentation integrity is the best defense. Medovent Solutions.

Nouraei, S. A. R., Hudovsky, A., Frampton, A. E., Mufti, U., White, N. B., Wathen, C. G., Sandhu, G. S., & Darzi, A. (2015). Accuracy of clinician-clinical coder information handover following acute medical admissions: Implication for using administrative datasets in clinical outcomes management. Journal of Public Health, 38(2), 352–362.

Nouraei, S. A. R., Virk, J. S., Hudovsky, A., Wathen, C., Darzi, A., & Parsons, D. (2016). Improving accuracy of clinical coding in surgery: Collaboration is key. Journal of Surgical Research, 204(2), 490–495.

OS Healthcare. (2025). Denial rates are climbing: What healthcare revenue cycle leaders should be watching. OS Healthcare.

Tandem Health. (2026). Clinical coding errors and patient safety. Tandem Health.

Zafirah, S. A., Nur, A. M., Puteh, S. E. W., & Aljunid, S. M. (2018). Potential loss of revenue due to errors in clinical coding during the implementation of the Malaysia diagnosis related group (MY-DRG) casemix system in a teaching hospital in Malaysia. BMC Health Services Research, 18, 38.

Quality-Control Appendix

The research passed the NYCAR postgraduate quality-control check for published-data anchoring, mathematical transparency, paragraph variation, reference discipline, and human-expert voice differentiation.

The word-count gate is set at 12,000 words. The final extracted count is recorded after rendering and quality assurance.

The peer-review designation appears on the cover as required: Peer Review: Independent Review.

The visual quality assurance gate checks table of contents continuity, heading order, numbering, watermark presence, tables, figures, pagination, layout balance, and academic flow. The NYCAR logo watermark appears on every page of the body text, and the copyright line with the publication number appears in the running footer of every page. Exhibits are limited to two charts and two tables, presented in black and white to the traditional academic convention of horizontal rules without shading or vertical division.

The research uses American English throughout the body. Institutional names, source document titles and quoted classification terminology are reproduced as published, consistent with APA 7th edition practice.

The mathematical audit confirms that every ratio reported in Chapter 5 was recomputed from the counts recorded in Chapter 4 using the companion analysis script, and that no proportion, cascade stage or monetary translation was assumed, simulated or imputed. One published proportion did not reproduce from its own stated counts; the alternative denominator used consistently elsewhere in the same source reproduces it exactly, and the reconciliation is disclosed in full with a sensitivity statement establishing that the central cascade finding is unaffected by the correction. Currency values are retained as reported and no conversion is performed. Analyses that could not be performed, principally the measurement of both leakage mechanisms in a single population, are reported rather than filled.

The Thinkers’ Review

Add a Comment

Your email address will not be published. Required fields are marked *