Healthcare & Clinical Trials Guide

AI in Medical Coding

How ICD-10, CPT, MedDRA, and WHODrug coding is being rebuilt around AI, what works in production, what still needs a credentialed human, and how to roll it out without breaking your audit trail.

BY CAROLINA BOSCH · DIRECTOR, R&D CLINICAL DATA

What AI Medical Coding Is

Medical coding turns clinical language into standardized codes, ICD-10-CM/PCS, CPT, HCPCS for reimbursement and health records, MedDRA for adverse events, and WHODrug for medications in clinical trials. It is the layer between what a clinician wrote and what a payer, regulator, or statistician can act on.

AI-assisted coding uses NLP and large language models to read the source document and propose the codes. A credentialed coder reviews, edits, and signs off. Autonomous coding, no human in the loop for a defined set of encounter types, is production reality now for a narrow slice; assistive coding is the norm everywhere else.

Healthcare vs. Clinical Trials

On the healthcare side the pressure is reimbursement: correct DRG assignment, clean claims, fewer denials, and audit defensibility against payers and CMS. Speed matters, but accuracy under audit matters more.

On the clinical trials side the pressure is regulatory: consistent MedDRA and WHODrug coding across a global study, clean safety signals, and a sponsor-auditable trail for the FDA, EMA, and PMDA. The vocabularies are stricter and the sign-off model is non-negotiable.

The AI stack is similar. The governance is not.

Where AI Actually Fits

Autonomous coding (narrow)

Radiology, pathology, and routine outpatient encounters where the note structure is tight and code space is bounded. High-volume, low-ambiguity, high-confidence auto-submit.

Assistive inpatient coding

Complex admissions get AI-proposed principal and secondary diagnoses ranked with rationale and evidence spans. The coder edits, sequences, and signs off.

MedDRA auto-coding

Verbatim AE and medical history terms mapped to MedDRA LLTs with confidence scores. The medical coder confirms or overrides; overrides become training signal.

WHODrug mapping

Concomitant medications normalized to WHODrug preferred names, resolving trade names, formulations, and misspellings a rules engine misses.

CDI at the point of care

Clinical documentation integrity suggestions surfaced to the clinician while the note is still being written, reducing rework and query volume downstream.

Denial prediction & appeals

Models flag claims likely to be denied before submission and draft appeal letters citing chart evidence when they are.

Accuracy, Audits & Denials

The right accuracy target is not "beats the average coder", it is "beats the audit". That means measuring code-level precision and recall against a gold-standard human panel, tracked by encounter type, by service line, and over time as guidelines change.

Denials tell you the same story from the payer side. Track denial rate, denial reason, and appeal success on AI-coded versus human-coded claims. If AI-coded denials shift toward medical-necessity or documentation categories, the fix is upstream, CDI and clinician prompts, not the coder.

"The model proposes. The credentialed coder disposes. Every code that leaves the building still needs a human signature and a trail back to the note."

Governance & Compliance

In healthcare the guardrails are HIPAA, the OIG compliance program guidance, and payer-specific audit rules. In clinical trials the guardrails are ICH E6, 21 CFR Part 11, EU Annex 11, and the sponsor's MedDRA and WHODrug versioning policy.

Across both, the non-negotiables are the same: models are validated for their intended use, every AI-proposed code is reviewable and reversible by a human, and the system emits an audit trail that ties each code to the source text, the model version, and the reviewer. Dictionary versions (MedDRA quarterly, WHODrug biannually, ICD-10 annually) must be pinned per study or per coding period, silent upgrades break comparability.

Getting Started

  1. 1

    Pick one encounter type

    Radiology, pathology, or a single therapeutic area for MedDRA. Bounded scope is the difference between a pilot that ships and a pilot that stalls.

  2. 2

    Build a gold-standard panel

    Two or three credentialed coders re-code a stratified sample. That is your ground truth for accuracy, drift, and vendor comparison, not the vendor's marketing benchmark.

  3. 3

    Deploy assistive first

    Every AI-proposed code goes through a human. Measure agreement, edit distance, and time saved. Only expand to autonomous coding where sustained agreement clears your audit bar.

  4. 4

    Pin your dictionaries

    MedDRA version, WHODrug version, ICD-10 year, CPT year, locked per study or per coding period, with a documented upgrade protocol. This is where quiet regressions live.

  5. 5

    Instrument the audit trail

    Source span, proposed code, confidence, model version, reviewer, final code, timestamp. If a payer or regulator asked tomorrow, you should be able to replay a single code end to end.

Rolling out AI coding at your org?

I am open to senior leadership roles in AI, Healthcare & Life Sciences, and Developer Ecosystems, plus select consulting engagements on AI-assisted coding, MedDRA/WHODrug governance, and audit-ready deployment.