AI-Driven TLF
Automation
How Tables, Listings, and Figures are moving from a six-week biostatistics sprint to a days-long, agent-assisted pipeline without giving up double programming or the audit trail.
BY CAROLINA BOSCH · DIRECTOR, R&D CLINICAL DATA
What TLFs Actually Are
TLFs, Tables, Listings, and Figures, are the standardized statistical outputs that make up the bulk of a Clinical Study Report. Demographics tables, adverse-event summaries, subject-level listings, Kaplan–Meier curves: hundreds per study, each traced back to a SAP-defined shell and produced from CDISC ADaM datasets.
They are the deliverable regulators actually read. Any change in how they are produced needs to preserve the same statistical logic, the same double programming discipline, and the same audit trail.
Why TLFs Are Slow
Most sponsors still translate TLF shells into SAS by hand, then double-program each output for QC. Every SAP amendment, every new safety cut, every regulator question triggers another cycle. It is not unusual for a Phase 3 CSR to burn six to eight weeks between database lock and final TLF package, with biostatistics as the critical-path bottleneck.
The work is repetitive, template-heavy, and unusually well-specified. That combination is exactly where AI is credible today.
Where AI Fits in TLF Production
Shell-to-code translation
LLMs turn TLF shells and mock-ups into first-draft SAS or R programs against ADaM specs, hours of scaffolding compressed to minutes, with a human refining the statistical logic.
Automated QC drafting
Models cross-check outputs against ADaM datasets, spot mismatches in counts and percentages, and draft QC comments for the double programmer to accept or reject.
SAP-to-shell drafting
Statistical Analysis Plan sections become draft TLF shells with variable mappings, the biostatistician reviews and locks the shells before any code runs.
Cycle-time forecasting
Historical study data trains models that predict TLF package completion given SAP complexity and headcount, surfacing risk before database lock, not after.
Narrative drafting
For patient narratives and safety summaries, LLMs draft the prose from ADaM plus SDTM, a medical writer reviews, edits, and signs off.
Traceability graphs
Agents maintain the SAP → shell → program → output chain automatically, so any regulator question about a specific number can be answered with the exact code path in seconds.
Agentic TLF Pipelines
A useful mental model: the TLF pipeline stops being a Gantt chart of hand-offs and becomes a queue of proposed outputs. Signals flow in from the SAP, the ADaM specs, and prior study patterns. Agents draft the shell, draft the program, run it, cross-check it against a second draft, and attach a QC report. A human biostatistician triages the queue, approves, edits, or rejects.
What that changes operationally: fewer end-of-study fire drills, weeks of calendar time recovered before submission, and, crucially, a traceability graph that shows why a number is what it is, not just that it was signed off.
"The model drafts. The biostatistician signs. Every TLF still needs a human, a signature, and a trail."
CDISC & 21 CFR Part 11 Expectations
Regulators care about the validated output and its traceability, not whether a human or a model drafted the first line of code. The same discipline that governs hand-written SAS applies to AI-drafted code: independent double programming, defined test data, versioned scripts, and reproducible outputs from locked ADaM.
For AI-assisted TLF automation that means three non-negotiables: outputs are reproducible from CDISC-conformant datasets, every AI-drafted program is reviewed and validated by a qualified programmer, and the whole pipeline emits an audit trail compatible with 21 CFR Part 11 and Annex 11. Skip any of those and you have a demo, not a submission.
Getting Started
- 1
Pick one output family
Demographics and disposition tables are the honest starting point: high volume, well-specified, low statistical risk. Prove the pipeline there before touching efficacy tables.
- 2
Lock your ADaM first
AI drafts collapse when specs move. Freeze ADaM structures for the pilot; save creative interpretation for after the pipeline is trusted.
- 3
Keep double programming
Have the AI draft one arm and a human draft the other, or vice versa. Compare outputs the same way you always have. Adoption comes from matching, not marketing.
- 4
Measure the right thing
Not tokens per second. Measure cycle time from SAP approval to final TLF package, and QC catch rate. Those move VP-level decisions.
- 5
Instrument the audit trail
Every prompt version, every drafted program, every human edit, every final output. If a regulator asked tomorrow, could you reconstruct the story of a single table cell?
Automating TLFs at your org?
I am open to senior leadership roles in AI, Healthcare & Life Sciences, and Developer Ecosystems, plus select consulting engagements on clinical reporting automation, biostatistics tooling, and AI governance in regulated environments.