29 August 2026 · Human-AI Teaming × Clinical Trials

Why did AI improve trial-screening accuracy without saving time?

A 30-minute pack on randomized human-AI prescreening, criterion-level error analysis, automation bias, workflow efficiency and how to evaluate AI in statistical programming.

DifficultyC1
Time30 minutes
Main study355 patients
OutputEvaluation design

Why this reading

A better model does not automatically create a faster workflow.

Human+AI prescreening improved chart-level accuracy relative to Human-alone, but average review time remained almost unchanged.

For SP automation, that is a valuable warning: quality, efficiency, safety and traceability are separate endpoints and should be measured separately.

Reading order

Your 30-minute plan.

0-3 minPreview

Find the primary accuracy endpoint and secondary efficiency endpoint.

3-14 minMain study

Read randomization, gold-standard review, accuracy, review time and automation bias.

14-20 minTrialMatchAI

Compare end-to-end retrieval and criterion-level classification.

20-25 minCDISC

Connect evaluation to reproducibility, traceability and measurable impact.

25-30 minOutput

Design quality and efficiency endpoints for one SP AI feature.

Open-access sources

Randomized workflow evidence plus an end-to-end AI system and current standards context.

Brief background

Accuracy improved. Time did not.

The randomized evaluation used 355 de-identified oncology charts from patients with non-small cell lung or colorectal cancer.

Clinical research coordinators reviewed charts with AI augmentation or without it, and their decisions were compared against an expert-developed gold standard.

Human+AI chart-level accuracy was about 76% compared with about 71% for Human-alone. The improvement was strongest for biomarker, staging and response-related criteria.

Average review time was roughly 37 minutes per chart in both groups. The expected efficiency gain did not materialize.

The study also highlights automation bias: augmentation can introduce new human error pathways if reviewers over-trust machine suggestions.

For SP teams, evaluate quality, efficiency, failure modes and traceability independently.

Key vocabulary

Fifteen terms for human-AI workflow evaluation.

Term中文Meaning / use
prescreening预筛选An early review step used to identify patients who may satisfy clinical-trial eligibility criteria.
eligibility criterion入排标准条目A rule determining whether a patient can enter a clinical trial.
noninferiority trial非劣效性试验A design testing whether a new approach is not unacceptably worse than a comparator.
superiority test优效性检验A test asking whether one method performs better than another.
chart-level accuracy病历层面准确率The proportion of complete patient charts classified correctly against a gold standard.
gold standard金标准The trusted reference used to judge whether an AI or human assessment is correct.
neurosymbolic AI神经符号人工智能AI combining neural pattern recognition with explicit symbolic or rule-based reasoning.
automation bias自动化偏误The tendency to over-trust an automated recommendation even when it is wrong.
human-AI teaming人机协作A workflow in which humans and AI jointly complete a task rather than one replacing the other.
workflow efficiency工作流效率How much time or operational effort a process requires.
criterion-level classification条目级分类A decision made separately for each eligibility rule rather than only at whole-chart level.
biomarker criterion生物标志物标准An eligibility rule based on a molecular or laboratory characteristic.
staging criterion分期标准An eligibility rule based on the extent or stage of disease.
traceability可追溯性The ability to connect an output to its source evidence, rules, metadata, and review history.
measurable impact可衡量影响A demonstrable improvement in quality, efficiency, reuse, conformance, or another defined outcome.

Useful phrases

Language for workflow-validation discussions.

  1. improve accuracy without reducing review time - The system improved accuracy without reducing review time.
  2. evaluate the combined human-AI workflow - The trial evaluated the combined human-AI workflow.
  3. compare performance against a gold standard - All prescreening decisions were compared against a gold standard.
  4. define accuracy and efficiency as separate endpoints - Accuracy and efficiency should be defined as separate endpoints.
  5. identify where automation bias can occur - Teams should identify where automation bias can occur.
  6. keep a human reviewer in the decision loop - The workflow keeps a human reviewer in the decision loop.
  7. measure impact beyond model-level performance - Production validation should measure impact beyond model-level performance.
  8. use criterion-level error analysis - Criterion-level error analysis reveals where the system actually helps.
  9. preserve traceability from source evidence to decision - The system should preserve traceability from source evidence to decision.
  10. treat time saved as an empirical outcome - Time saved should be treated as an empirical outcome, not an assumption.

Comprehension

Five questions.

  1. Why were accuracy and review time treated as separate endpoints?
  2. What does unchanged review time reveal about model capability versus workflow efficiency?
  3. Why might biomarker and staging criteria benefit more from AI assistance?
  4. How can automation bias reduce the value of an accurate system?
  5. Which metrics should be used for an AI-assisted SAS or ADaM review workflow?

Retelling

Say it three times.

  • 30 seconds · Study design → accuracy result → efficiency result.
  • 45 seconds · Explain why Human+AI can be better without being faster.
  • 60 seconds · Apply the lesson to SAS log review, SDTM mapping, ADaM review or TFL QC.

5-minute output task

Design an evaluation for one SP AI feature.

  1. Minute 1: Choose log review, SDTM mapping, ADaM review, TFL QC, review triage or code generation.
  2. Minutes 2-3: Define separate quality, efficiency, safety and traceability endpoints.
  3. Minute 4: Choose Human-alone, AI-alone and/or Human+AI comparison arms.
  4. Minute 5: State the threshold required before production scaling.

One sentence to keep

An AI workflow is valuable only when the combined human-machine system improves the outcomes that matter; model accuracy, time saved, safety, and traceability must be measured separately rather than assumed to move together.