Why this reading
A better model does not automatically create a faster workflow.
Human+AI prescreening improved chart-level accuracy relative to Human-alone, but average review time remained almost unchanged.
For SP automation, that is a valuable warning: quality, efficiency, safety and traceability are separate endpoints and should be measured separately.
Reading order
Your 30-minute plan.
Find the primary accuracy endpoint and secondary efficiency endpoint.
Read randomization, gold-standard review, accuracy, review time and automation bias.
Compare end-to-end retrieval and criterion-level classification.
Connect evaluation to reproducibility, traceability and measurable impact.
Design quality and efficiency endpoints for one SP AI feature.
Open-access sources
Randomized workflow evidence plus an end-to-end AI system and current standards context.
Brief background
Accuracy improved. Time did not.
The randomized evaluation used 355 de-identified oncology charts from patients with non-small cell lung or colorectal cancer.
Clinical research coordinators reviewed charts with AI augmentation or without it, and their decisions were compared against an expert-developed gold standard.
Human+AI chart-level accuracy was about 76% compared with about 71% for Human-alone. The improvement was strongest for biomarker, staging and response-related criteria.
Average review time was roughly 37 minutes per chart in both groups. The expected efficiency gain did not materialize.
The study also highlights automation bias: augmentation can introduce new human error pathways if reviewers over-trust machine suggestions.
For SP teams, evaluate quality, efficiency, failure modes and traceability independently.
Key vocabulary
Fifteen terms for human-AI workflow evaluation.
| Term | 中文 | Meaning / use |
|---|---|---|
| prescreening | 预筛选 | An early review step used to identify patients who may satisfy clinical-trial eligibility criteria. |
| eligibility criterion | 入排标准条目 | A rule determining whether a patient can enter a clinical trial. |
| noninferiority trial | 非劣效性试验 | A design testing whether a new approach is not unacceptably worse than a comparator. |
| superiority test | 优效性检验 | A test asking whether one method performs better than another. |
| chart-level accuracy | 病历层面准确率 | The proportion of complete patient charts classified correctly against a gold standard. |
| gold standard | 金标准 | The trusted reference used to judge whether an AI or human assessment is correct. |
| neurosymbolic AI | 神经符号人工智能 | AI combining neural pattern recognition with explicit symbolic or rule-based reasoning. |
| automation bias | 自动化偏误 | The tendency to over-trust an automated recommendation even when it is wrong. |
| human-AI teaming | 人机协作 | A workflow in which humans and AI jointly complete a task rather than one replacing the other. |
| workflow efficiency | 工作流效率 | How much time or operational effort a process requires. |
| criterion-level classification | 条目级分类 | A decision made separately for each eligibility rule rather than only at whole-chart level. |
| biomarker criterion | 生物标志物标准 | An eligibility rule based on a molecular or laboratory characteristic. |
| staging criterion | 分期标准 | An eligibility rule based on the extent or stage of disease. |
| traceability | 可追溯性 | The ability to connect an output to its source evidence, rules, metadata, and review history. |
| measurable impact | 可衡量影响 | A demonstrable improvement in quality, efficiency, reuse, conformance, or another defined outcome. |
Useful phrases
Language for workflow-validation discussions.
- improve accuracy without reducing review time - The system improved accuracy without reducing review time.
- evaluate the combined human-AI workflow - The trial evaluated the combined human-AI workflow.
- compare performance against a gold standard - All prescreening decisions were compared against a gold standard.
- define accuracy and efficiency as separate endpoints - Accuracy and efficiency should be defined as separate endpoints.
- identify where automation bias can occur - Teams should identify where automation bias can occur.
- keep a human reviewer in the decision loop - The workflow keeps a human reviewer in the decision loop.
- measure impact beyond model-level performance - Production validation should measure impact beyond model-level performance.
- use criterion-level error analysis - Criterion-level error analysis reveals where the system actually helps.
- preserve traceability from source evidence to decision - The system should preserve traceability from source evidence to decision.
- treat time saved as an empirical outcome - Time saved should be treated as an empirical outcome, not an assumption.
Comprehension
Five questions.
- Why were accuracy and review time treated as separate endpoints?
- What does unchanged review time reveal about model capability versus workflow efficiency?
- Why might biomarker and staging criteria benefit more from AI assistance?
- How can automation bias reduce the value of an accurate system?
- Which metrics should be used for an AI-assisted SAS or ADaM review workflow?
Retelling
Say it three times.
- 30 seconds · Study design → accuracy result → efficiency result.
- 45 seconds · Explain why Human+AI can be better without being faster.
- 60 seconds · Apply the lesson to SAS log review, SDTM mapping, ADaM review or TFL QC.
5-minute output task
Design an evaluation for one SP AI feature.
- Minute 1: Choose log review, SDTM mapping, ADaM review, TFL QC, review triage or code generation.
- Minutes 2-3: Define separate quality, efficiency, safety and traceability endpoints.
- Minute 4: Choose Human-alone, AI-alone and/or Human+AI comparison arms.
- Minute 5: State the threshold required before production scaling.
One sentence to keep
An AI workflow is valuable only when the combined human-machine system improves the outcomes that matter; model accuracy, time saved, safety, and traceability must be measured separately rather than assumed to move together.