Problems → Prototypes · AILE
A Clinical Reasoning + AI Calibration Coach
Clinical trainees increasingly encounter AI-supported decisions, but learning to use AI safely requires more than knowing what a tool can output. Learners must integrate foundational knowledge, uncertainty, patient context, and the possibility that the AI is wrong.
This exercise runs locally using simple rules. Your entries stay on this page and clear when you reload. No AI service is called.
Critical review of this working prototype
Strongest assumption: Independent assessment before advice teaches calibrated source checking rather than categorical trust or distrust.
Likely failure: Learners always challenge AI, mistake an administrative record match for clinical competence, or revise the supposedly independent response after viewing advice.
Test that could change the design: Use unseen fictional packets with supported, incomplete and conflicting advice; compare with unaided source checking and blind human review of explanations. Reject if participants always challenge advice regardless of evidence.
Non-AI comparison: Paper handoff packets with deliberately fallible peer summaries and facilitated source checking.
Evidence to collect: Structured record comparisons, source citations, confidence changes and independently reviewed transfer reflections. No clinical-competence or free-text score; no patient-outcome inference.
Human control and access: Entirely fictional administrative records and scripted advice. Explicitly not clinician validated; no diagnoses or treatment guidance. No patient data requested, no network storage, and exports require deliberate action. Native controls and untimed practice support keyboard access.
Changes made during review
- Included a supported bounded claim alongside misleading and incomplete suggestions to counter blanket distrust.
- Required three independent reflections plus confidence before advice; disabled committed fields and clear downstream results when revising.
- Marked post-advice reassessments in exports, including revisits to a case within the same run; reset explicitly cannot undo exposure.
- Added structured source checks, unscored rationale, revision/export and explicit clinician-validation limitation.
- Next smallest experiment: clinician and educator review of source-check wording before a supervised unseen-record transfer study.
This is a design and implementation review, not an empirical validation of learning outcomes.
Take the brief into your own AI environment
Copy a self-contained review prompt, then open the AI workspace you already use. The prompt asks for a falsification test, a non-AI alternative, evidence that distinguishes activity from learning, and human-system risks.
Preview the prompt
Act as a critical learning-engineering reviewer. Review the prototype brief below as a thin-slice experiment, not as a finished product. Return: 1. The strongest design assumption. 2. The most important plausible failure mode. 3. The highest-value falsification test for the next cycle. 4. One credible non-AI alternative that could address the same capability gap. 5. The next smallest prototype worth building. 6. What should be measured to distinguish engagement from learning, transfer, and real performance. 7. Human-agency, equity, accessibility, privacy, or governance concerns that should change the design. Be concrete. Separate evidence-backed claims from hypotheses. Do not reward novelty for its own sake, and do not assume more AI is better. COLLECTION: AILE TITLE: A Clinical Reasoning + AI Calibration Coach CAPABILITY GAP: Clinical trainees increasingly encounter AI-supported decisions, but learning to use AI safely requires more than knowing what a tool can output. Learners must integrate foundational knowledge, uncertainty, patient context, and the possibility that the AI is wrong. FIRST PROTOTYPE: Create a clinical reasoning simulator in which an AI assistant sometimes offers useful, incomplete, or misleading suggestions. The learner must state an independent assessment, identify what evidence would change the decision, and explicitly answer 'what might I be missing?' before seeing expert feedback. REPRESENTATIVE USE CASE: A simulated patient case includes an AI-generated differential diagnosis that omits a less common but consequential possibility. The learner decides whether to accept, challenge, or investigate the suggestion. LEARNING / HUMAN-SYSTEM FRAME: Simulation with deliberate practice can support complex professional judgment when feedback is explicit and repeated. Metacognitive prompts make uncertainty and self-monitoring visible; the AI is deliberately treated as fallible evidence, not an authority. QUESTIONS ALREADY IDENTIFIED: 1. How often should the AI be wrong to teach calibration without creating distrust? 2. Which measures capture reasoning quality rather than answer matching? 3. Does practice change behavior in authentic clinical settings? FIRST-CYCLE FRAMING: Understand: Validate that the real gap is the ability to integrate AI advice into independent, evidence-based professional judgment, not simply low engagement, low tool use, or a workflow inconvenience. Map: Map the system: trainee, supervisor, patient case, clinical knowledge, AI assistant, institutional policy, and safety constraints. Identify where the capability currently succeeds, breaks down, or is masked by other constraints. Instrument: Predefine evidence: diagnostic reasoning, appropriate challenge of AI, uncertainty calibration, transfer, safety-critical misses, and explanation quality. A key disconfirming signal is: learners learn to distrust or obey the AI categorically instead of calibrating trust to evidence.
Open the underlying design brief and eight-step path
A first prototype
Create a clinical reasoning simulator in which an AI assistant sometimes offers useful, incomplete, or misleading suggestions. The learner must state an independent assessment, identify what evidence would change the decision, and explicitly answer 'what might I be missing?' before seeing expert feedback.
Representative use case
A simulated patient case includes an AI-generated differential diagnosis that omits a less common but consequential possibility. The learner decides whether to accept, challenge, or investigate the suggestion.
Questions worth carrying forward
- How often should the AI be wrong to teach calibration without creating distrust?
- Which measures capture reasoning quality rather than answer matching?
- Does practice change behavior in authentic clinical settings?
Evidence anchors
These sources motivate the design; they do not validate this prototype.
- McGaghie, W. C., et al. Does simulation-based medical education with deliberate practice yield better results than traditional clinical education? Academic Medicine. (2011)
Simulation, deliberate practice, and repeated performance with feedback.
- Hattie, J., & Timperley, H. The Power of Feedback. Review of Educational Research. (2007)
Feedback design and evidence-producing practice.
- Education Endowment Foundation. Metacognition and Self-Regulated Learning, 2nd ed. (2025)
Metacognitive prompts, self-monitoring, and fading support.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). (2023)
Risk, governance, measurement, and accountable deployment.
- UNESCO. AI Competency Framework for Students. (2024)
Student AI literacy, critical judgment, and co-creation.
One possible first pass
- 01 Understand
Validate that the real gap is the ability to integrate AI advice into independent, evidence-based professional judgment, not simply low engagement, low tool use, or a workflow inconvenience.
- 02 Map
Map the system: trainee, supervisor, patient case, clinical knowledge, AI assistant, institutional policy, and safety constraints. Identify where the capability currently succeeds, breaks down, or is masked by other constraints.
- 03 Design
Compare an AI intervention with simpler non-AI options. Specify what the human decides, what the AI may suggest, and which tradeoffs are acceptable.
- 04 Build
Create the smallest usable prototype around one representative task, with expert-curated content, visible uncertainty, and an easy human override.
- 05 Instrument
Predefine evidence: diagnostic reasoning, appropriate challenge of AI, uncertainty calibration, transfer, safety-critical misses, and explanation quality. A key disconfirming signal is: learners learn to distrust or obey the AI categorically instead of calibrating trust to evidence.
- 06 Deploy
Pilot with a small, representative group in a near-real setting; preserve a baseline or comparison condition and document implementation conditions.
- 07 Evaluate
Look for capability growth and transfer, not just satisfaction or activity. Inspect subgroup patterns, human workload, errors, and unintended adaptations.
- 08 Refine
Let the evidence change the problem model. Default back to UNDERSTAND if the assumed gap was wrong; otherwise revisit the earliest step invalidated by the evidence.