Problems → Prototypes · AILE

← Back to this problem brief

A Clinical Reasoning + AI Calibration Coach

Clinical trainees increasingly encounter AI-supported decisions, but learning to use AI safely requires more than knowing what a tool can output. Learners must integrate foundational knowledge, uncertainty, patient context, and the possibility that the AI is wrong.

Try it

This exercise runs locally using simple rules. Your entries stay on this page and clear when you reload. No AI service is called.

Critical review of this working prototype

Strongest assumption: Independent assessment before advice teaches calibrated source checking rather than categorical trust or distrust.

Likely failure: Learners always challenge AI, mistake an administrative record match for clinical competence, or revise the supposedly independent response after viewing advice.

Test that could change the design: Use unseen fictional packets with supported, incomplete and conflicting advice; compare with unaided source checking and blind human review of explanations. Reject if participants always challenge advice regardless of evidence.

Non-AI comparison: Paper handoff packets with deliberately fallible peer summaries and facilitated source checking.

Evidence to collect: Structured record comparisons, source citations, confidence changes and independently reviewed transfer reflections. No clinical-competence or free-text score; no patient-outcome inference.

Human control and access: Entirely fictional administrative records and scripted advice. Explicitly not clinician validated; no diagnoses or treatment guidance. No patient data requested, no network storage, and exports require deliberate action. Native controls and untimed practice support keyboard access.

Changes made during review

  • Included a supported bounded claim alongside misleading and incomplete suggestions to counter blanket distrust.
  • Required three independent reflections plus confidence before advice; disabled committed fields and clear downstream results when revising.
  • Marked post-advice reassessments in exports, including revisits to a case within the same run; reset explicitly cannot undo exposure.
  • Added structured source checks, unscored rationale, revision/export and explicit clinician-validation limitation.
  • Next smallest experiment: clinician and educator review of source-check wording before a supervised unseen-record transfer study.

This is a design and implementation review, not an empirical validation of learning outcomes.

Optional AI critique

Take the brief into your own AI environment

Nothing is sent until you choose.

Copy a self-contained review prompt, then open the AI workspace you already use. The prompt asks for a falsification test, a non-AI alternative, evidence that distinguishes activity from learning, and human-system risks.

Preview the prompt
Act as a critical learning-engineering reviewer. Review the prototype brief below as a thin-slice experiment, not as a finished product.

Return:
1. The strongest design assumption.
2. The most important plausible failure mode.
3. The highest-value falsification test for the next cycle.
4. One credible non-AI alternative that could address the same capability gap.
5. The next smallest prototype worth building.
6. What should be measured to distinguish engagement from learning, transfer, and real performance.
7. Human-agency, equity, accessibility, privacy, or governance concerns that should change the design.

Be concrete. Separate evidence-backed claims from hypotheses. Do not reward novelty for its own sake, and do not assume more AI is better.

COLLECTION: AILE
TITLE: A Clinical Reasoning + AI Calibration Coach

CAPABILITY GAP:
Clinical trainees increasingly encounter AI-supported decisions, but learning to use AI safely requires more than knowing what a tool can output. Learners must integrate foundational knowledge, uncertainty, patient context, and the possibility that the AI is wrong.

FIRST PROTOTYPE:
Create a clinical reasoning simulator in which an AI assistant sometimes offers useful, incomplete, or misleading suggestions. The learner must state an independent assessment, identify what evidence would change the decision, and explicitly answer 'what might I be missing?' before seeing expert feedback.

REPRESENTATIVE USE CASE:
A simulated patient case includes an AI-generated differential diagnosis that omits a less common but consequential possibility. The learner decides whether to accept, challenge, or investigate the suggestion.

LEARNING / HUMAN-SYSTEM FRAME:
Simulation with deliberate practice can support complex professional judgment when feedback is explicit and repeated. Metacognitive prompts make uncertainty and self-monitoring visible; the AI is deliberately treated as fallible evidence, not an authority.

QUESTIONS ALREADY IDENTIFIED:
1. How often should the AI be wrong to teach calibration without creating distrust?
2. Which measures capture reasoning quality rather than answer matching?
3. Does practice change behavior in authentic clinical settings?

FIRST-CYCLE FRAMING:
Understand: Validate that the real gap is the ability to integrate AI advice into independent, evidence-based professional judgment, not simply low engagement, low tool use, or a workflow inconvenience.
Map: Map the system: trainee, supervisor, patient case, clinical knowledge, AI assistant, institutional policy, and safety constraints. Identify where the capability currently succeeds, breaks down, or is masked by other constraints.
Instrument: Predefine evidence: diagnostic reasoning, appropriate challenge of AI, uncertainty calibration, transfer, safety-critical misses, and explanation quality. A key disconfirming signal is: learners learn to distrust or obey the AI categorically instead of calibrating trust to evidence.
Open the underlying design brief and eight-step path

A first prototype

Create a clinical reasoning simulator in which an AI assistant sometimes offers useful, incomplete, or misleading suggestions. The learner must state an independent assessment, identify what evidence would change the decision, and explicitly answer 'what might I be missing?' before seeing expert feedback.

Representative use case

A simulated patient case includes an AI-generated differential diagnosis that omits a less common but consequential possibility. The learner decides whether to accept, challenge, or investigate the suggestion.

Questions worth carrying forward

  • How often should the AI be wrong to teach calibration without creating distrust?
  • Which measures capture reasoning quality rather than answer matching?
  • Does practice change behavior in authentic clinical settings?

Evidence anchors

These sources motivate the design; they do not validate this prototype.

LENS iteration cycle

One possible first pass

Refine → Understand
  1. 01
    Understand

    Validate that the real gap is the ability to integrate AI advice into independent, evidence-based professional judgment, not simply low engagement, low tool use, or a workflow inconvenience.

  2. 02
    Map

    Map the system: trainee, supervisor, patient case, clinical knowledge, AI assistant, institutional policy, and safety constraints. Identify where the capability currently succeeds, breaks down, or is masked by other constraints.

  3. 03
    Design

    Compare an AI intervention with simpler non-AI options. Specify what the human decides, what the AI may suggest, and which tradeoffs are acceptable.

  4. 04
    Build

    Create the smallest usable prototype around one representative task, with expert-curated content, visible uncertainty, and an easy human override.

  5. 05
    Instrument

    Predefine evidence: diagnostic reasoning, appropriate challenge of AI, uncertainty calibration, transfer, safety-critical misses, and explanation quality. A key disconfirming signal is: learners learn to distrust or obey the AI categorically instead of calibrating trust to evidence.

  6. 06
    Deploy

    Pilot with a small, representative group in a near-real setting; preserve a baseline or comparison condition and document implementation conditions.

  7. 07
    Evaluate

    Look for capability growth and transfer, not just satisfaction or activity. Inspect subgroup patterns, human workload, errors, and unintended adaptations.

  8. 08
    Refine

    Let the evidence change the problem model. Default back to UNDERSTAND if the assumed gap was wrong; otherwise revisit the earliest step invalidated by the evidence.