Problems → Prototypes · LENS Canonical
Human-AI Delegation Simulator
When AI enters a workflow, the capability requirement changes: people must decide when to rely, when to verify, when to override, and when to revoke delegation. Accuracy alone does not establish that the human-AI team is safe or effective.
This exercise runs locally using simple rules. Your entries stay on this page and clear when you reload. No AI service is called.
Three fictional workplace-support cases. Advice, confidence cues, policies, and consequences are authored examples. No model is running. Your rationale is recorded for reflection and is not machine-scored.
Your authority in this case
AI recommendation
What happened in this fictional case
Review your decisions
Compare your reasons with the evidence each case revealed. Which information changed the appropriate level of delegation? Would the same choice hold in a new case?
Critical review of this working prototype
Strongest assumption: Contrasting cases help learners distinguish source quality, confidence wording, and delegated authority.
Likely failure: Learners can memorize action labels or escalate every case without learning calibrated reliance.
Test that could change the design: Use unseen cases with different confidence cues and reversed surface details; compare reasoning and decisions to a paper scenario exercise.
Non-AI comparison: Facilitated tabletop cases with printed evidence cards and a decision log.
Evidence to collect: Record action, rationale, source inspection before decision, confidence, unnecessary handoffs, and performance on unfamiliar cases.
Human control and access: The authored consequences are illustrative, not empirical predictions; free-text rationale is not automatically scored.
Changes made during review
- Include a low-confidence routine case where accepting advice is reasonable, a stale-source case, and an authority-boundary case.
- Require independent written reasoning before revealing consequences.
- Freeze submitted choices and record whether the source was inspected beforehand.
- Export the decision record with the fictional-scenario boundary and provide a complete reset.
This is a design and implementation review, not an empirical validation of learning outcomes.
Take the brief into your own AI environment
Copy a self-contained review prompt, then open the AI workspace you already use. The prompt asks for a falsification test, a non-AI alternative, evidence that distinguishes activity from learning, and human-system risks.
Preview the prompt
Act as a critical learning-engineering reviewer. Review the prototype brief below as a thin-slice experiment, not as a finished product. Return: 1. The strongest design assumption. 2. The most important plausible failure mode. 3. The highest-value falsification test for the next cycle. 4. One credible non-AI alternative that could address the same capability gap. 5. The next smallest prototype worth building. 6. What should be measured to distinguish engagement from learning, transfer, and real performance. 7. Human-agency, equity, accessibility, privacy, or governance concerns that should change the design. Be concrete. Separate evidence-backed claims from hypotheses. Do not reward novelty for its own sake, and do not assume more AI is better. COLLECTION: LENS Canonical TITLE: Human-AI Delegation Simulator CAPABILITY GAP: When AI enters a workflow, the capability requirement changes: people must decide when to rely, when to verify, when to override, and when to revoke delegation. Accuracy alone does not establish that the human-AI team is safe or effective. FIRST PROTOTYPE: Create a scenario simulator in which AI advice varies in confidence, quality, and context. Learners practice delegation, verification, override, and escalation while the system records not only outcomes but the reasoning behind reliance decisions. REPRESENTATIVE USE CASE: An analyst receives AI-ranked cases with explanations of uneven quality. Some recommendations are right for the wrong reason; others are uncertain. The learner must decide what to accept, inspect, or escalate before seeing consequences. LEARNING / HUMAN-SYSTEM FRAME: Human agency is an engineered property of the team. The design should make authority, uncertainty, and correction visible, and test calibrated reliance rather than simple compliance or distrust. QUESTIONS ALREADY IDENTIFIED: 1. Which decisions may be delegated and under what conditions? 2. Can the human detect when the AI is outside its competence? 3. Does practice improve calibrated reliance without increasing unnecessary workload? FIRST-CYCLE FRAMING: Understand: Define the shared task, stakes, authority boundaries, and failure modes of the human-AI team. Map: Map information flow, AI capabilities and limits, human expertise, time pressure, accountability, override paths, and recovery mechanisms. Instrument: Measure decision quality, calibration of reliance, verification behavior, override appropriateness, recovery from AI error, workload, and performance when the AI is unavailable.
Open the underlying design brief and eight-step path
A first prototype
Create a scenario simulator in which AI advice varies in confidence, quality, and context. Learners practice delegation, verification, override, and escalation while the system records not only outcomes but the reasoning behind reliance decisions.
Representative use case
An analyst receives AI-ranked cases with explanations of uneven quality. Some recommendations are right for the wrong reason; others are uncertain. The learner must decide what to accept, inspect, or escalate before seeing consequences.
Questions worth carrying forward
- Which decisions may be delegated and under what conditions?
- Can the human detect when the AI is outside its competence?
- Does practice improve calibrated reliance without increasing unnecessary workload?
Evidence anchors
These sources motivate the design; they do not validate this prototype.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). (2023)
Risk, governance, measurement, and accountable deployment.
- Amershi, S., et al. Guidelines for Human-AI Interaction. CHI 2019. (2019)
Human-AI interaction design and calibrated reliance.
- U.S. Department of Education, Office of Educational Technology. Artificial Intelligence and the Future of Teaching and Learning: Insights and Recommendations. (2023)
Human-centered AI in education; governance and instructional use.
- UNESCO. AI Competency Framework for Students. (2024)
Student AI literacy, critical judgment, and co-creation.
- Hattie, J., & Timperley, H. The Power of Feedback. Review of Educational Research. (2007)
Feedback design and evidence-producing practice.
One possible first pass
- 01 Understand
Define the shared task, stakes, authority boundaries, and failure modes of the human-AI team.
- 02 Map
Map information flow, AI capabilities and limits, human expertise, time pressure, accountability, override paths, and recovery mechanisms.
- 03 Design
Compare an AI intervention with simpler non-AI options. Specify what the human decides, what the AI may suggest, and which tradeoffs are acceptable.
- 04 Build
Create the smallest usable prototype around one representative task, with expert-curated content, visible uncertainty, and an easy human override.
- 05 Instrument
Measure decision quality, calibration of reliance, verification behavior, override appropriateness, recovery from AI error, workload, and performance when the AI is unavailable.
- 06 Deploy
Pilot with a small, representative group in a near-real setting; preserve a baseline or comparison condition and document implementation conditions.
- 07 Evaluate
Look for capability growth and transfer, not just satisfaction or activity. Inspect subgroup patterns, human workload, errors, and unintended adaptations.
- 08 Refine
Let the evidence change the problem model. Default back to UNDERSTAND if the assumed gap was wrong; otherwise revisit the earliest step invalidated by the evidence.