Problems → Prototypes · LENS Canonical

← Back to this problem brief

Human-AI Delegation Simulator

When AI enters a workflow, the capability requirement changes: people must decide when to rely, when to verify, when to override, and when to revoke delegation. Accuracy alone does not establish that the human-AI team is safe or effective.

Try it

This exercise runs locally using simple rules. Your entries stay on this page and clear when you reload. No AI service is called.

Three fictional workplace-support cases. Advice, confidence cues, policies, and consequences are authored examples. No model is running. Your rationale is recorded for reflection and is not machine-scored.

Your authority in this case

AI recommendation

Your decision

Critical review of this working prototype

Strongest assumption: Contrasting cases help learners distinguish source quality, confidence wording, and delegated authority.

Likely failure: Learners can memorize action labels or escalate every case without learning calibrated reliance.

Test that could change the design: Use unseen cases with different confidence cues and reversed surface details; compare reasoning and decisions to a paper scenario exercise.

Non-AI comparison: Facilitated tabletop cases with printed evidence cards and a decision log.

Evidence to collect: Record action, rationale, source inspection before decision, confidence, unnecessary handoffs, and performance on unfamiliar cases.

Human control and access: The authored consequences are illustrative, not empirical predictions; free-text rationale is not automatically scored.

Changes made during review

  • Include a low-confidence routine case where accepting advice is reasonable, a stale-source case, and an authority-boundary case.
  • Require independent written reasoning before revealing consequences.
  • Freeze submitted choices and record whether the source was inspected beforehand.
  • Export the decision record with the fictional-scenario boundary and provide a complete reset.

This is a design and implementation review, not an empirical validation of learning outcomes.

Optional AI critique

Take the brief into your own AI environment

Nothing is sent until you choose.

Copy a self-contained review prompt, then open the AI workspace you already use. The prompt asks for a falsification test, a non-AI alternative, evidence that distinguishes activity from learning, and human-system risks.

Preview the prompt
Act as a critical learning-engineering reviewer. Review the prototype brief below as a thin-slice experiment, not as a finished product.

Return:
1. The strongest design assumption.
2. The most important plausible failure mode.
3. The highest-value falsification test for the next cycle.
4. One credible non-AI alternative that could address the same capability gap.
5. The next smallest prototype worth building.
6. What should be measured to distinguish engagement from learning, transfer, and real performance.
7. Human-agency, equity, accessibility, privacy, or governance concerns that should change the design.

Be concrete. Separate evidence-backed claims from hypotheses. Do not reward novelty for its own sake, and do not assume more AI is better.

COLLECTION: LENS Canonical
TITLE: Human-AI Delegation Simulator

CAPABILITY GAP:
When AI enters a workflow, the capability requirement changes: people must decide when to rely, when to verify, when to override, and when to revoke delegation. Accuracy alone does not establish that the human-AI team is safe or effective.

FIRST PROTOTYPE:
Create a scenario simulator in which AI advice varies in confidence, quality, and context. Learners practice delegation, verification, override, and escalation while the system records not only outcomes but the reasoning behind reliance decisions.

REPRESENTATIVE USE CASE:
An analyst receives AI-ranked cases with explanations of uneven quality. Some recommendations are right for the wrong reason; others are uncertain. The learner must decide what to accept, inspect, or escalate before seeing consequences.

LEARNING / HUMAN-SYSTEM FRAME:
Human agency is an engineered property of the team. The design should make authority, uncertainty, and correction visible, and test calibrated reliance rather than simple compliance or distrust.

QUESTIONS ALREADY IDENTIFIED:
1. Which decisions may be delegated and under what conditions?
2. Can the human detect when the AI is outside its competence?
3. Does practice improve calibrated reliance without increasing unnecessary workload?

FIRST-CYCLE FRAMING:
Understand: Define the shared task, stakes, authority boundaries, and failure modes of the human-AI team.
Map: Map information flow, AI capabilities and limits, human expertise, time pressure, accountability, override paths, and recovery mechanisms.
Instrument: Measure decision quality, calibration of reliance, verification behavior, override appropriateness, recovery from AI error, workload, and performance when the AI is unavailable.
Open the underlying design brief and eight-step path

A first prototype

Create a scenario simulator in which AI advice varies in confidence, quality, and context. Learners practice delegation, verification, override, and escalation while the system records not only outcomes but the reasoning behind reliance decisions.

Representative use case

An analyst receives AI-ranked cases with explanations of uneven quality. Some recommendations are right for the wrong reason; others are uncertain. The learner must decide what to accept, inspect, or escalate before seeing consequences.

Questions worth carrying forward

  • Which decisions may be delegated and under what conditions?
  • Can the human detect when the AI is outside its competence?
  • Does practice improve calibrated reliance without increasing unnecessary workload?

Evidence anchors

These sources motivate the design; they do not validate this prototype.

LENS iteration cycle

One possible first pass

Refine → Understand
  1. 01
    Understand

    Define the shared task, stakes, authority boundaries, and failure modes of the human-AI team.

  2. 02
    Map

    Map information flow, AI capabilities and limits, human expertise, time pressure, accountability, override paths, and recovery mechanisms.

  3. 03
    Design

    Compare an AI intervention with simpler non-AI options. Specify what the human decides, what the AI may suggest, and which tradeoffs are acceptable.

  4. 04
    Build

    Create the smallest usable prototype around one representative task, with expert-curated content, visible uncertainty, and an easy human override.

  5. 05
    Instrument

    Measure decision quality, calibration of reliance, verification behavior, override appropriateness, recovery from AI error, workload, and performance when the AI is unavailable.

  6. 06
    Deploy

    Pilot with a small, representative group in a near-real setting; preserve a baseline or comparison condition and document implementation conditions.

  7. 07
    Evaluate

    Look for capability growth and transfer, not just satisfaction or activity. Inspect subgroup patterns, human workload, errors, and unintended adaptations.

  8. 08
    Refine

    Let the evidence change the problem model. Default back to UNDERSTAND if the assumed gap was wrong; otherwise revisit the earliest step invalidated by the evidence.