Problems → Prototypes · LENS Canonical
Evidence-to-Impact Mapper
Organizations often know who completed training but cannot tell whether the targeted capability changed, transferred to work, or contributed to the operational outcome that justified the intervention.
This exercise runs locally using simple rules. Your entries stay on this page and clear when you reload. No AI service is called.
Plan what would count as evidence
Start with a decision your team needs to make. Map what learners will demonstrate, then name what could make the apparent effect misleading.
The filled example is fictional. This planner organizes your entries and checks your selected evidence types. It does not evaluate the quality of your measures or estimate causal effects.
Your proposed evidence chain
Each link is a hypothesis to examine; this is not a demonstrated causal chain.
- Activity:
- Capability:
- Near-term observation:
- Transfer:
- Operational outcome:
What this plan leaves unresolved
Comparison limits
Smallest next improvement
Preview the complete plan
Critical review of this working prototype
Strongest assumption: An explicit evidence chain makes unsupported impact claims easier to recognize before deployment.
Likely failure: A completed form or a selected performance label can create an appearance of rigorous measurement without an adequate task or comparison.
Test that could change the design: Compare original and revised measurement plans using independent reviewers, then check whether teams collect disconfirming evidence and narrow claims when needed.
Non-AI comparison: A paper logic model with a measurement reviewer and a rival-explanations checklist.
Evidence to collect: Measure changes in work samples, transfer checks, comparison logic, and claim restraint, not form completion.
Human control and access: User selections are self-descriptions; the tool does not validate a measure, estimate an effect, or establish attribution.
Changes made during review
- Separate participation, self-report, and direct performance evidence.
- Require a described transfer observation when one is selected.
- State limits for absent, before/after, and concurrent comparisons.
- Invalidate the export when inputs change; include the stated performance criterion in the download.
This is a design and implementation review, not an empirical validation of learning outcomes.
Take the brief into your own AI environment
Copy a self-contained review prompt, then open the AI workspace you already use. The prompt asks for a falsification test, a non-AI alternative, evidence that distinguishes activity from learning, and human-system risks.
Preview the prompt
Act as a critical learning-engineering reviewer. Review the prototype brief below as a thin-slice experiment, not as a finished product. Return: 1. The strongest design assumption. 2. The most important plausible failure mode. 3. The highest-value falsification test for the next cycle. 4. One credible non-AI alternative that could address the same capability gap. 5. The next smallest prototype worth building. 6. What should be measured to distinguish engagement from learning, transfer, and real performance. 7. Human-agency, equity, accessibility, privacy, or governance concerns that should change the design. Be concrete. Separate evidence-backed claims from hypotheses. Do not reward novelty for its own sake, and do not assume more AI is better. COLLECTION: LENS Canonical TITLE: Evidence-to-Impact Mapper CAPABILITY GAP: Organizations often know who completed training but cannot tell whether the targeted capability changed, transferred to work, or contributed to the operational outcome that justified the intervention. FIRST PROTOTYPE: Build an AI-assisted measurement planner that begins with a system outcome, maps the human capability hypothesized to influence it, proposes observable proximal and transfer measures, and surfaces rival explanations before deployment. REPRESENTATIVE USE CASE: A service team wants to reduce repeat customer escalations. Instead of treating course completion as evidence, the tool links a specific diagnostic capability to scored work samples, field observations, and downstream escalation patterns. LEARNING / HUMAN-SYSTEM FRAME: Measurement is part of the design, not a post hoc dashboard. The prototype should make causal claims smaller and clearer, distinguish capability evidence from operational outcomes, and preserve uncertainty when attribution is weak. QUESTIONS ALREADY IDENTIFIED: 1. What evidence would convince us the capability changed? 2. What outcome is close enough to the intervention to attribute reasonably? 3. Which rival explanations would make the apparent effect disappear? FIRST-CYCLE FRAMING: Understand: Specify the operational decision the evidence must support and the claim being tested. Map: Build a causal map linking intervention, learner behavior, capability, context, and system outcome; name rival explanations and measurement threats. Instrument: Predefine work samples, transfer measures, operational indicators, comparison logic, subgroup checks, and stopping rules for claims the evidence cannot support.
Open the underlying design brief and eight-step path
A first prototype
Build an AI-assisted measurement planner that begins with a system outcome, maps the human capability hypothesized to influence it, proposes observable proximal and transfer measures, and surfaces rival explanations before deployment.
Representative use case
A service team wants to reduce repeat customer escalations. Instead of treating course completion as evidence, the tool links a specific diagnostic capability to scored work samples, field observations, and downstream escalation patterns.
Questions worth carrying forward
- What evidence would convince us the capability changed?
- What outcome is close enough to the intervention to attribute reasonably?
- Which rival explanations would make the apparent effect disappear?
Evidence anchors
These sources motivate the design; they do not validate this prototype.
- National Academies of Sciences, Engineering, and Medicine. How People Learn II: Learners, Contexts, and Cultures. (2018)
Broad learning-science anchor for learner/context/system modeling and transfer.
- Hattie, J., & Timperley, H. The Power of Feedback. Review of Educational Research. (2007)
Feedback design and evidence-producing practice.
- Barnett, S. M., & Ceci, S. J. When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin. (2002)
Transfer as a design and evaluation target.
- National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0). (2023)
Risk, governance, measurement, and accountable deployment.
- U.S. Department of Education, Office of Educational Technology. Artificial Intelligence and the Future of Teaching and Learning: Insights and Recommendations. (2023)
Human-centered AI in education; governance and instructional use.
One possible first pass
- 01 Understand
Specify the operational decision the evidence must support and the claim being tested.
- 02 Map
Build a causal map linking intervention, learner behavior, capability, context, and system outcome; name rival explanations and measurement threats.
- 03 Design
Compare an AI intervention with simpler non-AI options. Specify what the human decides, what the AI may suggest, and which tradeoffs are acceptable.
- 04 Build
Create the smallest usable prototype around one representative task, with expert-curated content, visible uncertainty, and an easy human override.
- 05 Instrument
Predefine work samples, transfer measures, operational indicators, comparison logic, subgroup checks, and stopping rules for claims the evidence cannot support.
- 06 Deploy
Pilot with a small, representative group in a near-real setting; preserve a baseline or comparison condition and document implementation conditions.
- 07 Evaluate
Look for capability growth and transfer, not just satisfaction or activity. Inspect subgroup patterns, human workload, errors, and unintended adaptations.
- 08 Refine
Let the evidence change the problem model. Default back to UNDERSTAND if the assumed gap was wrong; otherwise revisit the earliest step invalidated by the evidence.