Show Your Work
LENS: Coupling Learner and System Performance for Capability and Human Flourishing
William Gray-Roncal and James Diamond
Working paper for community discussion. Learning Engineering for Next-Generation Systems (LENS) offers a framework for defining, developing, and evaluating capability. This paper presents a synthesis and a design proposal for the learning-engineering community.
Abstract
Learning engineering seeks to help people accomplish meaningful work in the settings where that work matters. Learning Engineering for Next-Generation Systems (LENS) frames capability as the capacity of a person, team, or human-system arrangement to accomplish valued work under stated conditions, with usable opportunities to act. Capability depends on the coupling of learner performance with system performance: knowledge and judgment must meet accessible tools, reliable information, adequate resources, authority, and workable coordination. Human flourishing supplies the purpose for this coupling, directing attention to agency, well-being, participation, and the distribution of benefits and burdens. We connect this position to assessment research and work-system design, distinguish output quality, capability, learning, and intervention impact, and propose a practical evidence record for evaluating their relationships. An iterative process and developing examples show how learning-engineering teams can examine both people and systems. The contribution is a framework for community scrutiny and a research agenda; the examples do not establish learning gains or improved flourishing. To show the work is to make the capability claim, its conditions, its evidence, and its consequences open to examination.
1. LENS begins with capability
A learner can know what to do and still be unable to do it. A system can deliver its intended output and still leave people unable to act. Learning engineering has to examine the relationship between these two forms of performance. Otherwise, it can produce a well-taught course inside a work system that prevents competent action, or a technically successful system whose users cannot understand, direct, or recover from its behavior.
Consider a hypothetical effort to reduce missed alerts. A learner may recognize an alert and explain the appropriate response. The operational system must also provide a timely, interpretable signal, route it to someone with authority, and make a response feasible. Successful alert handling depends on that coupling. If a notification arrives after the decision, if the recipient lacks permission, or if the workload makes action impossible, more instruction may leave the capability gap unchanged. If the system works but the learner cannot recognize a dangerous exception, redesign alone may also be insufficient.
Learning Engineering for Next-Generation Systems (LENS) asks the team to begin with the valued work and the conditions that make it possible. What should people be able to accomplish? What must the system provide? How will the two perform together when conditions vary? What benefit is worth pursuing, and who bears the burden of achieving it? The intervention may involve instruction, practice, a job aid, a redesigned interface, a different handoff, staffing, authority, or a combination of these.
The intended end is human flourishing. A capability is valuable because of what it enables people to do and become, and the lives and relationships it helps them sustain. Completing more tasks or shipping a product can contribute to that purpose, but neither establishes it. The community must be able to question whether the work expands meaningful opportunity, preserves agency, improves well-being, and serves the people affected.
This position requires functional evidence. A grade, certificate, portfolio, model benchmark, or finished product may contribute useful information, but each needs an interpretation. Assessment research treats the connection between observed performance and a proposed use as an argument whose assumptions must be examined. Evidence-centered design links the claim we want to make, the observations that would support it, and the tasks that could elicit those observations. [1, 2] LENS extends the practical question to the learner-system arrangement: what worked, for whom, with what support, and under which conditions?
2. Defining capability and its evidence
We propose capability in context as the capacity of a person, team, or human-system arrangement to accomplish specified, valued work under stated conditions, with usable opportunities to act. This is a working definition for design and evaluation. It is not a universal proficiency scale or a claim that one successful performance establishes enduring capacity. The unit of analysis must be named: a learner, a team, a technical component, or the coupled arrangement.
Learner performance concerns what a person demonstrates: interpreting information, making a judgment, taking an action, explaining a choice, or recovering from an error. System performance concerns whether the surrounding arrangement supplies what the work requires: usable information, reliable tools, access, staffing, support, authority, and coordination. A technical benchmark covers only the component it tests. A learner score covers only the performance its task and conditions elicit.
Coupling concerns the relationship between them. Can the learner use the information the system supplies? Can a decision become an authorized action? Can an error be noticed and corrected? Do handoffs preserve what the next person needs? Capability is not established by adding a human score to a system score, or by assuming that individually strong components will coordinate successfully. The relationship needs its own evidence from representative work.
Four claims organize that evidence. They concern different objects and require different observations.
| Claim | What must be established | What one good artifact cannot establish |
|---|---|---|
| Output quality | The product meets stated requirements in the tested conditions. | Who supplied the expertise or whether performance will repeat. |
| Capability in context | A person or team can perform representative work with specified support and constraints. | Performance with materially different tools, tasks, or conditions. |
| Learning | A relatively durable change attributable to experience, examined through later performance and appropriate comparisons. | Retention, transfer, or improvement from a single final submission. |
| Intervention impact | The intervention contributed to a valued outcome relative to a credible alternative. | Causation from an observed success or a before-and-after difference alone. |
These distinctions are practical. A team may deliver a useful dashboard while relying on one expert for all substantive interpretation. A learner may understand a problem but fail because the interface hides essential information. A tool may raise immediate performance without developing the capability needed when the tool is unavailable. Each situation deserves a different diagnosis.
Learning and immediate performance can diverge. The literature reviewed by Soderstrom and Bjork shows why smooth practice performance is an unreliable stand-in for durable learning. [3] For our purposes, this means that a successful demonstration supports a bounded performance claim. A learning claim also needs evidence of change, persistence, and the relevant form of transfer. No single waiting period or transfer task is adequate for every domain; the choice must follow the intended use.
The capability claim must therefore state both the demonstrated performance and the conditions that made it possible. A claim about a learner using a reliable support system differs from a claim about independent performance or recovery when support fails. None should silently stand in for the others.
3. Coupling learner performance with system performance
Performance depends on more than what a person knows. Tools, staffing, time, information, incentives, authority, and coordination can enable competent work or prevent it. A training intervention is justified when the evidence supports a learning problem that training can address. It is not the default response to every disappointing result.
This systems perspective has substantial precedent. Carayon and colleagues’ Systems Engineering Initiative for Patient Safety (SEIPS) model relates work-system components, processes, and outcomes, explicitly attending to their interactions. [4] LENS brings this systems perspective into the same design and assessment conversation as claims about learning and professional competence. The relationship between people, technology, and organization has established precedents; our proposal is a practical synthesis for learning-engineering work.
Consider a proposed training course intended to reduce missed alerts. Before designing the course, examine whether alerts are accurate, interpretable, timely, routed to someone with authority, and feasible to act on. Training may help recognition or response. It cannot create missing staff, repair an unreliable data feed, or give someone authority that the organization has withheld. These are competing explanations to investigate, not conclusions to infer from the outcome alone.
We call this gap attribution: constructing and testing explanations for the difference between intended and observed performance. Useful evidence could include observations, interviews, system logs, work samples, and targeted changes to the workflow. The evaluator should record which explanations remain plausible. Finding a human contribution to failure does not remove a system contribution, and finding a system contribution does not establish that every individual skill is adequate.
The team should examine three levels together. At the learner level, it can observe interpretation, judgment, action, and recovery. At the system level, it can inspect delivery, accessibility, reliability, workload, and authority. At the coupling level, it can observe the complete task: whether the right information reaches someone who can understand it, act on it, coordinate with others, and correct a failure. A system log showing delivery and a learner test showing recognition do not, by themselves, establish successful action in the intended setting.
Coupling also creates a learning opportunity. The system can make explanations, feedback, practice, and correction available while preserving meaningful decisions for the learner. It can also conceal important work or reward dependence on assistance. The design question is how support contributes to current performance while developing the judgment and recovery capability needed for future conditions. Which support should remain, which should fade, and which needs a fallback depends on the intended work.
The practical consequence is a wider set of permissible responses. Develop a skill, redesign a tool, change a handoff, provide a job aid, narrow the task, acquire better evidence, or stop. A justified decision not to deploy can be an example of capable work. An assessment that rewards only shipping a product would miss it.
4. Tasks connect competencies to performance
Knowledge, Skills, Abilities, and Tasks (KSAT) is one vocabulary for distinguishing enabling capacities from the work itself. The vocabulary is not universal. The National Institute of Standards and Technology (NIST) Workforce Framework for Cybersecurity, from the National Initiative for Cybersecurity Education (NICE), uses Task, Knowledge, and Skill statements in its 2020 revision. [5] For LENS, the practical requirement is to connect a competency claim to observable work and its conditions, regardless of the taxonomy used.
A task analysis should identify the work to be accomplished, its important variations, acceptable performance, and the resources ordinarily available. It can then identify the knowledge and skills likely to support that work. This gives curriculum and assessment a common reference. It does not justify the inference that completing a task proves possession of every hypothesized enabler.
There are many paths to the same output. A checklist can substitute for recall. A colleague can catch a mistake. AI can write code that its user could not write independently. A learner can reproduce a familiar solution without understanding when it applies. The resulting work may still be useful, but the claim about the person must be narrower than the claim about the product.
For example, someone who creates a functioning browser game with AI may demonstrate the ability to describe a desired experience, notice a defect, and guide a revision. That is meaningful work. It does not establish independent programming skill. Conversely, removing all support would be a poor assessment if the intended role is to direct and verify tool-assisted production. The assessment conditions should match the capability being claimed.
We therefore propose recording support explicitly: what the person did, what collaborators or tools supplied, what was checked, and what still depends on external help. Targeted explanation, diagnosis, and modification tasks can test uncertain inferences. These probes should be short and relevant; the aim is to establish the claim, not to require every person to reproduce every layer of the system.
5. Human flourishing sets the purpose
Human flourishing provides a reason to develop capability and a basis for questioning which capability deserves priority. We use it here as a normative commitment to agency, well-being, meaningful participation, and the opportunity to pursue valued purposes. These dimensions require interpretation with the people affected. They cannot be inferred from a productivity score or collapsed into one universal index.
The United Nations Development Programme (UNDP) frames human development around expanding people’s choices and distinguishes developing capacities from having opportunities to use them. [15] This perspective supports our concern with both learning and the conditions for action. LENS does not claim to reproduce a complete theory of human development. It asks a learning-engineering team to make the intended human benefit explicit and to examine whether its intervention supports it.
A measurable outcome is not necessarily the right outcome. Completion, satisfaction, speed, eye contact, and a polished product can each be useful observations while remaining poor substitutes for the capability of interest. Choosing a criterion is a substantive decision about whose purposes the system serves and what tradeoffs are acceptable.
Begin with the work and the people affected. Define the intended benefit, the conditions in which it matters, unacceptable failure modes, and the parts of performance that cannot be reduced to a single total. A faster decision may be worse if it hides uncertainty or transfers work to someone else. A technically correct tool can fail if its intended users cannot access or understand it. The outcome should include what makes the work useful and responsible.
Make the flourishing claim concrete enough to challenge. In the alert example, a team might seek more timely care while preserving patient choice, workable staff demands, and access to correction. It should collect appropriate evidence about those aims rather than use faster acknowledgment as a substitute for all of them. In education, completing a supported task may be useful, while confidence to act, opportunities to participate, and later performance on new tasks remain separate questions. The people affected should help define the benefit and acceptable tradeoffs; the sponsor’s target is one input into that decision.
A system may improve its average result while increasing burdens on a smaller group, reducing discretion, or making access more fragile. Conversely, giving people more control may introduce time or coordination costs. A LENS evaluation should state those tensions, inspect their distribution, and identify who can revise the design. Human flourishing guides the choice of outcomes and the interpretation of consequences; it is not an automatic result of improved task performance.
Distal outcomes, such as patient benefit or sustained workplace performance, matter. They are also influenced by many factors beyond a particular learner or intervention. Proximal evidence, such as a diagnostic explanation or a correctly executed procedure, can help locate mechanisms and guide feedback. Neither should automatically displace the other. Trace the proposed relationship between them and state which links have actually been examined.
The assessment should also distinguish a construct from its chosen measure. Engagement, for example, is not identical to looking at a screen or speaking frequently. Our MicroGPT teaching experiment makes this distinction visible through synthetic training data and deliberately constructed labels. It invites learners to inspect what a model is being asked to predict. It is an illustration of measurement choices, not empirical evidence about any culture or group. [10]
Fairness belongs in criterion selection, task design, access, and interpretation. Ask whether the task requires irrelevant language fluency, expensive equipment, prior coaching, or behavior associated with a narrow cultural expectation. Provide access and accommodations that preserve the intended construct. Examine subgroup results with attention to sample size and uncertainty, and provide a way to challenge the interpretation. Demonstration can support fairer decisions, but it is not automatically fair simply because the work is visible.
6. What counts as enough evidence?
Decision-grade evidence is evidence sufficient for a specified decision, given its consequences, uncertainty, reversibility, and available alternatives. It is not a universal score or a claim that maximal rigor is always affordable. A low-stakes prototype can proceed on limited evidence if the next step is bounded and informative. A consequential deployment requires a stronger account of reliability, failure, and recovery.
Evidence quality cannot be read from the source label alone. A randomized study, an investigation, an interview, and a software test answer different questions. A randomized study of the wrong outcome may contribute little to the decision at hand. A carefully documented failure can be enough to stop a release even when it cannot estimate the average effect of the intervention. Record study design, relevance, independence, limitations, and uncertainty alongside provenance.
Two bounded examples show why the type and scope of the claim matter. They are illustrations, not a systematic review or a representative sample of successes and failures.
A deployed model needs external evaluation. Wong and colleagues evaluated the Epic Sepsis Model using 38,455 hospitalizations at Michigan Medicine. The hospitalization-level area under the receiver operating characteristic curve was 0.63; at a threshold of 6 or higher, the model failed to identify 67 percent of sepsis cases in that cohort. [6] These findings concern the model, setting, time period, and evaluation described in the study. They do not establish how every later version performs. They show why adoption and vendor-reported performance cannot replace evaluation in the setting where a decision will be made.
A learning intervention is a package implemented over time. The randomized evaluation of Cognitive Tutor Algebra I found no first-year effect and evidence of second-year benefit, statistically significant in high schools but not middle schools. [7] The intervention combined curriculum, software, and implementation. The study does not isolate the causal contribution of one learner-modeling algorithm. Nor does the year-two pattern, by itself, tell us which implementation mechanism produced it. A justified account keeps the package, timing, population, and outcome attached to the result.
For an individual assessment, sufficiency usually requires more than one polished work sample. Programmatic assessment provides a precedent for combining observations and judgments into a defensible decision. [8] Our proposed application is to sample meaningful variation across tasks, occasions, and support conditions, increasing the evidence burden with the breadth and consequences of the claim. Agreement among assessors helps, but agreement on a flawed criterion does not establish validity.
7. AI changes what we must observe
AI can contribute as a creative partner, a source of suggestions, a coding assistant, a tutor, or a provisional evaluator. In each role it changes the relationship between the visible artifact and the human contribution. The response should be to clarify the claim and inspect the relevant work, while preserving useful assistance where it belongs in the task.
Three conditions are often worth distinguishing. Supported performance asks what a person can accomplish with the tools available in the intended setting. Independent performance asks what that person can do without specified assistance when the role requires it. Recovery performance asks what happens when the support is wrong, missing, or misleading. These are proposed assessment conditions, not a requirement to remove all tools from every evaluation.
Process evidence can help explain an output. A sequence of revisions may show that a learner identified an error, checked a source, or rejected an inappropriate suggestion. It is still an incomplete record. An absent explanation in a transcript is not proof of absent understanding; important work may occur elsewhere. A longer transcript is not necessarily better evidence, and requesting private internal reasoning is neither necessary nor an appropriate substitute for observable explanations and decisions.
AI-assisted scoring adds another inference that needs evaluation. A plausible explanation for a score does not establish that the score is correct. We propose testing scorers against independently adjudicated examples, preserving disagreement, auditing performance on unfamiliar tasks and relevant groups, and rechecking after changes to the model, prompt, or rubric. Sources quoted as evidence must actually support the criterion. The human decision-maker needs a usable correction and appeal process, not merely the theoretical ability to override a score.
Calibrating that human judgment is part of the work. Raters should score common examples, explain consequential disagreements, and revisit the criterion when the disagreement exposes ambiguity. An expert reference set is a documented judgment, not an infallible answer key. Multiple AI ratings are not automatically independent evidence. These practices make uncertainty inspectable; their effects on decision quality still require evaluation.
8. Examples of learner-system coupling
The following projects make the proposal concrete. They are author-associated artifacts and design examples, not independent validations of this paper. No learner-effect estimate is claimed for them here. Their role is to expose design choices and generate testable questions about learner performance, system performance, and the relationship between them.
Rainbow Bug Dash: creative judgment with implementation support
Rainbow Bug Dash was co-created by Julian, age 6, with AI as a creative partner. The browser game asks the player to collect rainbows and avoid bugs. The companion Vibe Coding job aid, by James Diamond and Will Gray-Roncal, describes the cycle of expressing an idea, reacting to a working result, specifying changes, and testing the revision. [9]
The example broadens what can count as a contribution. Specifying that a character should begin in a safe place is a design judgment even when the person making it cannot implement the collision logic. The functioning game is evidence of a resulting artifact. The job aid documents aspects of an iterative process. Neither is a controlled measure of Julian’s learning, proof of independent programming skill, or evidence that the approach works equally well for all children. A learning assessment would require an age-appropriate new task and observations of what the child can explain, choose, and revise over time.
MicroGPT: a system that makes its behavior inspectable
The MicroGPT companion adapts Andrej Karpathy’s small transformer implementation for a teaching interface. Its sequence moves from letters to haiku text to a synthetic bias example; users can inspect training data and compare character, word, and line tokenization. Live training and stored checkpoints make changes in model behavior available for inspection. [10]
The assessment opportunity is to ask learners to predict, test, and explain a change. What changed when the tokens changed? What did the training labels reward? Which result would justify revising the data rather than training longer? Successful interaction with the demonstration does not establish understanding. A useful follow-up presents different data and asks the learner to identify the same kind of problem without being shown the answer.
Calibrated Judgment: evidence and correction at the assessment interface
Calibrated Judgment’s design compares evidence in a finished essay with evidence available in the accompanying AI dialogue, linking criterion judgments to quoted material and routing selected decisions to an instructor. [11] The aim is to make the basis of assessment inspectable and allow corrections to inform later calibration.
A discrepancy between the two assessments is a reason to ask a better question. It is not a validated measure of over-reliance, deception, or lack of understanding. Missing transcript context, task differences, scorer error, and legitimately different evidence can also produce a gap. The next step should be a targeted clarification or performance probe before making a consequential claim about the learner.
ExpertTrace: supported practice and unfamiliar tasks
ExpertTrace’s design uses an operational knowledge corpus to support scenario-based practice and probe explanations at different levels of expertise. [12] It illustrates how an assessment can examine symptom interpretation, evidence use, uncertainty, and corrective action rather than reward recall of one finished answer.
The proposed evidence of learning would come from performance on unfamiliar scenarios, including cases in which a familiar cue is misleading or the support cannot be relied on. Repeating a practiced scenario can be useful instruction while remaining weak evidence of transfer. Working software and coherent simulated interactions establish implementation progress; learner gains require a study with learners and an appropriate comparison.
9. The LENS process and community competencies
The LENS process treats capability development as an iterative investigation. The team begins with a seed idea or concern, defines the problem with the people affected, maps capabilities and operating conditions, compares interventions, and revises its decisions as evidence becomes available. A product is one possible result. A changed workflow, better-supported practice, a narrower scope, or a justified decision to stop may also be an appropriate result.
The eight activities keep learner and system requirements connected:
| Activity | Question for the coupled arrangement |
|---|---|
| Understand | What valued work is difficult, for whom, and what human benefit would closing the gap serve? |
| Map | What must people and systems each be able to do, and which conditions or dependencies enable their work together? |
| Design | Which intervention addresses a plausible cause, within available authority, with acceptable benefits and burdens? |
| Build | Can the candidate, its support, and its recovery paths operate during a representative task? |
| Instrument | What evidence will distinguish learner performance, system performance, their coupling, and the intended benefit? |
| Deploy | Under what bounds, staffing, access, and stop conditions can the arrangement be used responsibly? |
| Evaluate | What changed, for whom, compared with what, and which explanations and consequences remain uncertain? |
| Refine | Which assumption or requirement needs revision, and what evidence would justify the next step? |
These are revisitable activities. New evidence may change the problem definition, expose an authority constraint, or show that a successful build does not support the intended work. Refine normally returns to Understand; a direct return to another activity needs a stated reason. The Capability Pipeline presents these choices through five worked cases and a fictional adventure, with teaching points to support deliberation. Its simulated outcomes are authored assumptions rather than empirical estimates. The companion casebook supplies additional contexts for discussing the work. [13, 14]
LENS organizes the relevant practice around five competency domains: Systems Analysis, Iterative Development, Human-System Collaboration, Test and Evaluation, and Sociotechnical Constraints. In this paper, these are organizing domains for the LENS framework, offered for community discussion. They do not constitute an adopted professional standard. A project can use them to name the expertise it needs and identify missing contributions.
A learning engineer can orchestrate that work by keeping the problem, intended benefit, evidence, and next decision connected. Domain practitioners, learning designers, technical specialists, evaluators, and affected participants contribute different expertise. One person need not perform every role, and naming a role does not demonstrate proficiency. A proficiency claim needs reviewed evidence of the relevant task under stated conditions, including judgment about when to seek help, revise, or stop.
This gives the community a concrete basis for discussing learning engineering as a profession, a team practice, and a process. The framework names work that can be shared across a team, documented in an iterative process, and used to examine an individual’s contribution. Competency definitions can develop through that examination rather than through a list detached from what teams actually do.
10. A practical evidence record
We propose a small evidence record for each consequential capability claim. It should be short enough to use and complete enough for another evaluator to question. The record is an author proposal developed from the assessment and systems principles above; it has not been validated as an assessment instrument.
| Element | What to record |
|---|---|
| Claim and use | The work, who or what is being assessed, and the decision the result will inform. |
| Conditions | Task variation, time, tools, AI, collaborators, accommodations, and relevant constraints. |
| Standard | Observable criteria, unacceptable errors, and how the performance level was established. |
| Learner evidence | Observed interpretation, judgment, action, explanation, transfer, and recovery, with provenance. |
| System evidence | Accessibility, information quality, delivery, reliability, resources, staffing, and authority relevant to the task. |
| Coupling evidence | Complete-task observations, handoffs, support use, exception handling, and correction paths. |
| Human benefit | The intended contribution to flourishing, who helped define it, and observed benefits, burdens, and missing outcomes. |
| Interpretation | How observations support the claim; alternatives, missing evidence, and uncertainty. |
| Decision and review | The permitted next step, responsible decision-maker, appeal route, and reassessment trigger. |
A record might support this bounded claim: a learner can use an approved assistant to create a small browser tool, verify its stated behavior, explain two consequential design choices, and repair a seeded defect. Evidence would include the tool, its tests, the explanations, and the repair. The claim would not extend to independent software engineering, security assurance, or durable learning without additional evidence.
The same record can support a decision to withhold judgment. If authorship is unclear, a key test is missing, or the task is not representative, report the limitation rather than manufacture a precise score. Where a safety-critical criterion applies, a high total should not compensate for failing that criterion unless the decision rule explicitly permits that tradeoff.
The evidence record should travel with the project through the LENS cycle. Decisions, objections, tests, revisions, and unresolved questions help explain how the arrangement changed. Preserve relevant artifacts and concise decision rationales so another team can inspect what was available at the time. These records support examination; their volume does not establish capability or improved outcomes.
11. How to test the proposal
A position paper should specify what would make its claims less credible. The next studies should compare the proposed assessment approach with a clearly described alternative, rather than compare a richly supported intervention with an unspecified absence of support. The appropriate design depends on the question; the following is a research agenda, not a report of completed work.
First, test the coupling claim. Where feasible, compare changes in learning support, changes in the operating system, and their combination on representative tasks. Observe learner performance, system performance, and complete-task performance separately, including exceptions and recovery. A design with both kinds of change does not by itself establish their interaction; suitable comparisons and evidence about the mechanism are needed. The framework needs revision if the proposed coupling measures add little beyond the separate measures or fail to explain relevant variation in performance.
Second, test whether the evidence record improves the accuracy and usefulness of decisions. Independent raters could judge common work samples with and without access to the structured record. Outcomes should include justified decisions, consequential errors, unresolved cases, agreement, and time required. More agreement would not count as success if it simply reflected shared error. A separate adjudication procedure and tasks outside the development set are needed.
Third, distinguish supported productivity from learning. Assess relevant baseline performance, track what support was actually used, and examine later performance on new tasks. Use both supported and independent conditions where they match the intended claims. Include a recovery task when recognizing erroneous assistance is part of the work. Specify the primary outcome and analysis before examining results, and report uncertainty, missingness, and participation differences.
Fourth, examine the intended contribution to human flourishing. Specify the relevant benefit with affected participants, then observe opportunities to act, meaningful control, well-being, and the distribution of burdens where those are part of the claim. Establish a credible comparison before claiming intervention impact. A gain in completion would not be sufficient if it depended on unacceptable loss of agency, exclusion, or workload. Longer-term and external effects may require a different study from the immediate performance evaluation.
Fifth, examine whether the approach works for the people who would use it. Observe accessibility barriers, unequal preparation, rater disagreement, and differences in errors or burden across relevant groups. Invite affected learners and practitioners to review the construct and its operationalization. Small groups and sparse failures limit what can be concluded; an absence of detected disparity is not proof of fairness.
Sixth, test the operational cost. Record preparation time, scoring and adjudication effort, tool costs, appeals, privacy requirements, and maintenance after changes. A method that improves a narrow research task but cannot be sustained in the intended setting has not yet demonstrated the capability that adoption requires.
Finally, evaluate outside the development team. External assessors, new sites, unfamiliar tasks, and independent replication can reveal assumptions that internal testing misses. The proposal would need revision if added process evidence fails to improve decisions, imposes disproportionate burden, worsens disparities, or fails to predict relevant later performance. A useful outcome may be a narrower claim about where the method belongs.
12. Community commitments and responsibilities
Showing work creates records about people. Collect the evidence needed for the stated decision, explain its use, limit access and retention, and avoid treating continuous surveillance as the price of a credible assessment. Prefer selected work samples and targeted probes when they can answer the question. The ability to collect a behavioral trace does not establish a reason to use it.
Learning-engineering competence also includes the authority to act responsibly within a system. Who may change a criterion, deploy a tool, access a learner record, or overrule an automated recommendation? Those roles should be explicit. A responsible process preserves the possibility of escalation, correction, and stopping when the evidence is inadequate.
This paper is a proposed synthesis. It does not establish a new psychometric theory, a universally valid rubric, a causal estimate of learning gains, or an adopted credentialing standard. It does not represent a formal position of the Institute of Electrical and Electronics Engineers (IEEE) or its International Consortium for Innovation and Collaboration in Learning Engineering (ICICLE). Contributions to that community require review and agreement through its own processes.
We offer LENS as a basis for community work on the meaning and evidence of capability. The invitation is to examine the definition, test the coupling of learner and system performance, make competency claims concrete, and question whether the resulting capability serves human flourishing. Different settings will require different tasks, standards, supports, and measures. The common commitment is to make those choices inspectable.
Show representative work. State what learners and systems each contribute, how they perform together, and the conditions that permit action. If the claim is learning, examine change and persistence. If it is operational capability, examine the complete work arrangement. If it is impact, justify the causal inference. If it is human flourishing, make the valued benefit and its distribution explicit. Learning engineering becomes accountable to the people it serves when each of these claims is open to examination and correction.
References and companion artifacts
-
Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1), 1-73. doi:10.1111/jedm.12000.
-
Mislevy, R. J., Almond, R. G., and Lukas, J. F. (2003). A brief introduction to evidence-centered design. ETS Research Report RR-03-16. doi:10.1002/j.2333-8504.2003.tb01908.x.
-
Soderstrom, N. C., and Bjork, R. A. (2015). Learning versus performance: An integrative review. Perspectives on Psychological Science, 10(2), 176-199. doi:10.1177/1745691615569000.
-
Carayon, P., Schoofs Hundt, A., Karsh, B.-T., Gurses, A. P., Alvarado, C. J., Smith, M., and Flatley Brennan, P. (2006). Work system design for patient safety: The SEIPS model. Quality and Safety in Health Care, 15(Suppl. 1), i50-i58. doi:10.1136/qshc.2005.015842.
-
Petersen, R., Santos, D., Wetzel, K., Smith, M., and Witte, G. (2020). Workforce Framework for Cybersecurity (NICE Framework). NIST SP 800-181 Rev. 1. doi:10.6028/NIST.SP.800-181r1. Cited for its Task, Knowledge, and Skill structure, not as a learning-engineering standard.
-
Wong, A., et al. (2021). External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Internal Medicine, 181(8), 1065-1070. doi:10.1001/jamainternmed.2021.2626.
-
Pane, J. F., Griffin, B. A., McCaffrey, D. F., and Karam, R. (2014). Effectiveness of Cognitive Tutor Algebra I at scale. Educational Evaluation and Policy Analysis, 36(2), 127-144. doi:10.3102/0162373713507480.
-
Van der Vleuten, C. P. M., and Schuwirth, L. W. T. (2005). Assessing professional competence: From methods to programmes. Medical Education, 39(3), 309-317. doi:10.1111/j.1365-2929.2005.02094.x.
-
Rainbow Bug Dash, co-created by Julian, age 6, with AI as a creative partner. Play the game. Diamond, J., and Gray-Roncal, W. Vibe Coding: Building functional tools without writing code. Companion job aid, Word document. User-supplied artifact and process illustration; no learning-effect study is asserted.
-
MicroGPT for AI leaders. Teaching companion, adapted from Andrej Karpathy’s microgpt. Author-associated teaching prototype with synthetic examples.
-
Calibrated Judgment. Project. Author-associated assessment prototype; cited for the design described here, not established effectiveness.
-
ExpertTrace. Project. Author-associated scenario-practice prototype; cited for the design described here, not established effectiveness.
-
The Capability Pipeline. Interactive explorer. Author-associated teaching simulation, not an empirical outcome model.
-
Gray-Roncal, W., and Diamond, J. (2026). Capability Matters: A Casebook. Companion draft. A source of case framing; empirical claims in this paper cite their primary studies directly.
-
United Nations Development Programme. (1990). Human Development Report 1990: Concept and Measurement of Human Development. Report and overview. Cited for the distinction between developing human capacities and opportunities to use them, not as validation of the LENS framework.
Acronym glossary
AI: artificial intelligence. ICICLE: International Consortium for Innovation and Collaboration in Learning Engineering. IEEE: Institute of Electrical and Electronics Engineers. KSAT: Knowledge, Skills, Abilities, and Tasks. LENS: Learning Engineering for Next-Generation Systems. NICE: National Initiative for Cybersecurity Education. NIST: National Institute of Standards and Technology. SEIPS: Systems Engineering Initiative for Patient Safety. UNDP: United Nations Development Programme.
Acknowledgment and community discussion
We acknowledge the learning-engineering community, including IEEE ICICLE and the Learning Engineering Body of Knowledge, as important settings for the ongoing conversation about professional practice. This working paper invites discussion of capability, learner-system coupling, and human flourishing; it does not speak for those groups. Prepared with AI assistance. Offered for author review and community discussion.