Problems → Prototypes · AILE
An AI Debate Coach for Evidence, Reasoning, and Voice
Students often need far more low-stakes rehearsal in argumentation and public speaking than a teacher or coach can provide. Generic generative AI can produce polished arguments for learners, however, which risks replacing the very reasoning and voice the practice is meant to develop.
This exercise runs locally using simple rules. Your entries stay on this page and clear when you reload. No AI service is called.
Schools should require community service for graduation.
Make the strongest case you can. The coach will not draft your speech; it will help you inspect the reasoning you already supplied.
Inspect what this prototype is actually doing
No language model is running in this thin slice. The local rules look for whether the learner supplied all three argument parts, whether the evidence contains source-like specificity, and whether the reasoning explicitly connects evidence to the claim. A production version would replace these heuristics with a bounded model and expert-validated rubric, while preserving the same learner-first sequence.
Critical review of this working prototype
Strongest assumption: A challenge after an independent argument helps learners revise their own reasoning.
Likely failure: Keyword cues can reward formulaic wording, and an edited input can corrupt the comparison with the original attempt.
Test that could change the design: Compare independent arguments on a new motion after practice with this coach versus a paper claim-evidence-reasoning checklist; use blinded human ratings.
Non-AI comparison: A peer or teacher uses the same claim-evidence-reasoning checklist and asks one skeptical question.
Evidence to collect: Inspect original arguments, revisions, source use, live transfer, confidence, and dependence on suggested wording. Word counts are not learning scores.
Human control and access: Surface words are weak proxies for argument quality; learners need readable feedback and control over their own voice.
Changes made during review
- Compare with a captured submitted attempt rather than mutable form values.
- Invalidate feedback, revision, and comparison whenever any initial input changes; require resubmission.
- Clear old comparison results when revising or resubmitting.
- Label keyword matches as cues rather than demonstrated reasoning quality.
- Keep feedback hidden until the learner submits an attempt.
This is a design and implementation review, not an empirical validation of learning outcomes.
Take the brief into your own AI environment
Copy a self-contained review prompt, then open the AI workspace you already use. The prompt asks for a falsification test, a non-AI alternative, evidence that distinguishes activity from learning, and human-system risks.
Preview the prompt
Act as a critical learning-engineering reviewer. Review the prototype brief below as a thin-slice experiment, not as a finished product. Return: 1. The strongest design assumption. 2. The most important plausible failure mode. 3. The highest-value falsification test for the next cycle. 4. One credible non-AI alternative that could address the same capability gap. 5. The next smallest prototype worth building. 6. What should be measured to distinguish engagement from learning, transfer, and real performance. 7. Human-agency, equity, accessibility, privacy, or governance concerns that should change the design. Be concrete. Separate evidence-backed claims from hypotheses. Do not reward novelty for its own sake, and do not assume more AI is better. COLLECTION: AILE TITLE: An AI Debate Coach for Evidence, Reasoning, and Voice CAPABILITY GAP: Students often need far more low-stakes rehearsal in argumentation and public speaking than a teacher or coach can provide. Generic generative AI can produce polished arguments for learners, however, which risks replacing the very reasoning and voice the practice is meant to develop. FIRST PROTOTYPE: Create an AI rehearsal partner that listens to or reads a learner's argument, identifies the claim-evidence-reasoning structure, poses a counterargument, and gives process-focused feedback. It should never write the final speech by default and should make uncertainty and source quality visible. REPRESENTATIVE USE CASE: A learner preparing a school debate gives a two-minute opening statement. The coach asks one skeptical follow-up, highlights an unsupported inference, and invites a revised response before showing any model language. LEARNING / HUMAN-SYSTEM FRAME: The core mechanism is deliberate, repeated practice with specific feedback and self-explanation. Human judgment remains central: a teacher or coach sets norms, evaluates rhetoric and ethics, and helps learners interpret feedback rather than treating an AI score as authoritative. QUESTIONS ALREADY IDENTIFIED: 1. Does repeated AI rehearsal improve live performance with a human audience? 2. Which feedback prompts strengthen reasoning without homogenizing student voice? 3. How should the system handle controversial topics or biased source material? FIRST-CYCLE FRAMING: Understand: Validate that the real gap is the ability to construct, defend, and revise an evidence-based argument in real time, not simply low engagement, low tool use, or a workflow inconvenience. Map: Map the system: learner, coach or teacher, peers/audience, source materials, AI rehearsal partner, and debate norms. Identify where the capability currently succeeds, breaks down, or is masked by other constraints. Instrument: Predefine evidence: argument quality, response to counterarguments, source use, oral confidence, transfer to live debate, and reliance on AI wording. A key disconfirming signal is: fluency improves but original reasoning, source judgment, or authentic voice declines.
Open the underlying design brief and eight-step path
A first prototype
Create an AI rehearsal partner that listens to or reads a learner's argument, identifies the claim-evidence-reasoning structure, poses a counterargument, and gives process-focused feedback. It should never write the final speech by default and should make uncertainty and source quality visible.
Representative use case
A learner preparing a school debate gives a two-minute opening statement. The coach asks one skeptical follow-up, highlights an unsupported inference, and invites a revised response before showing any model language.
Questions worth carrying forward
- Does repeated AI rehearsal improve live performance with a human audience?
- Which feedback prompts strengthen reasoning without homogenizing student voice?
- How should the system handle controversial topics or biased source material?
Evidence anchors
These sources motivate the design; they do not validate this prototype.
- Hattie, J., & Timperley, H. The Power of Feedback. Review of Educational Research. (2007)
Feedback design and evidence-producing practice.
- McGaghie, W. C., et al. Does simulation-based medical education with deliberate practice yield better results than traditional clinical education? Academic Medicine. (2011)
Simulation, deliberate practice, and repeated performance with feedback.
- Amershi, S., et al. Guidelines for Human-AI Interaction. CHI 2019. (2019)
Human-AI interaction design and calibrated reliance.
- U.S. Department of Education, Office of Educational Technology. Artificial Intelligence and the Future of Teaching and Learning: Insights and Recommendations. (2023)
Human-centered AI in education; governance and instructional use.
- UNESCO. AI Competency Framework for Students. (2024)
Student AI literacy, critical judgment, and co-creation.
One possible first pass
- 01 Understand
Validate that the real gap is the ability to construct, defend, and revise an evidence-based argument in real time, not simply low engagement, low tool use, or a workflow inconvenience.
- 02 Map
Map the system: learner, coach or teacher, peers/audience, source materials, AI rehearsal partner, and debate norms. Identify where the capability currently succeeds, breaks down, or is masked by other constraints.
- 03 Design
Compare an AI intervention with simpler non-AI options. Specify what the human decides, what the AI may suggest, and which tradeoffs are acceptable.
- 04 Build
Create the smallest usable prototype around one representative task, with expert-curated content, visible uncertainty, and an easy human override.
- 05 Instrument
Predefine evidence: argument quality, response to counterarguments, source use, oral confidence, transfer to live debate, and reliance on AI wording. A key disconfirming signal is: fluency improves but original reasoning, source judgment, or authentic voice declines.
- 06 Deploy
Pilot with a small, representative group in a near-real setting; preserve a baseline or comparison condition and document implementation conditions.
- 07 Evaluate
Look for capability growth and transfer, not just satisfaction or activity. Inspect subgroup patterns, human workload, errors, and unintended adaptations.
- 08 Refine
Let the evidence change the problem model. Default back to UNDERSTAND if the assumed gap was wrong; otherwise revisit the earliest step invalidated by the evidence.