AI leaders' perspectives · A guided companion

A small model.
Better questions.

What should understanding a language model change about the way you lead an AI project?

Start with Andrej Karpathy's microgpt. Then follow the decisions outward: from what the model learns, to how people use it, to what counts as evidence that the system works.

Work through three examples: letters, then words in example haikus (with character and whole-line comparisons), then bias in data and labels. Each step adds a question about what training can and cannot establish.

Example 1 · Letters

Learn a pattern, one letter at a time.

Begin with letters as tokens: the small units the model predicts. This model sees invented strings and practices predicting the next letter. It starts with random settings; training changes the probabilities it assigns to letters.

Try it: Inspect one training string, look at its letter tokens, then train 100 steps. Compare generated strings and the error on examples the model never practiced.
Inspect every training string and its tokens

Training samples with replacement. Each training example has the same chance of being drawn; test examples have zero chance. “Probe” marks the 16 training examples used for the displayed training loss. Test examples never update the weights.

Exact exampleUseChance per drawExplore

What the model actually receives

Select “Show tokens” beside any example above to inspect it here. The boundary token marks the beginning and end; the numbers identify tokens, not their meaning. Different tokenizations have different ID mappings. Scroll within the token lists to see every token.

 
Inspect the active model’s vocabulary

Loading this example's local training engine…

Training steps0
Training probe loss—
Held-out loss—

One step practices one complete invented string. Loss averages next-token error across its letters and the end token. The test contains 32 unseen strings from the same invented spelling rules. Lower is better on this task. Letter loss and word loss are different measurements; do not compare their numbers as a ranking of the two models.

Learning trace

4.000 steps100

Solid teal: training probeDashed rust: held-out

Exact measurements in this run
Measured states, including loaded checkpoints
StepTraining lossHeld-out loss

Generated letter sequences

Eight draws from the model. Plausible spelling is not proof of understanding. Output ends at the boundary token or after 12 letters. Sampling uses seed 2026 and never changes weights.

Start, save, or revisit a checkpoint

The ready-made states were trained with this engine and seed 42. Loading restores the actual weights, optimizer and random state. Each example runs independently; starting one does not reset the others. Download a state to keep it after leaving the page.

Ask of your project: what has actually improved? A lower training score shows a better fit to practiced examples. The held-out score checks a narrower claim: prediction on new strings from these rules.

Example 2 · Words in example haikus

One haiku. Three ways to define a token.

Train on the same example haikus using characters, whole words, or entire lines as tokens. The default is words, with a separate line-break token. Familiar-looking lines can emerge from word training because the same lines appear repeatedly in the examples.

The original teaching examples use an English 5/7/5 form, built from one-syllable words. There are 48 training haikus and 16 held-out combinations. Every source line is familiar; the test checks how the model predicts new combinations.

What changes? The word model can recombine words; the line model can only recombine the 12 complete source lines. The character model has to learn spelling too. Each mode uses its own vocabulary, context length and weights, so the raw loss numbers are not directly comparable.

Switching tokenization starts a fresh model. Download a checkpoint first to keep the current run. All three modes use the same 48 training poems and 16 test poems. For a quick comparison, load each mode's 100-step checkpoint. Equal steps mean equal sampled poems, but token counts, parameter counts and compute differ. Character training is slower.

Inspect every example haiku and its tokens

Training samples with replacement. Each training example has the same chance of being drawn; test examples have zero chance. “Probe” marks the 16 training examples used for the displayed training loss. Test examples never update the weights.

Exact exampleUseChance per drawExplore

What the model actually receives

Select “Show tokens” beside any example above to inspect it here. The boundary token marks the beginning and end; the numbers identify tokens, not their meaning. Different tokenizations have different ID mappings. Scroll within the token lists to see every token.

 
Inspect the active model’s vocabulary

Loading this example's local training engine…

Training steps0
Training probe loss—
Held-out loss—

One step practices one complete example haiku. Loss averages next-token error in the selected unit, including the end token. The test contains 16 new combinations of familiar lines; it does not test unfamiliar vocabulary or general poetic skill. Lower is better on this task. Letter loss and word loss are different measurements; do not compare their numbers as a ranking of the two models.

Learning trace

4.000 steps100

Solid teal: training probeDashed rust: held-out

Exact measurements in this run
Measured states, including loaded checkpoints
StepTraining lossHeld-out loss

Generated text · words model

Three draws from the model, without a poetry or syllable checker. Long lines scroll horizontally so actual line breaks stay visible. Character and word modes predict line breaks. Line mode samples complete fixed lines and displays a newline between them. Output can repeat training text, stop early or break the example form. Character mode can spell new words; word mode uses known words; line mode uses only known whole lines. Sampling uses seed 2026 and never changes weights.

Start, save, or revisit a checkpoint

The ready-made states were trained with this engine and seed 42. Loading restores the actual weights, optimizer and random state. Each example runs independently; starting one does not reset the others. Download a state to keep it after leaving the page.

Ask of your project: what does the model treat as a unit, and what can its vocabulary never express? A pleasing recombination of familiar words is a starting point for evaluation, not proof of broad language capability.

Example 3 · Bias and fairness

What does the data teach, and who is poorly served?

First change which names the letter model encounters. Then change what a separate classifier is taught to call engagement. Representation and labeling are different choices; inspect both.

3A · Representation bias · Real training on your device

Whose names get more practice?

Start from random weights, or load a measured checkpoint. Train on names used in a historical discrimination study, using the same letter-level architecture as Example 1. The model learns to generate strings from the chosen examples. It has 2,520 adjustable parameters, one transformer block, and three attention heads. Every live training update is computed in your browser.

First visit? Try “Train 100 steps,” compare the two loss values, then open “Explain this code.” The help buttons explain each term without assuming a programming background.
  1. What the model sees
  2. What the weights produce
  3. How the weights change
  4. What the evidence supports

What does the training data make familiar?

A model can become better at patterns it encounters more often. That can happen through deliberate selection or accidental gaps in collection. Here, change exposure while keeping the model design fixed. Then ask who could be poorly served if this pattern appeared in a real application.

Choosing a scenario starts a fresh model with seed 42. Each run learns different weights from its data. Download a checkpoint first if you want to keep the current run.

Loading dataset…

Inspect every training and test example

These are the exact lowercase strings supplied to the model. No hidden examples. Training draws with replacement: a word can appear repeatedly. Probabilities below are per draw, not observed counts. Test words have zero training probability. “Probe” marks the first eight training words in each set used for the training-loss display.

Active model's complete dataset
Exact textSetUseChance per training draw
Names, stereotypes and fairness

The names scenario uses historical U.S. study cues associated with perceived race. It makes a culturally situated stereotype inspectable without declaring it a universal fact about names or people. The model learns letter patterns; it does not predict anyone's identity or suitability for a job.

Inspecting the list also reveals who is absent. This English-alphabet model cannot represent accents or other scripts. Equal exposure within these two selected sets does not fix that exclusion. Continue to the engagement lab to see how a misleading label can survive balanced training.

The rest comes from set B. Applies when you choose New run. The inspector shows the active model’s mix.

Loading the local training engine…

Training steps 0
Training probe loss —
Held-out loss —

Loss is mean negative log probability per character, averaged equally across words. Lower is better for this prediction task. The fixed training probe contains 16 words; the held-out set contains 32 unseen words. Evaluation weights both sets equally. These small samples do not establish generalization to other tasks.

Look beyond the average

Set A held-out loss —

Set B held-out loss —

Lower loss means better spelling prediction on this tiny test set. A difference can reflect exposure, spelling difficulty or the split. Neither a gap nor equal scores establishes fair treatment of people.

Compare three models trained on the study names

Same names, test split, initial weights, optimizer and 600 updates; different chances of drawing each set. The equal-exposure run is a comparison point, not a “fair model” certification. Positive changes below mean worse prediction than that run. Each condition uses one seed, so this is a worked example, not a population estimate.

Measured 600-step checkpoints · all tests weight A and B equally
Training A:BA lossB lossA change vs 50:50B change vs 50:50Explore

Loading measured comparisons…

Learning trace

Measured training probe and held-out losses across saved steps. Exact values are also listed below. 4.000 steps1200

Solid teal: training probeDashed rust: held-out

Exact measurements in this run
Measurements saved at each checkpoint
StepTrainingHeld-outSet ASet B

Generated words

The same sampling seed is used at every checkpoint. Temperature changes the draw, not the learned weights. An empty sample is possible. Output stops at the end token or after 12 characters.

Keep or revisit a checkpoint

Pause before restoring or downloading. Restoring an earlier state branches this run from that point. Downloads contain model state, not personal information. Leaving the page clears session states.

Compare ready-made checkpoints without waiting

These are real saved training states from this exact engine, with seed 42. Each card identifies its dataset and mix. The displayed samples use temperature 0.8 and sampling seed 2026. Load one to inspect or continue training.

Loading checkpoints…

Exactly what is being learned?

The first example and this names model use the same next-character architecture. The first example’s invented families have no demographic meaning. Here, the sourced names have documented, historically situated associations; the inspector explains their provenance and our split. Set labels control sampling and auditing, but are never fed into the model.

Try it: compare a fresh balanced run with a fresh 90:10 run at the same number of steps. Inspect held-out loss separately for both name sets. The evaluation mix stays balanced even when training changes. Samples are illustrative draws, not a reliable estimate of the population distribution.

Our adaptation: Karpathy's original concept and reference architecture, independently implemented in JavaScript with a smaller width, an indexed differentiation tape, inspectable teaching datasets, and different initialization and optimizer settings. Full attribution and differences.

3B · Label and proxy bias · Invented learners, real model fitting

What if the model learns the wrong definition of engagement?

A model can faithfully learn a labeling rule that misrepresents the people it is meant to support. This experiment makes that failure visible. It fits a separate, simple logistic classifier, not the word-generating transformer above.

Everything here is synthetic. Contexts A and B are fictional, with the same underlying engagement prevalence. They are not countries, cultures, diagnoses, or estimates of real students. The numbers were chosen to expose a measurement failure. No camera, microphone, or eye-tracking data is collected.

Keep three things separate

1. The concept we care about
Engagement with a learning task. For this demonstration only, the generator assigns an unobserved binary state, 50% engaged in each context. Real engagement does not come with this answer key.
2. The signals we record
Invented camera-facing gaze, body stillness, and task-evidence scores, each between 0 and 1. Context B has a different outward display pattern at the same underlying engagement state.
3. What we call a positive training label
Either the assumed state, deliberately altered labels, or a rule that calls camera-facing and stillness “engaged.” These targets are not interchangeable.
Inspect every synthetic observation and its label

Assumed state and label use 1 for engaged and 0 for not engaged. Highlighted pairs disagree. Training weight is each row's contribution to a full-batch update, not a sampling probability. Test rows have zero training weight. Changing this view does not retrain the models.

Exact generated observations · rounded display
IDContextGazeStillnessTask evidenceAssumed stateRule labelTraining weight
1 · Explicit

Deliberately alter the labels

Relabel half of engaged examples in B as not engaged. Both contexts have equal training weight; all three signals are available.

2 · Inadvertent

Use a convenient proxy

Define engagement as average gaze and stillness above 0.55. Train with 90% weight on A. Nobody explicitly writes a rule about B; the proxy creates the mismatch.

3 · Incomplete repair

Balance representation

Give A and B equal weight but retain the same proxy labels and features. Does more representative data fix an invalid target?

4 · Different measurement

Change target and evidence

Use the assumed engagement state as the label and the noisy task-evidence signal as the feature. This is an idealized comparison, not a validated real-world engagement instrument.

Predict before running: which model will look strongest if you only score agreement with its own labeling rule? Which will make fewer mistakes against the assumed engagement state?

No classifier has been trained in this session.

Held-out audit: 400 new synthetic observations, including 100 engaged and 100 not-engaged examples per context. No test rows are used for training.
Training setupLabel agreement Accuracy against assumed stateEngaged A missed Engaged B missed
Run the comparison to calculate these values.

Decision threshold: model score ≥ 0.5. “Missed” means a false negative among the 100 engaged examples in that context. These are deterministic simulation counts, not population estimates or calibrated human-engagement probabilities.

Inspect false positives, target loss, and learned weights
How does this relate to culture, body language, and eye tracking?

Akechi and colleagues (2013) found both shared responses and differences in how Japanese and Finnish participants evaluated direct gaze. That study does not validate an engagement detector or establish the numerical differences used here. It supports asking whether a behavioral signal has the same meaning across contexts. Read the study.

Eye contact with a person, gaze at a screen, fixation on relevant material, and attention to a learning task are different observations and constructs. Posture and movement also need contextual interpretation. A student may look away while reasoning, attend to an interpreter, use assistive technology, or take notes elsewhere. These are plausible alternative explanations to investigate, not a diagnostic checklist.

Engagement research distinguishes behavioral, emotional, and cognitive dimensions. Fredricks, Blumenfeld, and Paris (2004) review the concept and its measurement. A review by Barrett and colleagues (2019) likewise cautions against inferring emotion directly from facial movements; emotion is not the same construct as engagement.

Our implication: do not turn the simulated A/B difference into a cultural stereotype or a rule for scoring individuals. Test the meaning and reliability of candidate measures in the intended setting, include learner perspectives and independent task evidence, and allow uncertainty or abstention. Group balance alone cannot establish construct validity.

What should a leader ask before adopting an engagement dashboard?
  • Who defined “engaged,” and how were the training labels obtained?
  • What independent evidence supports that interpretation of gaze or body language?
  • Are we validating against the same proxy that generated the labels?
  • Which conditions, accessibility needs, or cultural norms might change the signal's meaning?
  • What happens to a learner who is wrongly flagged, and can they contest the interpretation?
  • Would a less intrusive measure answer the actual educational question?

Behind the examples · Make the mechanism inspectable

Read the model. Then interrogate the project.

Karpathy presents a roughly 200-line, dependency-free Python implementation that trains a small generative pretrained transformer (GPT) on names. It includes tokenization, automatic differentiation, a transformer, an optimizer, training, and generation. His tutorial explains the implementation; the right-hand questions below are our translation into leadership practice.

01

Data

Names supply the examples.

Ask of your projectWhich people, tasks, and conditions are absent from our evidence?

02

Tokens

Characters become numbered symbols.

Ask of your projectWhat gets lost when we turn the work into the model's input?

03

Model

Attention combines context; learned parameters produce next-token scores.

Ask of your projectWhat information will the system actually have at the moment of use?

04

Objective

Loss measures how poorly the model predicts the observed next token.

Ask of your projectDoes the score we optimize represent the outcome we need?

05

Training

Gradients and an optimizer adjust the parameters.

Ask of your projectWhat evidence would show improvement on unfamiliar cases?

06

Inference

Fixed parameters generate a sequence through repeated sampling.

Ask of your projectWho checks the result before it changes a decision?

The capability we want to build: explain which part of an AI system a proposed change affects, and ask for evidence at that level. A better training score, a better answer, and a better operational outcome are three different claims.

Further exploration · Change one thing

Temperature changes selection. Does it establish correctness?

Imagine a model choosing between four possible next tokens. Before moving the slider, predict what will happen to the most likely token when temperature falls. Then inspect the distribution.

Illustrative simulation · Fixed, invented scores · No trained model

0.1 · More concentrated2.0 · More spread out
Token A 64.4%
Token B 23.7%
Token C 8.7%
Token D 3.2%

Token A receives 64.4% of the probability. The ranking stays the same.

Inspect the calculation

The scores (logits) are [3, 2, 1, 0]. For each token, we calculate pᵢ = exp(zᵢ / T) / Σⱼ exp(zⱼ / T). The implementation subtracts the largest scaled score before exponentiating for numerical stability. Only T changes. Displayed percentages are rounded.

What should a leader take from this?

There is no answer key in this calculation. A concentrated distribution does not establish that the preferred token is true, appropriate, or useful. Lower temperature can make a system repeat an error more consistently.

Our application: if someone proposes lowering temperature to fix unreliable advice, ask them to measure factual error on representative tasks. The setting is a candidate intervention. It is not the evidence of success.

This original illustration isolates the sampling calculation described in microgpt. It does not run Karpathy's network, learn from your input, or show measured model outputs.

Connect the perspectives

Mechanism, application, and impact

These authors address different questions. The connections below are Capability Matters' interpretation of their published work, not a joint position or a ranking of their views.

Mechanism · Educational demonstration

Andrej Karpathy

Source perspective. Microgpt strips a trainable language model down to an inspectable implementation.

Our connection. Technical literacy should let a leader locate a claim: is the proposed improvement in the data, learning process, generation settings, or system around the model?

Read microgpt · February 2026

Application · Practitioner judgment

Andrew Ng

Source perspective. In his December 2024 letter, Ng argues that applications are accelerating especially quickly, including uses already technically possible with earlier models.

Our connection. A procurement decision should compare complete ways of doing the work. Test the existing process, a model-assisted process, and a simpler alternative. Include the time spent checking and correcting outputs.

Read Ng's letter · December 2024

Impact · Conditional forecast

Dario Amodei

Source perspective. In Machines of Loving Grace, Amodei considers how powerful AI could accelerate progress, while identifying limits such as missing data, physical timescales, and human constraints. The essay is explicitly speculative.

Our connection. Ask what currently limits the outcome. If the constraint is access, authority, equipment, or an unworkable process, better model performance may leave it largely intact.

Read the essay · October 2024

The Capability Matters position: make the mechanism understandable, evaluate the work as a system, and identify the constraint that matters. A leader's job is to make the link from model performance to human capability explicit and testable.

04 · Produce a piece of evidence

A team exercise with two routes

Choose the route that fits your group. Both end with a short decision note: what changed, what the observation supports, what it cannot establish, and what to test next.

Leadership route · About 15 minutes · No coding

Hypothetical scenario. A team proposes an AI assistant that drafts feedback on student work. Its demonstration looks convincing, and its model scores well on a general benchmark. The team wants to expand use across a program.

  1. Define the outcome. What should students or instructors be able to do better? Specify the task and the conditions. “Use AI” is not an outcome.
  2. Choose the comparison. What is the current feedback process? Could a rubric, worked example, or other modest change solve the same problem?
  3. Specify the evidence. Assess feedback accuracy and usefulness on unfamiliar work. Include different student groups and cases the system should refer to an instructor. Measure review time as well as drafting time.
  4. Make the decision conditional. State what would justify a limited pilot, what would stop it, and who has authority to act on the findings. Set those criteria before seeing the pilot results.

Discussion prompt: what did inspecting the mechanism help you decide, and which decisions still require evidence about learners, instructors, and the local workflow?

Builder route · Allow 30–60 minutes · Basic Python

Use Karpathy's original Colab notebook or his source code. Runtime varies by machine. Keep a copy of the source version you run.

  1. Record the baseline. Keep the data, random seed, training steps, and settings fixed. Save the training log and a sample of generated names. Plausible samples alone do not establish generalization.
  2. Separate training from sampling. After one training run, keep the learned parameters fixed and sample at two temperatures. Record the sampling seed and generate enough examples to see variation. Do not retrain between temperatures.
  3. Make an evaluation extension. Before a new training run, split the examples into training and held-out sets. Evaluate next-token loss on the held-out set without parameter updates, using the same temperature of 1.0 for comparisons. This is an added exercise, not a result reported by the original tutorial.
  4. Report the boundary. A held-out name score concerns this task and dataset. It does not measure instruction following, factual accuracy, or readiness to advise students.

Your decision note

“We need [people] to perform [task] under [conditions]. We will compare [alternatives], using [outcome and error measures]. We will proceed only if [criteria]. [Owner] will review failures and decide whether to continue, revise, or stop.”

Use an unfamiliar second scenario to check whether the group can repeat the reasoning. Completing this page is not evidence of that capability.

05 · Keep the claim proportional

What this small model cannot settle

Microgpt is a useful entry point into model mechanics. A production assistant may also involve post-training, retrieval, tools, permissions, monitoring, and human review. None of those capabilities is validated by successfully running this exercise.

Our educational hypothesis is that inspecting a small mechanism improves the questions people ask about larger systems. That is a hypothesis to evaluate through their decisions on new cases. We have not established a learning effect for this page.

Continue with LLM101's depth-adjustable explanation or use the Capability Pipeline to frame a project from an observed capability gap.

Attribution and sources

  1. Andrej Karpathy. “microgpt.” February 12, 2026. Original concept, Python implementation, and technical tutorial. Original code; original notebook.
  2. Andrew Ng. Opening letter in The Batch, issue 281, DeepLearning.AI. December 25, 2024. Cited for his perspective on application progress.
  3. Dario Amodei. “Machines of Loving Grace.” October 2024. Especially “Basic assumptions and framework.” Cited as a conditional argument about potential impact and limiting factors.

Adaptation note. The independently implemented JavaScript training engine, measured checkpoints, synthetic engagement-bias experiment, code explanations, leadership questions, comparison of perspectives, temperature simulation, scenario, and decision-note exercise were developed for Capability Matters with AI assistance. This page credits and links to the originals; it does not reproduce Karpathy's tutorial or Python source text. Implementation attribution and differences. No endorsement by the cited authors is implied.

Terms and acronyms
AI · Artificial intelligence
Computational systems used for tasks such as prediction, generation, and decision support.
GPT · Generative pretrained transformer
A family of models built around transformer architectures for generating sequences.
LLM · Large language model
A language model trained at large scale. Microgpt is deliberately small.
Token
A unit represented by the model. Here the training example uses characters; other models use different units.
Loss
A numerical objective used to evaluate predictions during training or testing.
Temperature
A sampling setting that rescales scores before converting them to probabilities.
Held-out data
Examples excluded from training and used to evaluate performance on unseen inputs.

Under the surface · The actual code

Explain this code

The model source is an independent JavaScript adaptation of Karpathy's teaching architecture. Attribution and differences.

In more straightforward language