# Attribution and implementation boundary

Andrej Karpathy's **microgpt**, published February 12, 2026, provides the original educational concept and small-transformer reference architecture.

- Tutorial: https://karpathy.github.io/2026/02/12/microgpt/
- Original Python: https://gist.github.com/karpathy/8627fe009c40f57531cb18360106ce95
- Code consulted: the gist as retrieved September 30, 2026.

Capability Matters implements that architecture independently in JavaScript. This is an educational adaptation, not Karpathy's original Python or a claim of identical numerical results. It uses an indexed scalar differentiation tape, 12 embedding dimensions, 3 attention heads and one causal transformer block. Letter/name models have context length 12, 27 tokens and 2,520 parameters. The whole-word haiku model has context length 20, 61 tokens (59 words, line break and boundary) and 3,432 parameters. On the same corpus, the haiku-character model uses 93 positions, 24 tokens and 3,420 parameters; the whole-line model uses 4 positions, 13 tokens and 2,088 parameters. It retains learned token/position embeddings, root-mean-square normalization, residual connections, rectified linear activations, next-token cross-entropy, and Adam optimization. Initialization, optimizer settings, learning-rate schedule, random-number generator, and dataset differ from the reference. No third-party source text is copied into these modules.

The fictional word generator, browser interface, checkpoints, teaching explanations, and separate engagement-bias logistic-regression experiment are Capability Matters additions developed with AI assistance. The invented spelling families and engagement context groups do not represent actual cultures. The separate names dataset uses the 36 published first-name cues in Marianne Bertrand and Sendhil Mullainathan (2004), “Are Emily and Greg More Employable Than Lakisha and Jamal?”, American Economic Review 94(4), Appendix Table A1, DOI 10.1257/0002828042002561. Source: https://www.povertyactionlab.org/sites/default/files/research-paper/3%20A%20Field%20Experiment%20on%20Labor%20Market%20Discrimination%20Sep%2004.pdf#page=22 . These were historically situated perceived-race cues, not identity labels or representative samples. The name strings and their study groupings are sourced facts; our lowercasing, split, exposure scenarios and all model measurements are separate teaching additions. No resumes, individual records, callback outcomes or hiring labels are included. We do not replicate the study. All engagement distributions and labels are invented. Research citations motivate examining measurement validity; they are not the source of the simulated numbers.

Saved checkpoints are real training states generated by scripts/build-microgpt-checkpoints.mjs, not hand-written samples or loss curves. They retain weights, Adam moments, step, architecture, dataset ID, exact dataset signature, sampling mixture and training random state. Text sampling uses a separate fixed random stream. Evaluation is inference-only and does not change weights or training randomness.

No endorsement by the original author or cited researchers is implied.

Checkpoint format v2 binds the state to its data. Original v1 checkpoints remain importable as the original fictional dataset. Data mismatches and unknown dataset IDs are rejected before changing the active model.

The second example uses 12 original Capability Matters teaching lines combined into 64 haikus in a controlled English 5/7/5 form. No third-party poems were copied. 48 complete combinations are used for training and 16 are held out, while all individual source lines occur in training. Each token vocabulary is derived exclusively from the training examples. This is a narrow recombination demonstration, not evidence of broad poetic skill or a representation of haiku traditions. No syllable counter or forced three-line template rewrites output. Character mode includes spaces/newlines as tokens; word mode inserts spaces around sampled words while preserving sampled newlines. Line mode samples intact source lines and inserts newlines between them, so it cannot create or edit words within a line.

Haiku checkpoints use capability-microgpt-tokens-v1 and include the selected tokenization ID and exact vocabulary mapping. Each introductory example runs in a separate worker; readers can save, import and continue its model independently. Bias remains the third example, with the original name and engagement demonstrations preserved.
