20 Functions Is a Result. Not General Intelligence.

Discovery-Tiny-001 reproduces 20 short Python functions from their six-token prefixes. The public result is 20 out of 20 exact matches and 20 out of 20 parse checks on the training set. That sentence is intentionally narrow. The model is not a general-purpose coding assistant, and a perfect score on these functions does not make it one.

What the experiment actually asks

The project explores discrete diffusion for code rather than autoregressive generation. Its documented architecture is absorbing discrete diffusion: an encoder reads a short prefix, and a decoder denoises the masked completion through iterative unmasking. The experiment asks whether a model trained from scratch can memorize and reproduce a small fixed set of code exactly.

The published configuration is about 92 million parameters, 256-dimensional representations, eight layers, and a roughly 151,000-token vocabulary. The README notes that the embedding table dominates the parameter count. That is a useful reminder that a headline parameter total does not, by itself, describe where a model spends its capacity.

Exact match has a job

The training run uses 20,000 steps on 20 short Python functions. The evaluation checks whether the output is byte-exact and whether it parses. Parse success and exact match answer different questions. Code can parse while being the wrong function. Exact match is a strict way to check reproduction when the target is already known.

The functions are small and fixed: arithmetic operations, comparisons, parity checks, and short transformations such as square and cube. The project lists them explicitly. There is no need to turn that finite set into a vague story about broad coding ability. The constraint is what makes the experiment readable.

Training-set success is not a held-out claim

The result belongs to the training set. Outside the documented prefixes, the README says output is undefined. The model does not chat, follow instructions, or establish generalization. Those are not footnotes to hide after a large score. They are part of the result itself.

A memorization artifact can still be useful. It shows that the training and generation path can meet a precise target, and it gives the next experiment a baseline. It does not prove that the same mechanism will handle new functions, longer programs, unfamiliar prefixes, or a different distribution.

Keep the artifact inspectable

The public repository includes inference code, the tokenizer, model definitions, and the canonical training script. The checkpoint is a release asset. The README also documents the seed and a generation-aware health gate used during training. Those details help someone distinguish the actual procedure from a retrospective success story.

This post does not claim an independent reproduction of that run. It explains the published artifact and its stated boundary. Reproducing it means obtaining the checkpoint, running the included evaluation path, and comparing actual generated functions with the known targets, not simply repeating the score.

The next stronger claim would require a stronger experiment: held-out functions, a defined split, disclosed failure cases, and a clear generation protocol. Until that evidence exists, the honest accomplishment is a from-scratch diffusion code model that does this small thing exactly. A research result gets stronger when the wording stops where the evidence stops.

Public source

Artifact, training configuration, function list, evaluation result, and limitations: https://github.com/yethikrishna/discovery-tiny-001 . The reported 20/20 result is the project's training-set result, not a new benchmark run performed for this article.

← back to the journal