Language comes from the world around us. So I made the world slightly more annoying.
In my fun alternate universe, human mouths cannot make an a-sound. People can still hear it. They can still remember it. Their mouths have simply resigned from that part of the job.
I started with the easiest-looking version: remove the written a. Deleting one keyboard character seemed easier than rebuilding a civilization.
The scorecard
Each round fixed one problem and exposed another:
| Round | What I tried | Honest result |
|---|---|---|
| First model sweep | Block the written a | 9/9 samples avoided a. The stories made little sense. |
| Text-packing contest | Fit more story into each token | The 16,000-piece tokenizer packed text best. It also made the model bigger. |
| Longer training | Turn legal text into readable stories | 3/3 samples avoided a. The raw stories were still broken. |
| Whoops test | Say whoops for mishaps; invent a word for everything else | 4/4 creation prompts produced legal words. 0/2 mishaps produced whoops. |
| Goblin stories | Write a complete story in the invented language | About 91% of the words were legal. 0/4 raw stories were complete. |
| English audit | Check whether the goblin words were secretly English | Only 9.09% matched the English list. Under the sound rule, only whoops survived. |
| Final English system | Tell a positive English story without the missing sound | Raw model: 0/4. Model plus sentence rails: 4/4. |
| Live demo | Run the final system from this page | Public request: 200. A cold start can take about 30 seconds. |
| Widened SFT | Make the raw model follow supported subjects | Language and subject pass rates each reached 96.53% on its 144-story evaluation. |
| VPO v2 | Make a group of stories more useful to search | Best@12 improved. Single-story reliability slipped. Three-detail prompts remained unsolved. |
The final model has 10,932,736 adjustable knobs. The runs through the first finished English system cost $1.16026396. The later VPO v2 smoke and full run added $0.07785851. We spent just over one dollar teaching a tiny computer that a box should not go to the box, then paid seven more cents to hold auditions.
How small is small?
I started with three models from “TinyStories: How Small Can Language Models Be and Still Speak Coherent English?”, by Ronen Eldan and Yuanzhi Li. The paper uses simple synthetic stories with young-child vocabulary, but the stories still have people, objects, problems, and endings. It is basically a tiny universe where someone always loses a toy near a tree.
The models were named TinyStories-1M, -3M, and -8M. Those names count only the central block. Once I included the large word tables at both ends, the real totals were about 3.75 million, 8.28 million, and 19.70 million knobs.
The 8M model was nearly 20M.
Tiny!
Each model began with an empty little brain and received 800 practice rounds. Every model saw the same stream of stories. Then I blocked all 19,636 text pieces containing a or A.
To test them, I gave every model a dog and a box. This was generous. The shared opening was:
The little dog looked into the box.
The smallest model immediately misplaced the dog, the box, and possibly the concept of indoors:
The little little girl is so proud. She would keep out! It lived in the boy.
The biggest model remembered the box. Unfortunately, it remembered only the box:
The girl's mom took the box to the box. Her mommy took the box to the box.
Every blocked sample avoided the forbidden letter. The plot was last seen entering a box.
This first sweep cost $0.18283310, including one failed startup. I had successfully purchased proof that a machine can obey one rule while understanding approximately none of the occasion.
Maybe the words were packed badly
The machine does not read words the way we do. It chops text into reusable pieces. If the pieces are awkward, a tiny model spends too much of itself carrying the dictionary and not enough remembering that the dog exists.
So I made three new text choppers: one with roughly 4,000 pieces, one with 8,000, and one with 16,000.
The 16,000-piece version packed ordinary stories best. It fit about 4.19 characters into each piece and used about 1.22 pieces per word. It also produced a 10.93-million-knob model, because a bigger dictionary is still bigger even when you call it compression.
I trained the winner for 10,000 rounds over 81.92 million packed text pieces. Its prediction score improved by 28.1 percent.
Wonderful. Surely it could now tell a story.
Suddenly, the old kid: there were two of three1O cents�Lr C©©
No.
Another attempt discovered a powerful new literary technique called saying “upon” until everyone leaves:
Once upon upon upon: there there were two boys were E�r C©©
My checker, built with Prime Intellect's verifiers library, still gave these blocked samples a high average score: 0.934583. It checked three things: no forbidden letter, enough length, and not too much repeated phrasing. It did not check whether the result resembled something a person might willingly finish reading.
The machine had found the exam answer sheet. The exam had forgotten to ask about stories.
There was also a small cloud incident. The final worker was supposed to continue from round 2,000. It saw an older view of the saved files and restarted from zero instead. The checkpoint existed. The worker simply behaved like a student staring directly at completed homework and opening a fresh document.
I fixed that by forcing the worker to refresh the shared storage before looking for its checkpoint. The whole compression run cost $0.86731718, still below the $1.50 limit.
Useful evidence! Slightly expensive evidence wearing a fake moustache, but useful.
Fine. Just say “whoops”
Full stories were going badly, so I made the job smaller.
For a mishap, say exactly whoops. For a creation, invent one pronounceable word. Never use the missing letter or sound.
This looked difficult to misunderstand.
I trained the model for 1,000 more rounds using 32,000 examples. It learned to produce neat little made-up words. It did not learn when to say whoops.
A clockwork bird fell apart. The model replied:
zudruf
A tiny rainbow vanished. The model replied:
foogloov
All four test replies were legal invented words. All four were invented words. The model reacted to every emergency like a toddler naming a new dinosaur.
The outside program now returns whoops whenever the caller marks a prompt as a mishap. That guarantee comes from the program around the model, not from the model suddenly developing concern.
The run cost $0.03551419. One earlier attempt finished training, then fell over during testing because I passed the wrong option into the text generator. It did not say whoops. I did.
The goblin novel phase
One invented word was not enough. I wanted a complete story, so I trained the same model on three-sentence tales made from whoops and a 512-word invented vocabulary.
It received 32,000 examples and 1,000 more practice rounds. This took 99.78 seconds and cost $0.02886985.
The raw model produced this:
veroovuzul veregop krootroot voovuzur lefifglooz. whoops skuzop soovuf vuruzuzux zegopetrepox reezuskux gree
Honestly, it has energy.
Every test story avoided a. About 91 percent of its words passed the invented-word rules. But the generator sometimes stopped in the middle of its final word, like someone unplugged the goblin during a speech. None of the four raw stories passed the complete test.
The downloadable runner repairs chopped words, adds enough words when a story ends early, and restores three sentence endings. With that guardrail, all four test stories passed.
This distinction matters. Training made good behavior likely. The outside program made the promise reliable.
The invented language still had no proven meaning. A word could return several times, but we had not shown that it kept referring to the same person, object, or event. We had created excellent noises. Society remained pending.
Could it please use English?
I then checked whether the goblin stories were secretly English.
They were not secretly English.
Only 9.09 percent of the raw words appeared in the English frequency list. Under the stricter pronunciation check, only whoops survived everything.
So the next request became: write a full, positive story using only verified English words, still without the missing sound.
This is how projects grow. You begin by deleting one letter. Soon you are operating a customs desk for every word in the language.
I went back to the clean 10,000-round model and trained it on 48,000 short English stories. Each story followed a simple shape: someone loves something, something goes wrong, the hero feels low, the hero takes one step, a friend helps, and everyone goes home feeling better.
That training took 137.77 seconds and cost $0.04572964 including final testing.
The final English model still became repetitive and failed every evaluation prompt without the sentence rails.
So I gave it sentence rails.
The model chooses six things: the hero, two places, an object, a helper, and a tool. The program supplies the safe sentence structure around those choices. It does not generate a broken story and swap in a secret backup. Every word enters through the checked route. If the result breaks a rule, the runner raises an error instead of pretending nothing happened.
Here is one result:
once there lived one robot in the river. the robot loved one key. one evening the key fell into the tree. the robot felt low but kept hope. the robot took one step. one queen met the robot. the queen used one hook to lift the key. the robot felt joy. both felt good. they went home together.
Every final test story passed once the sentence rails were active. Every output word was English. None used the written a or the forbidden pronunciation family. The final model has 10,932,736 adjustable knobs, and its downloadable weight file is about 42 MB.
It is not secretly writing an unrestricted masterpiece inside its head. This tiny model has no separate stream of hidden reasoning words. It performs math, then chooses output pieces. The rails control every piece that becomes text.
Also, the input can contain a. The restriction applies to the model’s voice, not to the poor person trying to ask it for a dog story.
Try the tiny model
Type any story idea below. The first request can take about 30 seconds because this tiny computer enjoys one tiny nap.
Give the tiny model a problem
Your prompt can use a. Its story cannot.
The first story can take about 30 seconds while the model wakes up.
Update: the model finally noticed the prompt
The first English model could make a legal story with sentence rails, but it could still ignore the requested subject. I could type computer and receive a fish. This is technically a story generator. It is also technically rude.
I made 32,000 wider training examples across 24 subjects. Some subjects needed legal substitutes: cat became kitty, airplane became jet, and computer became robot. After 1,000 supervised updates, the raw model reached a 96.53 percent language pass rate and a 96.53 percent subject pass rate on 144 held-out stories.
The boring thing worked. I showed the model examples. Disgusting. I had hoped for a more mysterious answer.
Then I gave VPO something to argue about
My first VPO run on the widened model changed almost nothing because the easy subject test was already nearly solved. The optimizer entered a room full of correct homework and tried to identify the most correct corner.
So I made the exam harder. Each new prompt requested three things: one subject, one object, and one place. Each received its own score. The model produced three stories in one chain, so the later stories could see the earlier ones and try something different.
Here is the held-out result. Both models received the same 12 prompts and produced 144 stories.
| What I measured | SFT only | SFT + VPO v2 |
|---|---|---|
| Valid single stories | 83 / 144 | 80 / 144 |
| Valid three-story groups | 48 / 48 | 47 / 48 |
| Best score after 12 tries | 0.730142 | 0.740155 |
| Reward-space spread | 1.777728 | 1.812602 |
| Stories using all three requested details | 0 | 0 |
If I need one random story, SFT is slightly safer. If I can request 12 stories and keep the best one, VPO v2 gives the verifier a slightly better pile to search.
VPO did not make one tiny writer clearly smarter. It made the audition room slightly better.
Also, neither model used all three requested details even once. The audition improved. Nobody got the role.
Try SFT and VPO v2
Give both raw models one supported subject, one object, and one place. They receive the same prompt and random seed. Each writes three stories in a chain. If none passes every rule, the demo shows the closest attempt instead of quietly inventing a winner.
SFT only
the widened model after 1,000 supervised updates.
its selected story will appear here.
Did I do it? Yes. Did the model do it alone? Less than I wanted, more than before.
The first finished text system passed 4/4 evaluation prompts. Its raw English model passed 0/4 without sentence rails. Later, wider supervised training made the raw model follow its supported subjects on a larger evaluation. VPO v2 made groups of candidates slightly better to search, but it did not solve the harder three-detail request.
That is the answer.
The finished system produced verified stories for all four original evaluation prompts. None wrote a or chose words from the tested a-sound family. The model picked six details. The checked sentence rails kept the story legal. It was a group project, and the model kept putting its name first.
Four original prompts and 12 later prompts are still tiny samples. This is not proof that the machine can handle every prompt ever typed by a human with Wi-Fi.
Did I teach one tiny model to write freely in a whole new language? No. It invented excellent goblin noises, failed an English inspection, and spent some time moving a box to the box.
Did I build a working text system for this exact silly rule? Yes. Did I build one raw model that follows every detail without help? No.
No mouths, dialects, or audio entered the experiment. A spoken version would need a voice system checking the actual noises coming out. That sounds like a problem for future me, who has done nothing to deserve it.
So yes. I did it at the system level, in text. The raw model learned more of the job later. It still has open positions.
That is an annoyingly specific yes, which is still yes.
The robot got its key back.