THE RESEARCH, IN REAL-WORLD UNITS · DATA AS OF 2026-08-25

What does it take to teach a phone first aid?

Everything below is computed from our own logs and ledgers — the same ones our published results come from. Rounded, but never invented.

The corpus we wrote in one August week

77,771 answers

Between Sunday evening and Tuesday evening, a 30-billion-parameter teacher model sat on rented H100s and wrote 77,771 emergency answers — 13.6 million words, about 140 paperbacks, at 175 words each. Not one of them was written for a person to read. They are the textbook a four-billion-parameter model studies so it can fit in a pocket and answer with the signal off.

Here is the pantry filling up, every fifteen minutes of it. The line is flat where the run stalled and steep where we added machines — we are showing you the stumbles because the shape of a real run is the interesting part.

This curve is drawn with JavaScript. Every number around it is here without it.

Hover or tap the curve for that quarter-hour’s numbers.

The flat stretch is the honest part. Overnight the work stopped — a dropped network, a closed laptop lid, and a proxy that quietly throttled the whole fleet to one answer a minute. Nothing was lost: every answer already written stayed written, and the run picked up from the last one on disk when the machines came back. Then we went from one rented GPU to three, from eight workers to twenty-four, and the curve turned upward — 6,433 answers in the peak hour, against about 1,400 on the first night.

We stopped 203 answers short of the plan — 99.73% of it — because the last quarter of one percent was not worth another hour of rented GPU. That gap is written down in our records rather than quietly rounded away, and it is why two honest numbers live on this page: 77,771 answers we have, and 99.73% of the answers we set out to write. They differ because the plan changed mid-run — we cut from three drafts per question to two, and kept the 1,539 third drafts already written.

77,771 complete answers counted from the run’s own output files — finished cleanly, non-empty, counted once per prompt-and-take, which is the same test the pipeline applies (an answer cut off mid-sentence is an attempt, not a result) · each row carries the minute it was written, which is where this curve comes from · 75,780,433 characters of answer text · teacher: an open 30B model, self-hosted · 48.2 hours wall clock · as of 2026-08-25

The thinking we threw away

292 novels, deleted

Before each answer, the teacher reasons — and we keep none of it. That discarded deliberation came to 157 million characters, about 26.2 million words: 2.1× more thinking than the corpus it produced, gone the moment it arrived. It is most of what we actually rented the GPUs for, and we measured it rather than guess, because it is the real price of a corpus like this one.

reasoning characters reported per answer by the serving stack, summed over 77,771 answers · the model has no off switch for it, only a dial

The bookshelf we trained on

3,846 novels

The training text this lab has built, filtered and frozen across every experiment — about 2.08 billion characters of emergency medicine, survival doctrine and hand-written examples. As paperbacks, that's a bookcase forty shelves long.

2,076,754,867 chars ÷ 6 chars/word ÷ 90,000 words/novel · counted from every frozen dataset on disk

The paper stack

115 meters

Print it all — 1.15 million pages — and the stack stands taller than the Statue of Liberty, torch included. A 35-story building of paper, distilled into a model that fits beside your camera roll.

1.15M pages × 0.1 mm/sheet · Statue of Liberty: 93 m

Books our machines wrote back

434 novels

To find out what the models actually learned, we make them answer — a lot. Their practice answers and teaching drafts total ~39 million words. Nobody reads them all; that's the next number's job.

234,104,341 generated chars, counted from every response and teacher file on disk

The strictest grader alive

15,879 graded answers

Every answer is scored by an AI judge against a written checklist for that exact emergency — the judge's grading notes on disk run to ~9.7 million words. A human instructor at three minutes per answer would need five months of full workdays. Ours does it in an afternoon, and the hard calls get a second, stricter judge.

counted from every verdict file on disk · 15,879 × 3 min ÷ 8h days

54 answers per emergency

346 scenarios × 54

Our test is 346 frozen emergency scenarios — never trained on, never edited to make a model look good. Across every model and every experiment, each one has now been answered about 54 different ways. We keep the failures on file with the wins.

18,538 evaluated answers ÷ 346 scenarios

An encyclopedia habit

37 years of reading

Our open research line indexed all of English Wikipedia — 7.6 million articles, sliced into 53 million searchable passages. Reading it yourself at a brisk 250 words a minute, around the clock, no sleep: see you in 2063.

≈4.9B words ÷ 250 wpm · corpus figures as published in our lab brief

The arithmetic

1,946 humanity-years

Our rented GPU time adds up to roughly 5×1020 calculations. If every person on Earth did one calculation per second — eight billion pencils scratching in unison — it would take about 1,946 years to match what we bought for less than the phone it has to run on.

136.5 rented GPU-hours (billing, 2026-08-13 to 2026-08-27) × ~10¹⁵ ops/s ÷ 8B people · GPU hours from our spend ledgers

The electric bill confession

7,644 phone charges

All that compute drew roughly 96 kWh — about what your phone uses in twenty years of nightly charging. Small on purpose: this lab's method is doing more with less, because the product has to run on the least hardware of all — yours.

≈96 kWh (GPU-hours × 700 W) ÷ 12.5 Wh per phone charge

The map of what training changed

3.49 billion dials

Training doesn't rewrite a model, it nudges it — and this is the nudge, counted. We put the finished model beside the open one we started from and compared all 723 weight tables, number against number. 3.49 billion of the model's 4.54 billion numbers came out different — 76.8% of them. Read the changed ones aloud at one a second and you'd finish in 110 years.

Every square below is one table — a row per layer, top to bottom, and a column for each kind of table a layer can hold. Darker means training moved that table further, measured against how big the table already was.

This map is drawn with JavaScript. Every number around it is here without it.

Hover or tap a square for its real numbers.

The lit squares run from 4.5% — layer 31's down-projection, the quietest thing that moved at all — to 13.4% at layer 17's gate. That is a three-fold spread across 200 tables, with no single layer running away from the rest. The blank squares aren't missing data: this model alternates two kinds of attention block, so 24 layers carry one set of tables and 8 carry the other. And the ten kinds of table that lit up are exactly the ten the training recipe declares it aims at — none extra, none missing.

What this map is not: a picture of where medicine lives. Nothing here says burns are stored in layer 12. It shows where the update concentrated — which parts of the machinery had to move for the model's answers to change — and says nothing about what any of them mean.

The update is also narrower than 3.49 billion sounds. The entire change is described by 61 million numbers: 1.3% of the model, a 244 MB file (233 MiB if your file manager is a pedant). We factored six of the changed tables to see how it is spread, and in every one the top 32 directions carry 99.92%–99.98% of the change. Billions of dials moved, but the whole update hums in a few dozen directions per table — which is how this training method is built to work, and now we have watched it do it.

The dark half is the point. 523 of the 723 tables came out carrying no change at all — 475 of them bit-for-bit identical to the file we started from, including the 636-million-number vocabulary table and the whole 334-million-number image-reading stack (the other 48 differ only in how the same values are stored). Identical is a claim you can check, and checking it is why this picture gets taken. A training run that quietly does nothing still reports success; the tell is a map with nothing lit on it. That has happened here — a merge once handed back a model byte-for-byte the same as the one it started from, while every log said the job finished clean. Nothing gets called finished now until it has sat for this portrait.

723 tables compared element-wise against the open base model we started from · shade = size of the change ÷ size of the table (Frobenius norms), linear scale — the lit range spans 3×, so nothing needed a log · “came out different” = a value moved by more than one part in a million · 523 identical: 475 byte-for-byte, 48 differing only by a precision cast that round-trips exactly · 15 base-only tables (a spare prediction head, dropped when the training was merged in) are recorded in the data and not mapped · change file on disk: 243,854,336 bytes · measured Aug 2026

And it all fits in your pocket

2.8 GB

The point of the bookcase, the paper tower and the humanity-years: a model smaller than an hour of 4K video that answers in the first 30 seconds — with no signal at all. The mountain goes in, the pocketknife comes out.

shipping model file size, quantized for iPhone · see the measurements
Counted, not conjured: sources are our committed eval files, spend ledgers and the lab brief. Rounding is generous, direction is honest.
← Red Kit · Red · Hardware
← Red Kit · Privacy · Terms · Support · Contact