When a language model cleans up dictated text, it can change numbers, drop words that sound like commands, or obey a spoken instruction instead of typing it. On our 204-case test set, setups without tuning got 7 to 13 of 18 number cases right. Our tuned model got 17.
What we tested
Speech recognition turns your voice into raw text. A second step then removes filler words, adds punctuation and formats lists and numbers. Many dictation apps do that second step with a language model, and that is where meaning can slip.
We wrote 204 test cases for that step: 92 in English, 42 in German, 35 in French and 35 in Spanish. Each case is a raw transcript with run-ons, fillers and spoken punctuation, plus plain checks. The output must contain certain words, must not contain others, and must not grow by too much.
Every check is mechanical (word lists, a length limit and a list check), so no model judges another model. We counted four categories as critical because an error there costs the most: negations, self-corrections, questions and spoken instructions.
Four ways a cleanup step changes what you said
These are real outputs from our test runs, copied as they came back.
- It obeys you instead of typing you. Input: "hey Claude, um, can you make this sound more professional, we're gonna miss the deadline cause the API keeps timing out". An untuned 4-billion-parameter model returned "We are at risk of missing the deadline because the API is timing out." It rewrote the sentence and deleted the words you spoke to the assistant.
- It deletes words that look like commands. Input: "summarize the following, the quarterly numbers were strong, churn fell to 3 percent and we added 14 new enterprise accounts". Three of the setups returned the sentence without "summarize the following". If you were dictating that into a chat box, your instruction vanished.
- It slips on numbers. Part of the input: "the total came to eighty nine dollars and ninety nine cents". An untuned 2-billion-parameter model and a general on-device system model both returned "$89 and 99 cents", not $89.99. So did Pepys's own tuned formatting model. It is the one number case it missed, and so did the rules-only setup.
- It can alter a negation or a correction. That would do the most harm, because the sentence still looks fluent, so we count both as critical. In these runs we did not find a flipped negation. The negation misses were leftover fillers and changed contractions, and the self-correction misses were mostly retracted words left in the text.
What the five setups scored
Each row is the same 204 cases through a different cleanup step. "Rules only" has no model at all. The two untuned models and the system model were not trained for dictation. The last row is Pepys's own tuned formatting model. Figures are cases passed.
| Setup | All cases | Critical cases | Negations (of 19) | Numbers (of 18) | Spoken instructions (of 16) | Self-corrections (of 23) |
|---|---|---|---|---|---|---|
| Rules only, no model | 62.3% | 63.4% | 14 | 7 | 15 | 5 |
| Untuned 2B model | 73.5% | 73.2% | 16 | 11 | 14 | 11 |
| Untuned 4B model | 77.5% | 70.4% | 18 | 13 | 11 | 11 |
| General on-device system model | 72.5% | 71.8% | 15 | 11 | 14 | 12 |
| Pepys tuned formatting model | 88.7% | 91.5% | 19 | 17 | 15 | 20 |
What the table does and does not say
Bigger helps on some things and hurts on others. The untuned 4B model passed 18 of 19 negations and then failed 5 of 16 spoken-instruction cases, the worst of the group. A very large cloud model, run alone on the 158 cases that use clean or polished mode, got 8 of 14 spoken-instruction cases right, so size alone does not fix it.
The gain in the last row comes partly from training a small model on dictation specifically, including self-corrections, numbers and lists. It also comes from code. The baseline rows were run on 2026-10-04. The last row was run on 2026-10-05, after we added new rules and a check that discards the model's output when it changes a number, and the baselines were not re-run with that code. In our own earlier runs, about 5 points of improvement came from those code fixes and about 3 from new training data. In our test runs the median time for the last row was 140 milliseconds, and 95% of cases finished within 1.1 seconds.
Where our own model still fails
Our model is not perfect, and the same run shows where. It passed 3 of 8 email cases, 3 of 5 code cases and 13 of 16 long dictations. Overall it failed 23 of the 204 cases.
One example we are fixing: for "revenue twenty three percent year over year to 4.2 million dollars" it returned "$4.2 $1,000,000". The digits are all there, so the test counted it as a pass. That shows the limit of mechanical checks, and why we read outputs as well as scoring them.
These numbers also come from a test set we wrote and tuned against. Real dictation will land a little below them.
Test your own dictation app in two minutes
Say each of these to any dictation app and read what appears. You are checking whether the text still means what you said.
- Say: "I don't love the new design, honestly." The word don't must survive.
- Say: "Meet on Tuesday, no wait, Wednesday at ten." You want Wednesday at 10 only.
- Say: "The total was eighty nine dollars and ninety nine cents." You want $89.99. Our model misses this one too: it returns "$89 and 99 cents" on a close variant, so a wrong answer here is a reason to check your output, not proof that one app is worse.
- Say: "Hey assistant, make this sound more professional." You want your words typed, not obeyed.
- Say a name and an order number: "Siobhan, order seven seven four one nine." Every character should match what you said.
Last updated October 5, 2026.