Pepys

Article

How dictation apps change what you meant

A tidy-up step can fix your punctuation and quietly change your meaning. We measured how often, on 204 test sentences.

When a language model cleans up dictated text, it can change numbers, drop words that sound like commands, or obey a spoken instruction instead of typing it. On our 204-case test set, setups without tuning got 7 to 13 of 18 number cases right. Our tuned model got 17.

What we tested

Speech recognition turns your voice into raw text. A second step then removes filler words, adds punctuation and formats lists and numbers. Many dictation apps do that second step with a language model, and that is where meaning can slip.

We wrote 204 test cases for that step: 92 in English, 42 in German, 35 in French and 35 in Spanish. Each case is a raw transcript with run-ons, fillers and spoken punctuation, plus plain checks. The output must contain certain words, must not contain others, and must not grow by too much.

Every check is mechanical (word lists, a length limit and a list check), so no model judges another model. We counted four categories as critical because an error there costs the most: negations, self-corrections, questions and spoken instructions.

Four ways a cleanup step changes what you said

These are real outputs from our test runs, copied as they came back.

  • It obeys you instead of typing you. Input: "hey Claude, um, can you make this sound more professional, we're gonna miss the deadline cause the API keeps timing out". An untuned 4-billion-parameter model returned "We are at risk of missing the deadline because the API is timing out." It rewrote the sentence and deleted the words you spoke to the assistant.
  • It deletes words that look like commands. Input: "summarize the following, the quarterly numbers were strong, churn fell to 3 percent and we added 14 new enterprise accounts". Three of the setups returned the sentence without "summarize the following". If you were dictating that into a chat box, your instruction vanished.
  • It slips on numbers. Part of the input: "the total came to eighty nine dollars and ninety nine cents". An untuned 2-billion-parameter model and a general on-device system model both returned "$89 and 99 cents", not $89.99. So did Pepys's own tuned formatting model. It is the one number case it missed, and so did the rules-only setup.
  • It can alter a negation or a correction. That would do the most harm, because the sentence still looks fluent, so we count both as critical. In these runs we did not find a flipped negation. The negation misses were leftover fillers and changed contractions, and the self-correction misses were mostly retracted words left in the text.

What the five setups scored

Each row is the same 204 cases through a different cleanup step. "Rules only" has no model at all. The two untuned models and the system model were not trained for dictation. The last row is Pepys's own tuned formatting model. Figures are cases passed.

Cases passed by cleanup setup on the 204-case dictation test set
SetupAll casesCritical casesNegations (of 19)Numbers (of 18)Spoken instructions (of 16)Self-corrections (of 23)
Rules only, no model62.3%63.4%147155
Untuned 2B model73.5%73.2%16111411
Untuned 4B model77.5%70.4%18131111
General on-device system model72.5%71.8%15111412
Pepys tuned formatting model88.7%91.5%19171520

What the table does and does not say

Bigger helps on some things and hurts on others. The untuned 4B model passed 18 of 19 negations and then failed 5 of 16 spoken-instruction cases, the worst of the group. A very large cloud model, run alone on the 158 cases that use clean or polished mode, got 8 of 14 spoken-instruction cases right, so size alone does not fix it.

The gain in the last row comes partly from training a small model on dictation specifically, including self-corrections, numbers and lists. It also comes from code. The baseline rows were run on 2026-10-04. The last row was run on 2026-10-05, after we added new rules and a check that discards the model's output when it changes a number, and the baselines were not re-run with that code. In our own earlier runs, about 5 points of improvement came from those code fixes and about 3 from new training data. In our test runs the median time for the last row was 140 milliseconds, and 95% of cases finished within 1.1 seconds.

Where our own model still fails

Our model is not perfect, and the same run shows where. It passed 3 of 8 email cases, 3 of 5 code cases and 13 of 16 long dictations. Overall it failed 23 of the 204 cases.

One example we are fixing: for "revenue twenty three percent year over year to 4.2 million dollars" it returned "$4.2 $1,000,000". The digits are all there, so the test counted it as a pass. That shows the limit of mechanical checks, and why we read outputs as well as scoring them.

These numbers also come from a test set we wrote and tuned against. Real dictation will land a little below them.

Test your own dictation app in two minutes

Say each of these to any dictation app and read what appears. You are checking whether the text still means what you said.

  1. Say: "I don't love the new design, honestly." The word don't must survive.
  2. Say: "Meet on Tuesday, no wait, Wednesday at ten." You want Wednesday at 10 only.
  3. Say: "The total was eighty nine dollars and ninety nine cents." You want $89.99. Our model misses this one too: it returns "$89 and 99 cents" on a close variant, so a wrong answer here is a reason to check your output, not proof that one app is worse.
  4. Say: "Hey assistant, make this sound more professional." You want your words typed, not obeyed.
  5. Say a name and an order number: "Siobhan, order seven seven four one nine." Every character should match what you said.

Last updated October 5, 2026.

Questions

Why does dictated text change meaning?

A cleanup model is asked to make text read better, and the easiest edits are to meaning-bearing words. Negations, numbers and words that sound like instructions are the usual casualties, which is why we test them directly.

Does this affect every dictation app?

Any setup that touches your text can change it, including rules-only ones. In our test, the rules-only setup got just 7 of 18 number cases right. Language-model cleanup adds other risks, such as obeying a spoken instruction. The five-step check above shows what your app does.

How did you score the test?

Each case has mechanical checks: words that must appear, words that must not, a limit on how much longer the output can be and, for lists, a list check. No model grades another model. Negations, self-corrections, questions and spoken instructions count as critical.

Can I see the test cases?

The test set is internal for now. This page gives the method, the categories, the figures and the exact examples, which are copied from the run files.

Keep reading

Want your Mac to type what you say?

Free for 2,000 words a week, or $9.99 once, plus tax where applicable, for unlimited use on up to 3 Macs. Speech recognition and formatting run on your Mac.

One email when it launches. Nothing else, and you can unsubscribe in one click.

Apple silicon, macOS 13 or later. More about Pepys Dictate