Talk Type · Episode 19 · 4 min ·
The Transcription Mountain Between Fieldwork and Your Thesis
A qualitative PhD runs a couple of dozen interviews, and typing them up by hand can eat weeks. Here is the real math, why your ethics board changes the tool you can use, and why transcription is a methods decision, not clerical work.
Transcript
This is Talk Type, from the team at Pepys, where we turn talk into text.
Here's the part of a PhD nobody warns you about. It sits between the fieldwork and the analysis, and it's just, typing. You've done the interviews. Now you have to turn them into text before you can do anything with them.
And there are more of those interviews than you'd think. Mark Mason went through five hundred and sixty doctoral dissertations that used qualitative interviews, and found the typical study wasn't three interviews, or thirty. It was a couple of dozen. A median of twenty-eight. The most common numbers were twenty and thirty.
Why so many? Because you keep interviewing until new interviews stop teaching you anything new. Researchers found code saturation, hearing the full range of issues, at around nine interviews. But truly understanding each issue took sixteen to twenty-four. So a thesis that wants depth lands right back in the dozens.
Now turn those interviews into hours, because hours are what you actually transcribe. A normal qualitative thesis is something like twenty to forty hours of recorded audio, sitting on your drive after fieldwork.
And typing it up by hand is slower than you remember. A University of Bath guide puts it at four to seven hours of work for one hour of audio. Run that against the pile. Twenty-eight hours of tape, at the fast end, is about a hundred and twelve hours of transcription. Three full working weeks at the keyboard, before you code a single line.
So most people reach for software, and that's the right instinct. But two things about a thesis make it different from any other transcription job.
The first is your ethics board. Interview recordings are treated as confidential by default. One university's research program says audio is always Confidential data. Another requires the files stay on managed devices, and that records be kept at least three years after the study ends. So "where does my audio go, and does anyone train a model on it" isn't paranoia. It's paperwork. This is exactly why we never train on your audio.
The second thing is that transcription itself is a methodological choice, not clerical work. Your committee expects you to name the convention you used and defend it. Naturalized verbatim keeps every pause and overlap, for conversation or discourse analysis. Denaturalized clean verbatim tidies the grammar, and it's usually enough for thematic coding. Pick one before you edit a line, and write a short paragraph in your methods chapter explaining why. Examiners look for it.
One more small thing that saves hours later. Keep the speaker labels clean. Coding software keys on who said what, so an interviewer's prompt shouldn't get coded as a participant's answer. Start each turn with the name and a colon, and NVivo or MAXQDA codes by speaker the moment you import.
The transcription mountain is real. But it's a season, not the whole climb. Get the pile off your drive and into text you can defend, and the analysis is finally allowed to begin.
That's this episode of Talk Type. The full write up, with the links and sources, is in the show notes. Pepys transcribes any file or link, any length, pay once, and we never train on your audio. Your first sixty minutes are free at pepys dot co. Thanks for listening, and we'll see you next time.