Pepys

Talk Type · Episode 13 · 4 min ·

From Recording to Coding-Ready: Transcribing Research Interviews

The end-to-end researcher workflow for turning interviews into coding-ready transcripts: capture clean, get an AI first pass, pick a verbatim style, anonymize, and export a file your analysis software actually opens.

Transcript

This is Talk Type, from the team at Pepys, where we turn talk into text.

Here's the part of a research project nobody budgets for. The hours after the interview ends, when a recording has to become text you can actually code.

Transcribing it yourself is the slow way. Researchers have measured it at up to six hours of typing for a single hour of audio. That's most of a working day spent not analyzing, just capturing. And here's the thing worth remembering. Transcription isn't clerical. The person deciding where a sentence ends, whether a pause counts, whether a laugh belongs on the page, is already doing analysis. The transcript is data, not a neutral record.

So start upstream, with capture. The recording sets a ceiling on everything after it. For a remote interview, record each participant on their own channel if the platform lets you. Separate tracks mean the tool isn't guessing who's talking when two voices overlap, and clean diarization, which is just labeling who said what, saves you fixing turns by hand later.

Then the first pass. Let a tool produce a speaker-labeled, timestamped draft in minutes, and spend your attention on cleanup instead of typing. Read it against the audio. The machine handles the bulk. You handle the load-bearing five percent. The names, the jargon, the acronyms, the moment two people talk at once. And if a stretch is genuinely unclear, bracket it as inaudible with the timestamp. A flagged gap is honest. A confident wrong quote is a correction waiting to happen.

Before you clean, pick a style, because it changes every line. A naturalized transcript keeps every utterance in detail, and that's what conversation and discourse work needs, where the pauses and overlaps are the data. A denaturalized transcript corrects grammar and drops the interview noise, and reads better for thematic or content analysis, where you're after meaning across cases. Neither is more correct. Just pick one, and apply it to every recording in the study, or your cases stop being comparable.

Then anonymize, because most qualitative recordings are human-subjects data. As you clean the draft, swap names, employers and places for role labels. The HIPAA Safe Harbor list of eighteen identifiers is a handy checklist for what to strip. And remember the EU line. Under GDPR, pseudonymized data still counts as personal data. So keep the real identities in a separate, access-controlled master, use a tool that never trains on your files, and you've kept your ethics board happy.

Last, export something your analysis software will actually open. A Word file, a DOCX, imports cleanly into the main coding packages, so one export covers whatever your team runs. Speaker labels intact. Timestamps intact.

Get those decisions right, capture, style, anonymization, format, and the transcript holds up when a reviewer asks how it was made. Because in qualitative work, how you made the transcript is part of the finding.

That's this episode of Talk Type. The full write up, with the links and sources, is in the show notes. Pepys transcribes any file or link, any length, pay once, and we never train on your audio. Your first sixty minutes are free at pepys dot co. Thanks for listening, and we'll see you next time.