# Pepys > Pay-as-you-go audio and video transcription with built-in AI. Drop a file or paste a link and get an accurate, timestamped, speaker-labeled transcript in minutes – plus AI summaries and chat. Credits never expire; no subscription. Pepys (operated by KMF Ventures LLC) converts audio and video to text. It offers automatic language detection across 99+ languages, speaker labels (diarization), word-level timestamps, and a native AI layer – summaries, chapters, action items, and a chat grounded in your transcript. Pricing is pay-as-you-go: one-time credit packs that never expire, about $1/hour with volume discounts on larger top-ups, and no subscription. New accounts get 60 minutes free, no card required. Pepys never trains on your audio or text. ## Core pages - [Home](https://pepys.co): Transcribe audio and video to text, pay as you go. - [Pricing](https://pepys.co/pricing): Pay-as-you-go credit packs, no subscription, credits never expire. About $1/hour, from $0.85/hr at volume; 60 free minutes on signup. - [Developers / API + MCP](https://pepys.co/developers): REST API and an MCP server for developers and AI agents – bearer API keys, direct uploads, transcription endpoints, signed webhooks, and podcast-feed ingestion, plus an MCP server (a hosted connector at https://pepys.co/api/mcp, or the npx pepys-mcp stdio package on npm) so agents like Claude and ChatGPT transcribe directly. Use Pepys as a programmatic transcription backend. - [MCP server](https://pepys.co/mcp): Pepys as an MCP (Model Context Protocol) server. Connect Claude, ChatGPT, or Cursor to the hosted connector at https://pepys.co/api/mcp (OAuth sign-in, no API key) and the agent transcribes audio, video, and whole podcast feeds on its own – speaker diarization, SRT/VTT export, paste-a-link, and transcript search. Also runnable locally via npx pepys-mcp with an API key. - [AI analysis](https://pepys.co/analyze): Six use-case AI frameworks built into transcription – short-form video, podcasts, meetings, lectures, interviews, and a general summary. - [Transcription by industry](https://pepys.co/transcribe): A dedicated transcription page per industry, with industry-specific samples, exports, and AI analysis. - [Free transcription tools](https://pepys.co/tools): Free format converters, platform-transcript tools, subtitle generators, and AI note tools. - [Pepys vs alternatives](https://pepys.co/alternatives): Honest, side-by-side comparisons with other transcription tools. - [How-to guides](https://pepys.co/how-to): Step-by-step transcription guides. - [Blog](https://pepys.co/blog): Original research and notes on transcription – what people actually struggle with, and the honest fixes. - [Templates](https://pepys.co/templates): Free, downloadable templates for the work around transcription – interview release forms, recording consent, podcast show notes. - [Best transcription software](https://pepys.co/best): Ranked roundups for specific needs – podcasters, qualitative research, no-training tools, and free tools. - [Glossary](https://pepys.co/glossary): Plain-language definitions of transcription terms. - [Security](https://pepys.co/security): How Pepys encrypts, stores, retains, and never trains on your data, plus named subprocessors. - [About](https://pepys.co/about): What Pepys is, who operates it (KMF Ventures LLC), and how it handles your data. - [Reviews](https://pepys.co/reviews): What customers say about Pepys. ## What you can do with Pepys - Transcribe uploaded audio and video files (common formats; long and large files supported). - Paste a link to transcribe (YouTube, podcasts and RSS, direct media URLs). - Get speaker labels, plus word- and segment-level timestamps. - Generate AI summaries, chapters, and action items; chat with your transcript for grounded answers. - Translate transcripts into other languages. - Export to TXT, Markdown, DOCX, PDF, SRT, VTT, and JSON (word-level JSON/SRT/VTT available on demand). - Integrate programmatically via the REST API (built for developers and AI agents): create API keys, upload files or submit URLs, poll transcriptions, receive signed webhooks, and auto-transcribe podcast feeds. See https://pepys.co/developers. - Connect Pepys as an MCP server so AI agents (Claude, ChatGPT, Cursor) transcribe on their own: a hosted remote connector at https://pepys.co/api/mcp (OAuth sign-in, no API key), or the npx pepys-mcp stdio server (bring your own API key) for local dev agents. Tools include transcribe, speaker diarization, SRT/VTT export, whole-podcast-feed batch, and transcript search. See https://pepys.co/developers. ## Pricing and privacy - Pricing model: one-time credits (1 credit = 1 minute) that never expire – no subscription, nothing to cancel. - Free tier: 60 minutes free on signup (one-time, never expires), no card required. - Privacy: your audio and text are never used to train AI models. ## Company and contact - Operated by KMF Ventures LLC. Questions, API access, or press: contact@pepys.co. - [Privacy Policy](https://pepys.co/privacy) - [Terms of Service](https://pepys.co/terms) ## Transcription by industry pages Each links to a page with an industry-specific sample, the AI analysis that fits it, and the relevant exports. - [Podcast transcription](https://pepys.co/transcribe/podcasters): To transcribe a podcast, upload the episode audio or paste its link and Pepys returns a speaker-labeled transcript in minutes – plus AI-generated show notes, key topics, and pull-quotes. It's pay-as-you-go with no subscription, and credits never expire. - [Interview transcription](https://pepys.co/transcribe/journalists): To transcribe an interview, upload the recording or paste its link and Pepys returns a speaker-labeled, fully searchable transcript in minutes, plus an AI recap that surfaces the themes, the most quotable lines, and a Q&A breakdown. It is pay-as-you-go with no subscription, and credits never expire. - [Newsroom transcription](https://pepys.co/transcribe/newsrooms): To transcribe newsroom audio, upload the interview, press conference, or meeting recording, or paste its link, and Pepys returns a speaker-labeled transcript in minutes, plus the key themes, verbatim on-the-record quotes, and a Q&A breakdown. It is pay-as-you-go with no subscription, and credits never expire. - [Video transcription](https://pepys.co/transcribe/videographers): To transcribe a video, upload the file or paste its link and Pepys returns a speaker-labeled, time-coded transcript in minutes – plus exportable SRT and VTT captions and a quick AI summary. It's pay-as-you-go with no subscription, and credits never expire. - [Film and documentary transcription](https://pepys.co/transcribe/filmmakers): To transcribe film and documentary footage, upload your interviews or dailies, or paste a link, and Pepys returns a speaker-labeled transcript in minutes, plus the standout quotes, recurring themes, and a clean Q&A breakdown for your paper edit. It's pay-as-you-go with no subscription, and credits never expire. - [Radio transcription](https://pepys.co/transcribe/radio-stations): To transcribe a radio broadcast, upload the recording or paste its link and Pepys returns a speaker-labeled transcript in minutes, plus an AI-drafted segment summary, topics, and pull-quotes you can post or archive. It's pay-as-you-go with no subscription, and credits never expire. - [Media monitoring transcription](https://pepys.co/transcribe/media-monitoring): To transcribe a media mention, upload the broadcast, radio, or podcast clip or paste its link and Pepys returns a speaker-labeled, searchable transcript in minutes, plus an AI summary, key points, and the verbatim quotes your report needs. It is pay-as-you-go with no subscription, and credits never expire. - [Transcription for accessibility](https://pepys.co/transcribe/accessibility): Transcription for accessibility means turning a video into a corrected, time-coded transcript you can export as caption files (SRT and VTT) plus a readable on-page transcript. Pepys returns a speaker-labeled draft in minutes that you fix and ship, so captions are accurate rather than auto-generated guesses. It's pay-as-you-go with no subscription, and credits never expire. - [Legal transcription](https://pepys.co/transcribe/lawyers): Upload a deposition, hearing, or recorded statement and Pepys returns a speaker-labeled transcript with word-level timestamps in minutes, then surfaces the themes, notable testimony, and a question-and-answer recap. It is pay-as-you-go with no subscription, credits never expire, and your audio is never used to train a model. - [Law firm transcription services](https://pepys.co/transcribe/legal-firms): Pepys transcribes depositions, hearings, client interviews, and recorded statements into speaker-labeled text in minutes, then surfaces the key themes, notable admissions, and a question-by-question recap of the testimony. It's pay-per-recording with no subscription, credits never expire, and your files are never used to train a model. - [Sales call transcription](https://pepys.co/transcribe/sales-teams): To transcribe a sales call, upload the recording or paste its link and Pepys returns a speaker-labeled transcript in minutes – plus an AI summary, the decisions made, owner-tagged action items, next steps, and the open questions to chase. It's pay-as-you-go with no subscription, and credits never expire. - [Webinar transcription](https://pepys.co/transcribe/marketers): To transcribe a webinar, upload the recording or paste the replay link and Pepys returns a speaker-labeled transcript in minutes – plus an AI summary, key takeaways, and quotable lines you can repurpose into a blog, recap email, and social posts. It's pay-as-you-go with no subscription, and credits never expire. - [Coaching call transcription](https://pepys.co/transcribe/coaches): To transcribe a coaching call, record the session and upload the audio or paste a link – Pepys returns a speaker-labeled transcript in minutes, plus the client's commitments, action items, and open threads to revisit next session. It's pay-as-you-go with no subscription, and credits never expire. - [Nonprofit transcription](https://pepys.co/transcribe/nonprofits): To transcribe a nonprofit field interview, upload the recording or paste its link and Pepys returns a speaker-labeled transcript in minutes, plus AI-pulled themes, notable quotes, and a Q&A recap ready for an impact report or grant narrative. It is pay-as-you-go with no subscription, and credits never expire. - [Lecture transcription](https://pepys.co/transcribe/students): To transcribe a lecture, record the class on your phone or upload the audio and Pepys returns a clean transcript in minutes, plus AI study notes: key concepts with definitions, an outline, takeaways, and practice questions. It's pay-as-you-go with no subscription, and credits never expire. - [Transcription for educators](https://pepys.co/transcribe/educators): To turn a class into notes, record the session and upload it (or paste a link) and Pepys returns a speaker-labeled transcript in minutes, plus AI-drafted key concepts, an outline, takeaways, and study questions pulled straight from the lecture. It's pay-as-you-go with no subscription, and credits never expire. - [Research interview transcription](https://pepys.co/transcribe/researchers): To transcribe a research interview, upload the audio or paste its link and Pepys returns a speaker-labeled, timestamped transcript in minutes – plus AI-surfaced themes, verbatim quotes, and a question-by-question breakdown you can drop straight into your coding and write-up. It's pay-as-you-go with no subscription, and credits never expire. - [Qualitative research transcription](https://pepys.co/transcribe/research-firms): To transcribe a research interview, upload the recording or paste its link and Pepys returns a speaker-labeled, timestamped transcript in minutes – plus AI-drafted themes, verbatim quotes, and a moderator/participant Q&A breakdown ready for coding. It's pay-as-you-go with no subscription, and credits never expire. ## Free transcription tools 188 free tools at [https://pepys.co/tools](https://pepys.co/tools), grouped by what they do: - File-format converters (40): [3ga to text](https://pepys.co/tools/3ga-to-text), [3gp to text](https://pepys.co/tools/3gp-to-text), [Aac to srt](https://pepys.co/tools/aac-to-srt), [Aac to text](https://pepys.co/tools/aac-to-text), and more - Platform transcripts (19): [Apple podcast transcript](https://pepys.co/tools/apple-podcast-transcript), [Chat with podcast](https://pepys.co/tools/chat-with-podcast), [Chat with tiktok](https://pepys.co/tools/chat-with-tiktok), [Chat with youtube video](https://pepys.co/tools/chat-with-youtube-video), and more - Subtitles and captions (18): [Interview transcript formatter](https://pepys.co/tools/interview-transcript-formatter), [Json to srt](https://pepys.co/tools/json-to-srt), [Podcast subtitle generator](https://pepys.co/tools/podcast-subtitle-generator), [Subtitle generator](https://pepys.co/tools/subtitle-generator), and more - By use case (17): [Action item extractor](https://pepys.co/tools/action-item-extractor), [Conference talk transcription](https://pepys.co/tools/conference-talk-transcription), [Google meet transcription](https://pepys.co/tools/google-meet-transcription), [Interview transcription](https://pepys.co/tools/interview-transcription), and more - AI notes and summaries (45): [Ai notes from audio](https://pepys.co/tools/ai-notes-from-audio), [Audio summarizer](https://pepys.co/tools/audio-summarizer), [Audio to blog post](https://pepys.co/tools/audio-to-blog-post), [Audiobook summarizer](https://pepys.co/tools/audiobook-summarizer), and more - By language (49): [Arabic audio to text](https://pepys.co/tools/arabic-audio-to-text), [Arabic subtitle generator](https://pepys.co/tools/arabic-subtitle-generator), [Chinese audio to text](https://pepys.co/tools/chinese-audio-to-text), [Chinese subtitle generator](https://pepys.co/tools/chinese-subtitle-generator), and more ## Pepys vs other transcription tools Honest, public-fact comparisons (the one place we name competitors): - [TurboScribe alternative](https://pepys.co/alternatives/turboscribe-alternative): Same fast, accurate transcripts – without the monthly subscription. Pay once for the minutes you use, and your credits never expire. - [Otter.ai alternative](https://pepys.co/alternatives/otter-alternative): Transcribe files you already have without a per-seat subscription or a monthly minute cap that resets. Buy credits once – they never expire. - [NotebookLM alternative](https://pepys.co/alternatives/notebooklm-alternative): NotebookLM helps you understand your sources. Pepys turns audio and video into an accurate, speaker-labeled, timestamped transcript you can edit, own, and export – with AI chat on top, so you give up nothing. - [Rev alternative](https://pepys.co/alternatives/rev-alternative): Accurate AI transcripts without a subscription or a per-minute meter. Buy credits once, spend them whenever, and they never expire. - [Descript alternative](https://pepys.co/alternatives/descript-alternative): Descript is a brilliant editor if you cut video and podcasts from the text. If you mainly need the accurate, speaker-labeled transcript – and the exports – Pepys gives you that for a one-time credit purchase, no subscription. - [Cockatoo alternative](https://pepys.co/alternatives/cockatoo-alternative): Same fast, accurate transcripts – without the monthly subscription. Pay once for the minutes you use, and your credits never expire. - [Notta alternative](https://pepys.co/alternatives/notta-alternative): Accurate, speaker-labeled transcripts without a monthly seat. Pay once for the minutes you use, paste a link or upload a file, and export everywhere – your credits never expire. - [Happy Scribe alternative](https://pepys.co/alternatives/happy-scribe-alternative): Accurate transcripts, speaker labels, and 99+ languages – without a monthly plan or minutes that reset. Buy credits once, and they never expire. - [Sonix alternative](https://pepys.co/alternatives/sonix-alternative): Accurate, speaker-labeled transcripts with AI built in – without a subscription. Pay once for the minutes you use, and your credits never expire. - [Temi alternative](https://pepys.co/alternatives/temi-alternative): Same pay-per-file simplicity, at about a fifteenth of Temi's per-minute rate – plus 99+ languages, AI summaries, and credits that never expire. - [MacWhisper alternative](https://pepys.co/alternatives/macwhisper-alternative): MacWhisper transcribes locally on your Mac. Pepys is the web alternative when you're not on a Mac – pay once for the minutes you use, and your credits never expire. - [Trint alternative](https://pepys.co/alternatives/trint-alternative): Same speaker-labeled, timestamped transcripts, without the monthly subscription. Pay once for the minutes you use, and your credits never expire. - [OpenAI Whisper alternative](https://pepys.co/alternatives/openai-whisper-alternative): Get diarized, timestamped transcripts with summaries and exports – no Python, PyTorch or GPU to manage. Upload a file or paste a link, and pay once for credits that never expire. - [Microsoft Word alternative](https://pepys.co/alternatives/microsoft-word-transcribe-alternative): The same speaker-labeled transcripts, without a Microsoft 365 subscription or a 300-minute monthly cap. Pay once for the minutes you use, and your credits never expire. - [Good Tape alternative](https://pepys.co/alternatives/good-tape-alternative): Same accurate transcripts, without the monthly subscription. Pay once for the minutes you use, and your credits never expire. - [Transkriptor alternative](https://pepys.co/alternatives/transkriptor-alternative): Speaker labels and 99+ languages, auto-detected, without the monthly plan, the per-seat fees, or minutes that expire each cycle. Buy credits once, spend them whenever, and they never expire. ## How-to guides - [How to transcribe an interview](https://pepys.co/how-to/how-to-transcribe-an-interview): To transcribe an interview, start with a clean recording, then upload it to a transcription tool to get a speaker-labeled, timestamped draft in minutes. Read the draft against the audio, fix names, jargon, and the quotes you'll actually publish, and keep the timestamps so you can re-check any line. Doing the first pass by AI and the cleanup by hand is far faster than typing from scratch – and more accurate where it counts. - [How to transcribe a podcast](https://pepys.co/how-to/how-to-transcribe-a-podcast): To transcribe a podcast, upload the episode audio (or paste the link) to a transcription tool and get a clean, speaker-labeled, timestamped transcript in minutes. Then repurpose it: pull show notes and chapters from the timestamps, lift quotes for social, and publish the transcript on the episode page so search engines and listeners can find it. One episode becomes a transcript, notes, clips, and a searchable page from a single source file. - [Legal transcription](https://pepys.co/how-to/legal-transcription): Legal transcription is the word-for-word record of a legal proceeding. In federal court, only a transcript certified by an official reporter is the official record (28 U.S.C. 753(b)) – not an AI draft. Use automated transcription for fast, searchable working copies: discovery review, deposition summaries, quote-pulling – then route the filing-grade record to a certified reporter or transcriber. - [How to transcribe an interview for a dissertation](https://pepys.co/how-to/how-to-transcribe-an-interview-for-dissertation): Transcribe your dissertation interviews in two passes: get a speaker-labeled AI draft, then correct it against the audio to a convention your methodology can defend – naturalized (Jeffersonian) verbatim for conversation or discourse analysis, denaturalized clean verbatim for thematic coding. Structure each turn by speaker so NVivo, ATLAS.ti, or MAXQDA can import it, keep timestamps for re-checking, and justify the choice in your methods chapter. - [Qualitative research transcription](https://pepys.co/how-to/qualitative-research-transcription): Qualitative research transcription converts recorded interviews and focus groups into text for coding. Treat it as an interpretive first step of analysis, not clerical work. Choose a naturalized or denaturalized style to fit your research question. Document your conventions so transcripts stay consistent across a team. Anonymize direct identifiers to meet ethics and data-protection rules. Then export a speaker-labeled file your CAQDAS software can import. - [How to transcribe a zoom meeting](https://pepys.co/how-to/how-to-transcribe-a-zoom-meeting): To transcribe a Zoom meeting, record it first, then get a speaker-labeled, timestamped draft. Zoom's built-in cloud transcript needs a paid plan, so on a free Basic account you'll record the call yourself. For the cleanest labels, save a separate audio file per participant, upload those tracks to a transcription tool, then clean only the quotes you'll publish and keep the timestamps to re-check any line. - [Transcribe a focus group](https://pepys.co/how-to/how-to-transcribe-a-focus-group): To transcribe a focus group, record so each voice stays separable, then upload the file for a speaker-labeled, timestamped draft. Expect to fix more speaker turns by hand than in a one-on-one interview, because overlap and 6 to 10 voices push automatic labeling to its limit. Clean only the passages you'll code or quote, and anonymize names in the transcript itself. - [How to transcribe a lecture](https://pepys.co/how-to/how-to-transcribe-a-lecture): To transcribe a lecture, start with the cleanest recording you can get, then upload it for a speaker-labeled, timestamped draft in minutes instead of hours of typing. Read the draft against the audio to fix technical terms, names, and numbers, then keep the timestamps so every quote you cite traces back to the exact moment in the talk. - [How to transcribe a sermon](https://pepys.co/how-to/how-to-transcribe-a-sermon): To transcribe a sermon, start with the closest, cleanest recording you can get, then upload it to a transcription tool for a timestamped draft in minutes. Read the draft against the audio and fix what AI gets wrong in preaching: proper names, Scripture references, and theological terms. Then clean only the passages you'll publish or turn into a study guide. Keep the timestamps so every quote is re-checkable. - [Transcribe a sales call](https://pepys.co/how-to/how-to-transcribe-a-sales-call): To transcribe a sales call, get consent to record, then upload the audio to a transcription tool for a speaker-labeled, timestamped draft in minutes. Read it against the recording to fix names, prices, and product terms, then pull the objections, commitments, and next steps into your CRM. Keep the timestamps so every quoted commitment is re-checkable later. - [Transcribe a webinar](https://pepys.co/how-to/how-to-transcribe-a-webinar): To transcribe a webinar, export the recording's audio, upload it to a transcription tool, and get a speaker-labeled, timestamped draft in minutes. Webinar recordings usually give you one mixed track of the host and panelists, so plan to correct some speaker turns by hand. Clean only the quotes you'll publish, keep the timestamps, and remember the recording is copyrighted. - [How to turn a podcast into a blog post](https://pepys.co/how-to/how-to-turn-a-podcast-into-a-blog-post): To turn a podcast into a blog post, transcribe the episode first, then shape the transcript into a structured article rather than publishing it raw. Pull the strongest exact quotes, write an original narrative around them, add headings and links, and attribute every quoted line. The transcript is raw material; your editing and original framing are what make the post worth reading and worth indexing. - [How to make an srt file](https://pepys.co/how-to/how-to-make-an-srt-file): To make an SRT file, write plain text in four-part blocks: a sequence number, a timecode line using a comma before the milliseconds (00:00:01,000 --> 00:00:04,000), the caption text, then a blank line before the next cue. Save it with a .srt extension in UTF-8 so accented characters render. Faster still, upload your recording, export a ready-made SRT, and fix only the line breaks and timing. - [How to add subtitles to a video](https://pepys.co/how-to/how-to-add-subtitles-to-a-video): To add subtitles to a video, transcribe the audio into a timed text file – SRT or WebVTT – then either load it as a sidecar track your player toggles on, or burn it into the picture. Start with an AI transcript, correct the words and timing by hand, keep lines under about 37 characters and two per cue, then export SRT or VTT. - [How to transcribe multiple speakers](https://pepys.co/how-to/how-to-transcribe-multiple-speakers): To transcribe multiple speakers, record each person on a separate channel where you can, then upload the audio for an AI first pass that labels every speaker turn and adds timestamps. Read the draft against the audio and fix the overlaps by hand, because overlapping speech is where automatic speaker labeling fails most. Keep the labels through export so who-said-what survives into your notes. - [How to improve transcription accuracy](https://pepys.co/how-to/how-to-improve-transcription-accuracy): To improve transcription accuracy, fix the audio before you touch the software: mic each speaker close, cut background noise, and record per-channel so voices stay separable. Then run an AI first pass and verify by hand, because the errors that matter – names, numbers, and speaker labels – cluster in a small slice of the transcript. Read those lines against the recording and fix them. - [How to translate a transcript](https://pepys.co/how-to/how-to-translate-a-transcript): To translate a transcript, first transcribe it in the source language, then translate segment by segment so timestamps and speaker labels survive. Machine translation reads fluently but still makes meaning-level errors, so have a fluent speaker check the passages you'll quote or publish. For immigration, court, or research use, a certified or back-translated version is often required. - [Verbatim vs clean verbatim](https://pepys.co/how-to/verbatim-vs-clean-verbatim): Verbatim transcription captures every utterance – fillers, false starts, stutters, and repetitions all stay in. Clean verbatim keeps the speaker's exact words and meaning but removes that noise, so quotes read clearly. Use strict verbatim when how something was said matters, as in legal or discourse analysis; use clean verbatim for readable, accurate quotes in journalism and most research. - [How accurate is ai transcription](https://pepys.co/how-to/how-accurate-is-ai-transcription): How accurate AI transcription is depends on your audio, not the marketing number. Near-99% figures come from clean, read-aloud benchmarks. On real conversation, the best systems and professional human transcribers both land around 5–6% word error rate, roughly one word in twenty. Accents, crosstalk, and noise push errors higher, and models can occasionally hallucinate phrases no one actually said. - [Ai vs human transcription](https://pepys.co/how-to/ai-vs-human-transcription): Use AI as the default for interviews, lectures, and research: strong systems match human transcribers near 5.9% word-error rate at a fraction of the time and cost. Choose a certified human for legal, medical, or compliance records that could be challenged. The hybrid most professionals run: AI first pass, then human-verify every quote you publish. - [What is speaker diarization](https://pepys.co/how-to/what-is-speaker-diarization): Speaker diarization is the task of labeling a recording by speaker – 'who spoke when' – without knowing anyone's real identity or even how many people are talking. It splits audio into speech segments, groups them by voice, and tags each with a generic label like Speaker 1. It runs separately from speech recognition, which turns speech into words. - [How to anonymize interview transcript](https://pepys.co/how-to/how-to-anonymize-a-transcript): To anonymize an interview transcript, work on a copy, not the master. First strip direct identifiers – names, addresses, phone numbers, employers, ID numbers. Then generalize the quasi-identifiers that still single someone out: a city becomes a region, an exact date becomes a month, a rare job title becomes a category. Removing the name alone isn't anonymization – combined details re-identify people. - [Is ai transcription confidential](https://pepys.co/how-to/confidential-transcription): It depends on the vendor, not the technology. AI transcription is confidential when the provider encrypts your files in transit and at rest, deletes them on a clear schedule, and contractually won't train models on your content. None of that is automatic, so read the privacy policy for retention, subprocessors, and a written no-training term before you upload anything sensitive. - [Is it legal to record a meeting](https://pepys.co/how-to/is-it-legal-to-record-and-transcribe-a-meeting): Recording a meeting is legal in most of the US under federal one-party-consent law, meaning one participant's agreement is enough. But roughly 11 states require every party to consent, and recording someone in those states without permission risks criminal and civil penalties. AI notetakers that auto-join calls have drawn wiretap and biometric-privacy lawsuits; recording audio you captured yourself, with everyone's knowledge, is the clean path. - [Meeting notes without a bot](https://pepys.co/how-to/meeting-notes-without-a-bot): To get meeting notes without a bot, record the call with your platform's own recording, then upload that file to a transcription tool afterward for a speaker-labeled transcript and summary. Nothing joins the meeting live, so no third-party bot streams your audio mid-call. You control who gets captured, and you secure consent before the substance starts. - [Hipaa compliant transcription](https://pepys.co/how-to/hipaa-ai-transcription): HIPAA-compliant transcription only matters if HIPAA binds you. The Privacy Rule covers health plans, clearinghouses, most healthcare providers, and the vendors acting on their behalf. If that's you and a tool will handle patient data, you need a signed business associate agreement first. De-identify the data and it stops being PHI, so the requirement falls away. - [How to transcribe a therapy session](https://pepys.co/how-to/how-to-transcribe-a-therapy-session): To transcribe a therapy session, first get every participant's recorded consent, which professional ethics codes require. Upload the recording to a transcription tool that doesn't train on your files and auto-deletes them, and get a two-speaker, timestamped draft in minutes. Then de-identify the transcript, stripping names and other identifiers, before it goes into supervision notes, research, or the record. - [How much does transcription cost per minute](https://pepys.co/how-to/how-much-does-transcription-cost): Transcription typically costs $0.70 to $2.00 per audio minute for human services, and up to $3.50+ for legal or medical. Pure AI runs roughly $0.05 to $0.25 per minute, and automated pay-as-you-go tools about $0.17. Subscriptions and pay-once credit packs lower the effective rate only if you actually use the minutes. - [Cheapest way to transcribe audio](https://pepys.co/how-to/cheapest-way-to-transcribe-audio): The cheapest way to transcribe audio in raw dollars is open-source Whisper, which is free under an MIT license but needs your own GPU, setup, and time, and can't label speakers on its own. Free tiers cap you at short clips. For real volume, the cheapest reliable option is pay-once AI at roughly $1 an hour, with no subscription and no expiry. - [How much does it cost to transcribe an interview](https://pepys.co/how-to/how-much-to-transcribe-an-interview): Transcribing a one-hour interview costs roughly $48 to $119 with a professional human service (about $0.80 to $1.99 per audio minute), around $10 an hour with pay-as-you-go AI, or close to nothing but your time if you type it yourself, which can take up to six hours. Pay-once AI sits near $1 an hour. - [Is transcription software worth it](https://pepys.co/how-to/is-transcription-software-worth-it): Transcription software is worth it whenever you record regularly or work with files longer than a few minutes. Typing a transcript by hand takes up to six hours per audio hour, so even one interview a month justifies a tool. It stops being worth it only for rare, short, clean clips you could type in a single sitting. - [Free vs paid transcription](https://pepys.co/how-to/free-vs-paid-transcription): Free transcription covers real work but with hard limits: short per-file caps, no speaker labels or timestamps, and terms that may let the tool train on your audio. Paid buys a diarized, timestamped, exportable file and a no-training guarantee. For irregular, project-based volume, a pay-once tool with free starter minutes sits between the two. - [Transcription without a subscription](https://pepys.co/how-to/transcription-without-a-subscription): You have three subscription-free routes. Per-minute human services like Rev bill by audio minute ($1.99/min) with no monthly plan. One-time desktop apps like MacWhisper are a one-time buy (€64) and run locally on a Mac. Pay-once credits let you buy transcription minutes that never expire, so you pay nothing between projects. - [Can chatgpt transcribe audio](https://pepys.co/how-to/can-chatgpt-transcribe-audio): Partly. ChatGPT's consumer app can't take a direct audio upload – OpenAI's own docs list audio as an unsupported file type – so you can't drop in an MP3 and get a transcript. Voice Mode transcribes your live spoken conversation, and the API can transcribe short clips, but neither hands you a long, speaker-labeled, timestamped transcript file. For that, use a dedicated transcription tool. - [Apple voice memos transcription](https://pepys.co/how-to/apple-voice-memos-transcription): Apple's Voice Memos transcribes recordings to text on the device itself – a Mac on macOS 15 with Apple silicon, or an iPhone 12 or later on iOS 18 – but it adds no speaker labels, no timestamped export file, and keeps the transcript tied to the recording. For a diarized, timestamped, exportable transcript, export the recording as an .m4a and upload it to a transcription tool. - [Export transcript from notebooklm](https://pepys.co/how-to/how-to-export-transcript-from-notebooklm): NotebookLM has no one-click transcript export – no downloadable file with timestamps or speaker labels. To get text out, chat with your notebook, save the answer as a note, then export the note to a Google Doc. Treat that text as a paraphrase, not a verbatim record. For a transcript you can cite, with timestamps and speaker labels, upload the audio to a dedicated transcription tool instead. - [Get zoom transcript without being host](https://pepys.co/how-to/how-to-get-a-zoom-transcript-without-being-host): As a Zoom participant, you can't start a recording or pull the transcript yourself – both are host-controlled. You have two honest routes: ask the host to share the cloud recording and its transcript via a link, or record your own audio with everyone's consent and upload it to a transcription tool for a speaker-labeled draft. Covert recording isn't an option. - [Transcribe plaud recording](https://pepys.co/how-to/how-to-transcribe-ai-voice-recorder-audio): To transcribe a Plaud recording, open the Plaud app and export the original audio as an MP3 or WAV file. Then upload that file to a transcription tool for a speaker-labeled, timestamped transcript you can edit and export. This gets you the full verbatim text, not just the app's AI summary, and it sidesteps the monthly minute cap on the device's built-in transcription. - [How to write meeting minutes from a recording](https://pepys.co/how-to/how-to-write-meeting-minutes-from-a-recording): To write meeting minutes from a recording, record the meeting yourself, then upload the saved file for a speaker-labeled transcript – no bot joins the call. From the transcript, draft structured minutes: attendees, agenda, and each decision with its action items and owners. Under Robert's Rules, minutes record what was done, not what was said. The draft becomes official once members approve it, usually at the next meeting. - [What is transcription](https://pepys.co/how-to/what-is-transcription): Transcription is the process of turning spoken words into written text, and the written record it produces. A transcript can be strict verbatim, keeping every "um" and false start, or cleaned up for readability. The work is done by human typists, by automatic speech recognition, or by an AI draft that a person then corrects. - [What is speech to text](https://pepys.co/how-to/what-is-speech-to-text): Speech-to-text (STT), also called automatic speech recognition (ASR), is the automatic conversion of spoken audio into written words. Modern systems are trained neural models that learn the audio-to-text mapping directly. STT produces the raw text; a transcript is the finished, formatted artifact you edit, label by speaker, and cite. Accuracy depends heavily on your recording quality. - [Word error rate](https://pepys.co/how-to/word-error-rate-explained): Word error rate (WER) measures speech-to-text accuracy as the share of words a system gets wrong. You align its output against a correct reference transcript, then divide the substitutions, deletions, and insertions by the number of reference words. A 10% WER means one word in ten is wrong. Lower is better, and the score depends far more on the audio than on the tool. - [Timestamped transcript](https://pepys.co/how-to/what-is-a-timestamped-transcript): A timestamped transcript is an ordinary transcript with one addition: every segment or word is tagged with the exact time it occurs in the audio, written as hours:minutes:seconds – for example, 00:02:17,440. Those time codes let you jump back to any line, cite the precise moment of a quote, and turn plain text into synchronized captions. - [Best audio format for transcription](https://pepys.co/how-to/best-audio-format-for-transcription): For transcription, record lossless WAV or FLAC when you can, at 16 kHz or higher, 16-bit, with one channel per speaker. Modern engines like Whisper resample everything to 16 kHz anyway, so a clean 128 kbps MP3 transcribes almost as well. What actually moves accuracy is mic distance and low background noise, not the file format. - [Transcribe accented english](https://pepys.co/how-to/transcribe-accented-english): Accent and dialect measurably change transcription accuracy. In one benchmark the best model scored 19.7% word error rate on accented English versus 2.7% on US clean speech (Sanabria et al., 2023), and commercial systems have shown roughly double the error rate for some dialects (Koenecke et al., 2020). The gap comes from pronunciation and prosody, and clean audio plus human review of key quotes narrows it. - [Transcribe long audio files](https://pepys.co/how-to/how-to-transcribe-long-audio): To transcribe long audio, split the recording into chunks under your tool's caps, transcribe each piece, then re-offset the timestamps and stitch them into one file. OpenAI's API caps at 25 MB per file, and gpt-4o-transcribe rejects anything past ~25 minutes. The easiest path is a tool that chunks, transcribes, and re-joins automatically, so you upload a two-hour recording and get back a single transcript. - [Oral history transcription](https://pepys.co/how-to/how-to-transcribe-oral-history): To transcribe an oral history, start with an archival recording (WAV or BWF, at least 48 kHz, 24-bit), then get a speaker-labeled draft and audit-edit it line by line against the audio. Label the two voices narrator and interviewer, bracket editorial insertions and inaudible spots, and let the narrator review and correct the transcript before it's deposited with the recording. - [Import transcript into nvivo](https://pepys.co/how-to/how-to-import-a-transcript-into-nvivo): To import a transcript into NVivo, export it as a .docx with each speaker's name at the start of its own line, then choose Import, Documents, and select the file. NVivo also accepts .doc, .rtf, .txt, and .pdf. After importing, run Autocode to create a case for each speaker. Attaching a transcript to an audio or video file is a separate workflow that maps timespan, speaker, and content columns. - [How to analyze interview transcripts](https://pepys.co/how-to/how-to-analyze-interview-transcripts): To analyze interview transcripts, start from a clean, speaker-labeled transcript and code it: label the meaningful segments, then group those codes into themes. Braun and Clarke's six phases – familiarize, generate codes, search, review, define, and report – give the standard route. Evidence each theme with timestamped quotes drawn fairly across the whole data set, not just the striking lines. - [How to transcribe a phone interview](https://pepys.co/how-to/how-to-transcribe-a-phone-interview): To transcribe a phone interview, first confirm you can legally record the call – U.S. federal law allows one-party consent, but about a dozen states require everyone to agree. Capture the cleanest audio the line allows, ideally each speaker on a separate track. Then run an AI first pass for a speaker-labeled, timestamped draft and hand-check the quotes you'll publish. - [Clinical research interview transcription](https://pepys.co/how-to/how-to-transcribe-a-clinical-research-interview): To transcribe a clinical research interview, first confirm IRB-approved consent to record. Keep identifiable recordings out of any tool not built for PHI. For de-identified research use, get a speaker-labeled, timestamped draft, correct it verbatim against the audio. Then strip the 18 HIPAA Safe Harbor identifiers (voice prints included), and delete the raw audio under access control. - [How to cite an interview in apa](https://pepys.co/how-to/how-to-cite-an-interview): An interview goes in the reference list only if a reader can retrieve it. An interview you conducted that isn't recoverable is a personal communication: cite it in the text only. A published or recorded interview (podcast, article, broadcast) is cited by the format of its source, with the interviewee as author. APA, MLA, and Chicago each word this slightly differently. - [How to transcribe user interviews](https://pepys.co/how-to/how-to-transcribe-a-user-interview): To transcribe a user interview, record each speaker on a separate track, then upload the file to get a speaker-labeled, timestamped draft in minutes instead of hours of typing. Clean the moderator and participant labels, tag the transcript for themes, and export to DOCX or JSON so quotes and highlight clips drop into your research repository, each tied to the exact second it was said. - [Types of transcription](https://pepys.co/how-to/types-of-transcription): Transcription types split along two axes. Style covers how much you keep: strict verbatim records every filler and false start, while clean, edited, and intelligent verbatim progressively tidy the words for readability. Method covers who does the work: human, AI, or a hybrid first-pass-plus-cleanup. Phonetic transcription is separate, using the International Phonetic Alphabet (IPA) to represent speech sounds rather than words. Match the type to your use case. - [Transcription vs translation](https://pepys.co/how-to/transcription-vs-translation): Transcription turns speech into written text in the same language; translation converts text from one language into another. They run in sequence, not in place of each other: you transcribe a recording first to get source-language text, then translate that text if your audience needs another language. Transcription errors are mishearings; translation errors are meaning-level. Captions are same-language; subtitles are translated. - [Open vs closed captions](https://pepys.co/how-to/open-vs-closed-captions): Closed captions can be switched on or off by the viewer and travel in the signal or a separate file such as SRT, WebVTT, or CEA-608/708; open captions are burned permanently into the video picture and can't be turned off. Use open captions where you can't trust the player to render a track, and closed captions where you want viewer control and machine-readable text. - [Captions vs subtitles](https://pepys.co/how-to/transcription-vs-captions-vs-subtitles): Captions and subtitles are both time-synced text on video, but they solve different problems. Captions transcribe speech plus non-speech sound (speaker IDs, sound effects, music) for viewers who can't hear the audio. Subtitles translate only the dialogue for viewers who can hear but don't know the language. A transcript, by contrast, is a standalone document with no timing. - [Srt vs vtt](https://pepys.co/how-to/srt-vs-vtt): SRT and WebVTT both pair caption text with timecodes, but they differ in a few ways. SRT separates milliseconds with a comma (00:00:01,000); WebVTT uses a full stop (00:00:01.000). WebVTT files open with a WEBVTT header and support on-screen positioning and CSS styling. The HTML5 element reads WebVTT, not SRT. Use SRT for broad uploads, WebVTT for web players. - [How to transcribe a coaching session](https://pepys.co/how-to/how-to-transcribe-a-coaching-session): To transcribe a coaching session, get the client's clear consent to record, then upload the audio to a transcription tool for a speaker-labeled, timestamped draft in minutes. Read it against the recording to fix names and terms, pull out commitments and action items, and store it confidentially. Doing the first pass by AI and the notes by hand beats typing from scratch. - [How to transcribe genealogy recordings](https://pepys.co/how-to/how-to-transcribe-genealogy-recordings): Start by digitizing the tape: magnetic media degrades and its playback gear is going obsolete, so transfer to uncompressed WAV first. Run an AI first pass for a timestamped draft, then correct names and places against the audio by ear. Treat the recording as a genealogical source, cite it, and store three backup copies. - [Transcription mcp server](https://pepys.co/how-to/how-to-add-transcription-to-your-ai-agent-mcp): To add transcription to an AI agent, wrap a speech-to-text API as an MCP tool the agent can call. On the call, submit the audio and return a job ID immediately, then let the agent poll for status or wait on a webhook. Return diarized, timestamped segments rather than plain text, so the agent can say who spoke, when, and cite the exact moment. - [Transcribe audio for rag](https://pepys.co/how-to/how-to-transcribe-audio-for-rag): To transcribe audio for RAG, produce an accurate, speaker-labeled, timestamped transcript, then export it as JSON so each cue keeps its start, end, speaker, and text. Chunk on speaker turns at roughly 512 tokens with 25% overlap, attach that metadata to every chunk, and embed them into a vector store. The metadata is what lets each generated answer cite an exact source. - [How to transcribe a city council meeting](https://pepys.co/how-to/how-to-transcribe-a-city-council-meeting): To transcribe a city council meeting, record separated audio, run an AI first pass for a speaker-labeled, timestamped draft, then relabel each voice with the member's real name from the roll call. Verify the lines you'll quote against the audio, since automatic transcripts can hallucinate. Draft minutes from the actions taken, and post an accessible transcript or captions. - [How to add captions to a church service](https://pepys.co/how-to/how-to-add-captions-to-a-church-service): To caption a church service, record a clean feed from the soundboard, then run that audio through a transcription tool. Export the timestamped transcript as an SRT or VTT file. Correct the names and Scripture references by hand, then attach the caption track to your video or livestream. Caption live for a stream, or after the fact for the archived recording. - [Caption accuracy standards](https://pepys.co/how-to/caption-accuracy-standards): Accurate captions match the spoken words in the order spoken, keep proper names right, and convey non-speech information like speaker identity, music, and sound effects. In the US, the FCC judges captions on four standards – accuracy, synchronicity, completeness, and placement – while WCAG and Section 508 set the accessibility baseline. There's no single legal accuracy percentage; errorless captions are the stated goal. - [How to make a video ADA compliant](https://pepys.co/how-to/how-to-make-a-video-ada-compliant): To make a prerecorded video ADA compliant, add the three accessibility building blocks WCAG 2.1 AA calls for: synchronized captions for the audio (Success Criterion 1.2.2), a full text transcript as the media alternative (1.2.3), and audio description of the on-screen visuals (1.2.5). Start from an accurate transcript, time it into a caption file, then record the description. - [Ada title ii transcript requirements](https://pepys.co/how-to/ada-title-ii-audio-transcript-requirements): ADA Title II adopts WCAG 2.1 Level AA as the accessibility standard for state and local government web content and apps. In practice, prerecorded audio needs a text transcript, and video needs synchronized captions. Public entities with a total population of 50,000 or more must comply by April 26, 2027; smaller entities and special districts by April 26, 2028. - [European accessibility act transcript requirements](https://pepys.co/how-to/eaa-compliant-transcripts-captions): The European Accessibility Act (Directive (EU) 2019/882) applies from 28 June 2025 to a defined list of products and services, not to every website. For in-scope audio and video, that means captions and audio description for synchronised media, and a transcript or text alternative for audio-only content, matched to EN 301 549 and WCAG 2.1. - [Transcribe audio over 25mb](https://pepys.co/how-to/how-to-transcribe-audio-over-25mb): OpenAI's Whisper API hard-caps uploads at 25 MB. Fix it three ways: re-encode to 16 kHz mono MP3 (ffmpeg -i in.mp4 -vn -ac 1 -ar 16000 -b:a 64k out.mp3), which drops most files under the cap; split the audio on silence and stitch the text with re-offset timestamps; or hand it to a hosted transcriber that chunks server-side. - [How to make podcast show notes](https://pepys.co/how-to/how-to-make-podcast-show-notes): Good show notes have five parts: a 2 to 3 sentence summary, chaptered timestamps, 3 to 5 key takeaways, a guest bio with links, and 2 to 3 pull quotes. The fastest path is to run a timestamped transcript through a language model with one fixed prompt, or a tool that generates them in one click. The catch – the timecodes are only right if the transcript carries timestamps. - [How to transcribe messy audio](https://pepys.co/how-to/how-to-transcribe-messy-audio): Most errors on messy audio are baked in before you hit upload, so fix it upstream first. Record each speaker on a separate track, place mics close, and denoise the file. Then run a current large model like Whisper large-v3, add speaker labels for crosstalk, and do a human correction pass. No tool cleanly untangles heavy overlap, so capture and correction do the real work. ## Blog - [How Many Hours of Audio Does a Qualitative Study Produce?](https://pepys.co/blog/how-many-hours-of-audio-in-a-qualitative-study): A typical qualitative study yields ~9-17 hours of audio: 9-17 interviews to saturation at roughly an hour each. Focus-group studies run about 4-12 hours. - [How Long Is a Typical Research Interview? Benchmarks by Type](https://pepys.co/blog/how-long-is-a-research-interview): Research interview length runs ~20 minutes to several hours by type: surveys ~20 min, elite/expert 20–45 min, focus groups 60–90 min, in-depth 30 min+. - [How Long Is the Average Deposition (and How Many Transcript Pages)?](https://pepys.co/blog/how-long-is-the-average-deposition): A federal deposition is capped at 1 day of 7 hours (FRCP 30(d)(1)), about 420 transcript pages at most at 1 page per minute. No true average exists. - [How Much Audio Does a Podcast Produce in a Year?](https://pepys.co/blog/how-much-audio-do-podcasters-produce): A weekly podcast makes about 17–35 hours of audio a year: 52 episodes at the most common 20–40 min length (Buzzsprout, 2026). At 30 min, roughly 26 hours. - [How Long to Transcribe a Research Study by Hand?](https://pepys.co/blog/how-long-does-it-take-to-transcribe-a-research-study): Transcribing a full qualitative study by hand takes 120-210 hours: about 3-5 working weeks for a 30-interview study, at 4-7 hours per audio hour. - [How Long Is the Average Recording? Podcasts, Meetings, Interviews and More](https://pepys.co/blog/how-long-is-the-average-recording): Most recordings run 30–60 minutes: podcasts median 39 min, meetings 51.9 min, sermons 37 min, research interviews 64 min. What published data shows, by type. - [Subtitle Reading Speed Standards: What Netflix, the BBC, and WCAG Actually Require](https://pepys.co/blog/subtitle-reading-speed-standards): Caption style guides cap adult subtitle reading speed near 160-180 words per minute (about 15-20 characters per second): Netflix 20 CPS, BBC 160-180 wpm. - [How Many Filler Words Per Minute Is Normal? The Research on Um and Uh](https://pepys.co/blog/how-many-filler-words-per-minute): Published research puts spontaneous speech at ~6 disfluencies per 100 words – about 3 to 4 um/uh a minute at normal talking pace, one every 15 to 20 seconds. - [Reading Speed vs Listening Speed: You Read Faster Than Anyone Can Talk](https://pepys.co/blog/reading-speed-vs-listening-speed): Silent reading runs 238 wpm (non-fiction); audiobooks are spoken at 140-180 wpm. You read faster than anyone talks – here's the research on why a transcript wins. - [Transcription Market Size 2026: The Numbers, By Scope](https://pepys.co/blog/transcription-market-size-2026): The speech-to-text API market was ~$3.8B in 2024, heading to $8.6B by 2030 (14.4% CAGR) while traditional transcription grows ~5.2%. That gap is the AI shift. - [Does Audio Quality Affect Transcription Accuracy? What the Research Shows](https://pepys.co/blog/does-audio-quality-affect-transcription-accuracy): Audio quality sets the ceiling on accuracy: as signal-to-noise ratio falls 20 dB to 0 dB, measured word error rate climbs from ~3.5% to over 70%. The data. - [How Big Is an Hour of Audio? File Sizes for Every Format](https://pepys.co/blog/how-big-is-an-hour-of-audio): One hour of audio is about 58 MB as a 128 kbps MP3 and roughly 600 MB as an uncompressed CD-quality WAV. Here's the size for every format, and the 25 MB wall. - [Transcription in Academic Research: The Hours and Dollars Nobody Budgets For](https://pepys.co/blog/transcription-in-academic-research): A median PhD qualitative study is 28 interviews – about 20 to 40 hours of audio, and 4 to 7 hours of typing per hour to transcribe by hand. The real math. - [Transcription Accuracy by Domain: Why the Same AI Breaks on Jargon](https://pepys.co/blog/transcription-accuracy-by-domain): AI transcription hits 2.7% word error rate on clean speech but 4-51% worse on medical terms and over 50% in real clinical audio. Accuracy varies by domain. - [The Cost of Not Transcribing Meetings (in Hours, Dollars, and Forgotten Detail)](https://pepys.co/blog/the-cost-of-not-transcribing-meetings): The average employee spends ~18 hours a week in meetings and calls a third unnecessary; here's what not keeping a record actually costs, in hours and dollars. - [Dictation vs Typing Speed: Why Speaking Is Close to 3x Faster](https://pepys.co/blog/dictation-vs-typing-speed): Speaking runs about 150 words a minute; typing averages 40 to 52. A Stanford study clocked speech at 153 WPM vs 52 on a phone keyboard – close to 3x faster. - [AI Transcription Limits in 2026: The Real Numbers for Every Major Tool](https://pepys.co/blog/ai-transcription-limits-2026): ChatGPT caps audio at 25 MB, Claude's app refuses it, Gemini free allows 10 minutes per prompt. The real 2026 transcription limits of every major AI tool, verified. - [AI Transcription Accuracy in 2026: The Benchmark Numbers, Read Honestly](https://pepys.co/blog/ai-transcription-accuracy-benchmark): Published benchmarks show AI transcription hits about 2.7% word error rate on clean read English - near the human floor - but exceeds 30% on noisy group audio. - [How Many Words Are in an Hour of Audio? About 8,000-9,000.](https://pepys.co/blog/how-many-words-in-an-hour-of-audio): An hour of spoken audio is roughly 8,000-9,000 words. Conversation runs 120-150 words a minute, audiobooks 150-160 - and it takes about 35 minutes to read. - [Transcription Cost Data 2026: Every Published Rate on One Axis](https://pepys.co/blog/transcription-cost-data-2026): 2026 transcription costs span ~300x per audio minute: fractions of a cent for cloud APIs, cents for AI tools, $0.70-$2.00 for humans, and per-page for courts. - [Which Languages AI Transcription Struggles With (and Why)](https://pepys.co/blog/which-languages-ai-transcription-struggles-with): AI transcription accuracy swings ~40x by language: ~3-4% word error rate on English and Spanish, over 100% on low-resource languages like Amharic and Sindhi. - [How Well Does AI Handle Crosstalk? The Data Says: Badly, and Here's Why.](https://pepys.co/blog/how-well-does-ai-handle-crosstalk): AI isn't bad at telling voices apart – it's bad at hearing two at once. Published benchmarks show diarization error roughly doubles from clean to overlap-heavy audio. - [How Long Does It Take to Transcribe an Hour of Audio?](https://pepys.co/blog/how-long-to-transcribe-an-hour-of-audio): Manual transcription takes about 4 hours of work per hour of audio (the 4:1 rule), rising to 10 for verbatim. AI does it in minutes. Here's the full timeline. - [The State of Transcription in 2026: A Data Report](https://pepys.co/blog/state-of-transcription-2026): State of transcription 2026: AI nails clean English at 2.7% WER, but accuracy tracks training data, diarization breaks on crosstalk, and caps are pricing. - [We Read 100 Reddit Threads About Transcription. Here's What Everyone Hates.](https://pepys.co/blog/we-read-100-reddit-threads-about-transcription): We read ~100 Reddit threads where people begged for a better transcription tool. The same four complaints came up in almost every one. Here's what they are.