Transcribing a YouTube video is either a ten-second job or a real workflow, and the difference is what you need the text for. Pulling a quote from a talk is not the same task as turning fifty episodes of your own channel into blog posts. We do both at the agency, so this guide covers the free built-in route first, then the tools that earn their cost when the built-in route falls short.
The Quick Answer
Grab a quote or skim a video as text. Use YouTube's built-in transcript panel. Free, instant, works on any public video with captions.
Accurate, exportable transcripts of your own content. Run the audio through ElevenLabs Scribe. Speaker labels, timestamps, clean exports.
Transcribing because you're editing the video. Descript, where the transcript is the editing surface.
Audio that can't leave your machine. A local transcription app like MacWhisper.
Rights First, the Part Everyone Skips
One caveat before the methods, because it decides which methods you're allowed to use. Reading the transcript of any public video is fine, that's what the transcript panel is for. Downloading someone else's video or audio to run through a transcription tool is a different act. YouTube's Terms of Service prohibit downloading content without permission, and the third-party ripper sites that promise otherwise don't change that. The workflows below that involve uploading audio assume it's your own channel, your client's channel with a green light, or content you have permission to use. That covers the honest use cases, show notes, repurposing, translations, and archives.
Method 1. YouTube's Built-In Transcript Panel
YouTube generates captions for most videos automatically, and the transcript panel exposes them as scrollable, copyable text. No tool, no download, no account.
- Open the video on desktop.
- Expand the description with the ...more link.
- Click Show transcript. The transcript opens beside the player.
- Use the three-dot menu in the transcript panel to toggle timestamps off if you want clean text.
- Select the text and copy it.
The limits are real, though. Auto-captions fumble names, jargon, and heavy accents, there are no speaker labels, punctuation is thin, and copying a long video's transcript by hand-selecting is tedious. For skimming, quoting, and checking what was said, it's unbeatable at free. For anything you'd publish, it's a rough draft at best.
Method 2. ElevenLabs Scribe, the Accuracy Route
When the transcript is the deliverable, use a dedicated speech to text engine. ElevenLabs Scribe is the one in our stack. ElevenLabs bills it as industry-leading transcription accuracy across 99 languages, with speaker diarization and word-level timestamps built in, and quotes processing at 20-50x real-time, so an hour of audio comes back in a few minutes.
- Get the audio file. For your own channel, download the video from YouTube Studio, or better, use the original recording before it ever hit YouTube.
- Sign in to ElevenLabs and open the speech to text tool.
- Upload the file and pick the language, or let it detect one of the 99 it supports.
- Review the transcript. Speaker labels and timestamps are already in place.
- Export the text and use it anywhere.
Cost math. Transcription runs about 330 credits per minute, the free plan's 10,000 monthly credits cover roughly 12 minutes, and paid plans start at $6 a month. A weekly half-hour show fits comfortably inside the cheap tiers. We break down the whole ElevenLabs lineup, including where Scribe sits in it, in our ElevenLabs products explainer.
Method 3. Descript, When You're Also Editing
If you're transcribing your video because you're about to cut it into clips, repurpose it, or fix the audio, skip the standalone transcript and import the file into Descript. It transcribes on import, and then the transcript becomes the editor. Delete a sentence in the text, the cut happens in the video. Its free plan covers 60 minutes of media a month, and the paid Hobbyist tier runs $16 a month billed annually with 10 media hours.
The reason not to pick Descript is the same as the reason to pick it. You're paying for a full editor. If you never touch the media, that's the wrong subscription, a point we labor in our Descript vs Otter breakdown.
Method 4. Local Apps, When Privacy Decides
Some audio shouldn't touch anyone's cloud. Client calls under NDA, unreleased content, legal material. For those, local transcription apps like MacWhisper run the speech model entirely on your machine, the file never leaves your disk, and there's no per-minute meter. The trades are hardware speed, no collaborative features, and accuracy that depends on which local model your machine can run. For the private slice of the workload, it's the right answer. For everything else, the cloud tools above are faster to live with.
The Four Compared
| Method | Cost | Best for | The catch |
|---|---|---|---|
| Transcript panel | Free | Quotes and skimming any public video | Rough accuracy, no speaker labels, manual copying |
| ElevenLabs Scribe | Free ~12 min/mo, paid from $6/mo | Publishable transcripts of your own content | Credits meter the minutes |
| Descript | Free 60 min/mo, paid from $16/mo annual | Transcribing while editing the video | You're buying a whole editor |
| Local apps | Varies by app | Audio that can't leave your machine | Speed and accuracy depend on your hardware |
Putting Transcripts to Work
A transcript sitting in a folder is a receipt. The marketing value shows up when it becomes something searchable. The workflow we run for content clients looks like this. Transcribe the video, turn the transcript into a blog post that targets the query the video answers, pull the strongest lines as social copy, and feed the full text to whatever AI assistant drafts your show notes. Search engines and AI answer engines can't watch your video, but they can read its transcript, which is why the channels that transcribe consistently outrank the ones that don't. If you're picking which videos deserve the treatment first, our YouTube keyword research guide covers how we choose targets.
And if the transcription job you actually have is live speech rather than uploaded files, that's dictation, a different category with different tools, mapped in our ultimate guide to AI dictation.