Transcript generator › Script extractor
YouTube script extractor
Get what was actually said in a video, as readable prose rather than caption fragments. Paste a link and the script comes back in a few seconds.
What "script" means here
Worth being precise, because the word covers two different things and only one of them is obtainable.
What you can get is the spoken script: everything said aloud in the video, reconstructed from its caption track. That is what this page produces.
What you cannot get is the creator's original writing document — their draft with stage directions, b-roll notes and the lines they cut. That never leaves their machine. Any tool claiming to hand you a creator's actual script is giving you the spoken words with a more exciting label on them.
Captions are a record of delivery, not of intent
A script is what someone planned to say; captions are what came out. Presenters improvise, stumble, repeat themselves and cut things live. Read the extracted text as a faithful record of the delivery — which is usually what you want anyway, since that is the version that worked.
What people use it for
Studying how a video is built
Reading a talk as prose makes its structure visible in a way that watching does not. Where the hook lands, how long before the first real point, how transitions are handled, how the close is set up — all of it is much easier to see on a page than in a timeline. Creators studying videos that outperformed theirs are the most common users of this; reading a transcript for structure covers the method in detail.
Referencing and quoting
Pulling an exact line out of a 40-minute talk means scrubbing back and forth until you find it. With the script in front of you it is a text search. Switch to the Timestamps tab and every line links back to its moment, so you can verify the quote in context before publishing it.
Research across many videos
Comparing how a dozen videos explain the same concept is impractical by watching and straightforward as text. Extract each one, put them side by side, and the differences in framing show up immediately. For bulk work, the API does one request per video.
Accessibility and translation
A clean prose script is a better starting point for translation, summarising or reading with assistive software than a column of timed fragments.
Why the raw captions are not a script
Downloading captions gets you the words but not something readable. Subtitles are timed in short bursts so each fits on screen for a moment, which means the line breaks fall on timing boundaries rather than sentence boundaries. In our testing an 18-minute lecture came back as 286 separate caption lines, and a 4.4-hour course as 2,935.
Two things have to happen before that is a script:
- Rejoin the fragments into sentences and paragraphs, so it reads as prose.
- Decide where paragraphs break — the harder half. Where the captions carry punctuation, sentence ends work. Auto-generated captions frequently contain no punctuation at all, and splitting on full stops that do not exist yields one enormous block. A music video we tested collapsed into a single unbroken 2,089-character paragraph that way.
The fallback used here is speech pauses: when a paragraph runs long without a sentence end, it breaks at the largest gap between caption lines, because that gap is almost always where the speaker actually paused. That turned the same music video into two sensible paragraphs, and keeps the 4.4-hour course at an average of about 1,000 characters per paragraph with no runaway blocks.
Limits worth knowing
| Situation | Result |
|---|---|
| Video has captions | Full script, typically in a few seconds |
| Only auto-generated captions | Works, and is flagged as auto — expect misheard names and jargon |
| No caption track at all | Nothing to extract; this reads captions, it does not transcribe audio |
| Age-restricted or private | Requires a signed-in session, so captions are unreachable |
| Music, sound effects, silence | Appears as caption markers such as [♪♪♪], not as speech |
On auto-generated tracks specifically: the transcript tells you which kind it used, and if you are going to quote from one, click the timestamp and listen first. Speech recognition is reliably wrong about proper nouns and technical terms.
Using someone else's script
The words in a video belong to whoever made it. Extracting the script to study it, quote from it with credit, or research across sources is ordinary use.
Re-recording someone's script as your own video is not, and neither is republishing the full text as an article. Beyond the legal exposure, it is a poor strategy: the reason a video worked is rarely the words alone. Study the structure, then write your own.
Frequently asked
Can I get the creator's original script document?
No. What is available is the spoken script reconstructed from captions. The creator's own writing document is never published.
Does it work on videos in other languages?
Yes, when the video has a caption track in that language. Every available track is listed after extraction — see translated subtitles.
How accurate is it?
Human-authored caption tracks are essentially exact. Auto-generated ones are good on clear speech and unreliable on names, jargon and anything over music. The track type is always shown.
Can I extract scripts from a whole channel?
Not from this page, which handles one video at a time. Loop the API over the video IDs you care about.
Does it include stage directions or shot notes?
No. Captions record spoken words. Anything not said aloud is not in the caption track and cannot be recovered from it.
More guides
Independent tool. youtubegpt.ai is not affiliated with, endorsed by, or operated by YouTube or Google. It works with publicly available caption data from YouTube and does not host, mirror or re-upload any video. Transcript text belongs to the original video creators. "YouTube" is a trademark of Google LLC.