Creator-uploaded
Supplied or edited by the channel owner or production team.
YouTube caption source guide
Both track types can make a video searchable and accessible, but they are not interchangeable. Knowing the source helps you decide how much review a transcript needs.
Supplied or edited by the channel owner or production team.
Created by speech recognition from the video's audio.
Derived from another track and best labeled separately.
A production team can correct names, specialist terminology, punctuation, speaker changes, and meaningful sound descriptions. They can also time lines to match editorial intent. This usually makes a creator-uploaded track the preferred source.
Automatic captions provide coverage when a creator has not uploaded a file. They can work well with clear speech, but accuracy can fall with noise, music, accents, multiple speakers, weak audio, or uncommon terminology.
The extractor prioritizes a creator-uploaded track in the requested language. If none is accessible, it can use an available automatic track and labels the source in the result. Customers can decide whether the output suits casual reading or needs careful review.
Manual does not always mean professionally proofread, and automatic does not always mean unusable. The label describes provenance. Important quotations, research findings, legal language, and accessibility deliverables should receive human verification.
TXT, SRT, VTT, CSV, and JSON package the same selected caption source for different workflows. Choosing a different file format does not correct errors in the underlying caption text.
Frequently asked questions
No. They are usually preferable, but a creator may upload an unedited file or intentionally shorten spoken content.
Yes, when the video exposes an accessible automatic caption track.
The extraction workflow preserves source wording; readability grouping combines cues without paraphrasing them.
Creator-uploaded captions preferred. Clean reading and precise exports.
Download a transcript free →