Generating subtitles from speech: the pipeline and what hurts accuracy
In one line: subtitle generation turns speech into text with a timestamp for every line. The text depends on the recognition model; the timing depends on segmentation and alignment.
The pipeline
- Pick the audio track. On multi-track files, choose the right one (original, dub or commentary).
- Segment. Split on silence and voice activity — segments that are too long slow things down, too short breaks meaning.
- Recognise. The speech model emits candidate text, optionally with confidence scores.
- Punctuate. Turn the continuous transcript back into readable sentences.
- Time it. Attach start and end times so a player can display each line.
What hurts accuracy
| Factor | Effect | Mitigation |
|---|---|---|
| Music / background noise | Harder to separate speech | Prefer a clean track; denoise first if needed |
| Accents and dialects | More substitution errors | Choose a closer model, then proofread key terms |
| Proper nouns, names | Often misspelled | Do one find-and-replace pass on the export |
| Overlapping speakers | Merged or dropped lines | Work in segments and proofread each |
Worth doing after generation
- Read through proper nouns — people, places, technical terms.
- Check timing: lines that flash by too fast, or drift out of sync.
- For multiple languages, translate the proofread transcript, not the raw audio.
Frequently asked questions
Why does subtitle generation get words wrong?
Mostly background noise, accents, overlapping speech and proper nouns. Use the cleanest audio track you have and do one proofreading pass on names and terms.
Should I translate the audio or the subtitles?
Translate the subtitles. A proofread transcript produces a far better translation than raw speech.
Is the timing automatic?
Yes — segmentation and voice activity provide start and end times. If you re-edit the video afterwards, re-run the alignment.
More topics on the guides hub, or see what Lumasce does.