Fonts
34
Size is in pixels against your video's own resolution, so the preview matches a real burn-in.
Caption templates
Each one sets every value on the Text tab at once. Pick the closest, then fine-tune.
Caption script
Whisper always writes Urdu speech in Urdu script — Roman mode runs a local transliteration pass over it, mapping English loanwords (پریمیم → premium) back to real English spelling. Approximate by design; hand edits in the captions panel are always kept as typed.
Corrections & brand names
One per line as wrong = right. Applied after transliteration, so client names come out right every time. No speech model reliably spells brand names — this is how you pin them.
Recognition
Leave on auto for mixed Urdish. If a clip comes out as looping gibberish, forcing Urdu often fixes it — the tool also retries that automatically.
Model storage
Browser cache always lives in your Chrome profile on C: — a web page can't choose a drive. Turn this off when loading from your own folder and nothing is written to C: at all.
—
About
Built for MS Production — Creative Media Agency. Speech recognition runs locally in this tab via an open-source Whisper model, accelerated with WebGPU where available. The model downloads once from Hugging Face's CDN and is cached by the browser; your video never leaves this device.
Open this file directly in a browser — model downloads and local processing need normal browser permissions.
Subtitle files
Drop the .srt straight into CapCut, Premiere or Resolve.
Export video with captions
Video
12
Mbps. Browsers don't expose a clip's source frame rate, so pick the one your footage was shot at.
Audio
Load a video to see output settings.
Burn in with ffmpeg
Bakes your current style into the video. Styling maps as closely as ASS subtitles allow.