An AI-agent skill for creating, translating, formatting, and delivering subtitles for local audio and video.
It combines FFmpeg, the OOMOL Fusion API, an optional OOMOL-configured LLM, and a bundled Node.js helper to produce readable subtitle files and subtitled videos. Video delivery defaults to a burned-in MP4, with selectable soft-subtitle MKV as an alternative.
Subtitle presentation is designed around the readability principles in Netflix's Timed Text Style Guide, including concise cues, controlled line length, natural phrase-aware line breaks, and unobtrusive placement:
- Each cue targets about 2.8 seconds on screen and is capped at 4.2 seconds.
- Subtitles use no more than two lines, with one line preferred whenever possible.
- CJK subtitles default to 18 characters per line; Simplified Chinese can use the stricter 16-character limit from Netflix's Simplified Chinese guide.
- Long text is divided into consecutive cues instead of being squeezed onto the screen.
- Two-line subtitles favor natural phrase boundaries and a bottom-heavy shape.
- Burned-in subtitles use resolution-aware, bottom-centered white text with a black outline.
The goal is simple: spend less time reading and more time watching. These are Netflix-informed readability defaults, not a claim of Netflix delivery certification or full compliance with every language-specific specification.
Transcribe English narration, translate it into natural Simplified Chinese, format the cues for the screen, and burn the styled subtitles into an MP4.
▶ Play the original English → Chinese MP4
Turn Japanese narration into timestamp-preserving English and Simplified Chinese subtitle outputs. The preview below shows the Simplified Chinese burn-in result; the same workflow can produce the English subtitle version.
▶ Play the original Japanese subtitle workflow MP4
The demo recordings use “Generative AI explained in 2 minutes” by KI-Campus and “Job Interview” by Simpleshow Japan, both under CC BY-SA 4.0. The recordings show the agent workflow and modified subtitle outputs.
- Extract and normalize audio from local video with FFmpeg.
- Upload local media through the
ooCLI and transcribe it with Fusion API ASR. - Generate timed SRT and WebVTT subtitles from word-level timestamps.
- Remove narrowly scoped Fusion ASR artifacts without rewriting user-provided or translated subtitle text.
- Translate subtitle cues with an OpenAI-compatible LLM configured by
oo llm config. - Resume interrupted translations only when checkpoint cue indexes and source text still match.
- Reflow CJK and Latin subtitles into readable display lines.
- Generate resolution-aware ASS subtitles for consistent styling and positioning.
- Burn subtitles into MP4 by default, or package selectable SRT/ASS tracks into MKV.
flowchart TB
subgraph CORE["Core processing"]
direction LR
INPUT["Video / Audio"]
PREP["① Prepare audio<br/>FFmpeg + oo upload"]
ASR["② Transcribe<br/>Fusion API ASR"]
LOCAL["③ Format locally<br/>optional LLM → SRT / ASS"]
INPUT --> PREP --> ASR --> LOCAL
end
LOCAL --> DELIVERY{"④ Deliver"}
DELIVERY -->|"Default"| BURN["Burned-in MP4<br/>FFprobe + FFmpeg + libass"]
DELIVERY -->|"Optional"| MKV["Soft-subtitle MKV<br/>SRT / ASS track"]
BURN --> OUTPUT["Video + subtitle files"]
MKV --> OUTPUT
The main responsibility boundary is:
Cloud: media hosting, speech recognition, and optional translation
Local: timestamp normalization, ASR cleanup, cue layout, ASS styling, and video delivery
- Node.js 18 or newer
ffmpegandffprobe- The OOMOL
ooCLI (required), authenticated and able to use:oo file upload- the
fusion-apiconnector oo llm configwhen translation is requested
Check the local runtime:
node --version
ffmpeg -version
ffprobe -version
oo --versionThe skill cannot run its ASR and translation workflow without oo. If it is
not installed, the agent will stop and guide you through this one-time setup;
it will not install software without your approval.
macOS or Linux:
curl -fsSL https://cli.oomol.com/install.sh | bashWindows PowerShell:
irm https://cli.oomol.com/install.ps1 | iexThen open a new terminal or refresh PATH, and authenticate:
oo --version
oo auth loginSee the official installation guide for other installation options.
Clone the repository into an agent skills directory:
git clone https://github.com/oomol-lab/video-subtitle-translator.git \
~/.agents/skills/video-subtitle-translatorAlternatively, clone it anywhere and link it into the skills directory:
git clone https://github.com/oomol-lab/video-subtitle-translator.git
mkdir -p ~/.agents/skills
ln -s "$(pwd)/video-subtitle-translator" \
~/.agents/skills/video-subtitle-translatorRestart or refresh the agent environment after installation so it can discover SKILL.md.
The skill is designed to be invoked through an AI agent. Example requests:
Create Simplified Chinese subtitles for this video and burn them into an MP4.
Transcribe this audio and return SRT and VTT files without translation.
Translate this interview into Japanese and package the subtitles as a selectable MKV track.
Create natural Chinese subtitles for this technical course. Keep API names and code terms in English.
If translation is clearly requested but the target language is omitted, the skill infers the target from the language of the request. If it is unclear whether translation is wanted, the agent asks before starting the job.
For local video or media that needs normalization, the skill extracts a mono 16 kHz WAV:
ffmpeg -y -i "$INPUT_MEDIA" \
-vn -ac 1 -ar 16000 -c:a pcm_s16le \
"$WORK_DIR/audio.wav"The local file is uploaded before it is passed to the cloud ASR service:
oo file upload "$AUDIO_PATH" --jsonOnly the returned public downloadUrl is sent to Fusion API. Local filesystem paths are never passed directly to the cloud action.
The ASR workflow uses three connector actions:
qwen_asr_filetrans_submit
↓
qwen_asr_filetrans_state
↓
qwen_asr_filetrans_result
The full result is preserved as transcript.json. Fusion timestamps are treated as milliseconds by default and normalized to seconds locally.
The bundled helper provides four commands:
node scripts/subtitle-tools.mjs fusion-to-subtitles --help
node scripts/subtitle-tools.mjs translate-srt --help
node scripts/subtitle-tools.mjs prepare-display-srt --help
node scripts/subtitle-tools.mjs srt-to-burn-ass --helpTypical transcript conversion:
node scripts/subtitle-tools.mjs fusion-to-subtitles \
--input "$WORK_DIR/transcript.json" \
--out-dir "$WORK_DIR" \
--time-unit ms \
--formats srtTypical display formatting:
node scripts/subtitle-tools.mjs prepare-display-srt \
--input "$WORK_DIR/translation.zh.srt" \
--output "$WORK_DIR/translation.zh.display.srt" \
--cjk-line-length 18 \
--max-lines 2The final SRT is converted to ASS using the real video dimensions, then burned into the video:
ffmpeg -y -i "$INPUT_VIDEO" \
-vf "ass=$WORK_DIR/subtitles.burn.ass" \
-c:v libx264 -crf 18 -preset medium \
-c:a copy -sn \
"$OUTPUT_VIDEO.burned.mp4"Use this mode for social platforms, mobile playback, uploads, and other environments where subtitles must always be visible.
Package a selectable subtitle track without burning it into the image:
ffmpeg -y -i "$INPUT_VIDEO" -i "$SUBTITLE_SRT" \
-map 0:v? -map 0:a? -map 1:0 \
-c copy -c:s srt \
-disposition:s:0 default \
-metadata:s:s:0 language="$LANG_CODE" \
"$OUTPUT_VIDEO.subtitled.mkv"Use this mode when subtitles should remain selectable or editable, when multiple subtitle languages are needed, or when avoiding video re-encoding matters.
Translation uses the OpenAI-compatible model returned by:
oo llm config --jsonThe helper translates cue text in batches while preserving indexes and timestamps. It supports translation profiles for film and television, technical courses, interviews, news, business training, gaming, children’s content, and general video.
To resume an interrupted translation:
node scripts/subtitle-tools.mjs translate-srt \
--input "$WORK_DIR/transcript.srt" \
--out-dir "$WORK_DIR" \
--target-language "Simplified Chinese" \
--target-code zh \
--resumeCheckpoint entries are reused only when both the cue index and saved source text match the current SRT.
Noise cleanup is intentionally restricted to Fusion ASR input:
- A complete ASR token containing only three or more zeroes, such as
000, is treated as a probable artifact. - Duplicate punctuation is not appended when the ASR word already contains it.
- Normal expressive forms such as
...,??,!!, and?!are preserved. - Only clearly excessive punctuation runs are normalized.
- User-provided SRT, LLM translations, and checkpoint text are not subjected to destructive cleanup.
When cleanup occurs, the helper reports how many probable zero artifacts and punctuation runs were repaired.
A typical job can produce:
job.created.json
job.done.json
transcript.json
transcript.txt
transcript.srt
transcript.word-timed.srt
translation.<language>.json
translation.<language>.srt
translation.<language>.display.srt
translation.<language>.display.ass
<name>.burned.mp4
<name>.subtitled.mkv
- Local media is uploaded to obtain a URL that Fusion API can access.
- Translation text is sent to the LLM configured by
oo llm config. - API keys are read at runtime and must not be committed, printed, or written into output files.
- Raw transcript JSON is preserved locally for debugging and recovery, but should not be published when it contains sensitive speech.
Issues and pull requests are welcome. When changing subtitle conversion behavior, include a small fixture that covers the affected timestamp, punctuation, language, or checkpoint edge case.
Before submitting a change, at minimum run:
node --check scripts/subtitle-tools.mjs
node scripts/subtitle-tools.mjs --helpMIT © 2026 OOMOL Lab

