Captions
Karaoke captions with word timing.
project.dispatch({
type: "clip/add",
payload: {
kind: "caption",
id: "cap-1",
trackId: "captions",
startUs: 2_000_000,
durationUs: 1_800_000,
words: [
{ text: "made", startUs: 0, durationUs: 600_000 },
{ text: "with", startUs: 600_000, durationUs: 600_000 },
{ text: "miraiclip", startUs: 1_200_000, durationUs: 600_000 },
],
style: { preset: "karaoke", highlightColor: "#ffd400" },
transform: { y: 0.8 },
},
});A caption clip draws its words as one centered block. The words wrap at 80% of the composition width. The style preset decides how the word being spoken stands out. Caption clips go on video tracks.
Words
| Field | Meaning |
|---|---|
text | One word, as spoken |
startUs | Start, relative to the clip's start |
durationUs | How long the word is active |
A word is active from startUs to startUs + durationUs. Gaps between words are allowed. Core doesn't check that words fit inside the clip.
Presets
preset | Look |
|---|---|
plain | No emphasis |
highlight | The active word takes highlightColor (default) |
karaoke | Every word that has started stays in highlightColor |
pop | The active word takes highlightColor and grows |
reveal | Words appear as they start. The active word takes highlightColor |
display: "word" works with every preset: only the current word shows, and it stays up through gaps until the next word starts.
Style
| Field | Default | Meaning |
|---|---|---|
preset | "highlight" | See above |
fontFamily | "sans-serif" | A font asset's family, or a system font |
fontSizeFrac | 0.06 | Font size as a fraction of composition height |
color | "#ffffff" | Word color |
highlightColor | "#ffd400" | Emphasized word color |
display | "block" | "block" (all words) or "word" (current word only) |
backgroundColor | absent | Rounded box behind the block (behind the word with display: "word") |
activeBackgroundColor | absent | Rounded box behind each emphasized word |
textTransform | absent | "uppercase" or "lowercase" when drawing. The words keep their text |
strokeColor | absent | Outline color |
strokeWidthFrac | 0.08 | Outline width, as a fraction of font size |
shadowColor | absent | Drop shadow color |
shadowBlurFrac | 0.15 | Shadow blur, as a fraction of font size |
shadowOffsetFrac | 0.06 | Downward shadow offset, as a fraction of font size. 0 makes a glow |
fontWeight, fontStyle, lineHeight, letterSpacing | absent | Typography. lineHeight defaults to 1.3 here |
Captions are always centered, so they take no textAlign. Sizes are fractions, so a caption looks the same in a scaled preview and a full-size export.
// style merges: only the fields you pass change
project.dispatch({
type: "clip/set-property",
payload: { clipId: "cap-1", style: { preset: "pop", strokeColor: "#000000", textTransform: "uppercase" } },
});
// null clears an optional style field
project.dispatch({
type: "clip/set-property",
payload: { clipId: "cap-1", style: { strokeColor: null, backgroundColor: "#000000" } },
});preset, fontFamily, fontSizeFrac, color and highlightColor always have a value and don't accept null. Every other style field does.
Import subtitles
captionClipsFromSubtitles reads SRT or WebVTT and returns clip/add commands, one caption clip per cue. Words split each cue's time evenly.
import { applyCommands, captionClipsFromSubtitles } from "@miraiclip/core";
declare const srtText: string;
const commands = captionClipsFromSubtitles(srtText, {
trackId: "captions",
style: { preset: "karaoke" },
});
applyCommands(project, commands); // one undo stepImport ASR word timestamps
Speech-to-text word timings give real karaoke timing. captionClipsFromAsrWords takes Whisper-style words in seconds and groups them into clips.
import { applyCommands, captionClipsFromAsrWords, type AsrWord } from "@miraiclip/core";
const words: AsrWord[] = [
{ text: "So", startS: 0.0, endS: 0.2 },
{ text: "this", startS: 0.25, endS: 0.5 },
{ text: "works.", startS: 0.55, endS: 0.9 },
{ text: "Next", startS: 2.0, endS: 2.3 },
];
applyCommands(project, captionClipsFromAsrWords(words, { trackId: "captions", maxWordsPerGroup: 6 }));
// Two clips: "So this works." at 0 s and "Next" at 2 s (the 1.1 s gap starts a new clip)| Option | Default | Meaning |
|---|---|---|
trackId | required | Target track |
style | default style | Merged over the default caption style |
idPrefix | "caption" | Clip ids are ${idPrefix}-1, ${idPrefix}-2, … |
maxGapUs | 600_000 | ASR only: a longer silence starts a new clip |
maxWordsPerGroup | 8 | ASR only: maximum words per clip |
parseSubtitles(text) returns the raw cues ({ startUs, endUs, text }) if you want to build clips yourself. It ignores cue numbers, NOTE/STYLE blocks and cue settings, and strips tags like <i>.
Edit text, keep timing
import { isCaptionClip, retimeWords } from "@miraiclip/core";
const clip = project.getState().doc.clips["cap-1"]!;
if (isCaptionClip(clip)) {
const words = retimeWords(clip.words, "made with Miraiclip", clip.durationUs);
project.dispatch({ type: "clip/set-property", payload: { clipId: clip.id, words } });
}With the same word count, each word keeps its timing. With a different count, the words split the clip's length evenly. Empty text returns [].
Export a transcript
import { captionsToSrt, captionsToText, captionsToVtt, isCaptionClip } from "@miraiclip/core";
const captions = Object.values(project.getState().doc.clips).filter(isCaptionClip);
const srt = captionsToSrt(captions); // one cue per clip, in timeline order
const vtt = captionsToVtt(captions);
const text = captionsToText(captions); // one line per clipFonts
Register a font asset and use its family in style.fontFamily. See Text & typography.
project.dispatch({
type: "asset/add",
payload: { id: "brand-bold", kind: "font", src: "/fonts/Brand-Bold.woff2", family: "Brand", weight: 700 },
});
project.dispatch({
type: "clip/set-property",
payload: { clipId: "cap-1", style: { fontFamily: "Brand", fontWeight: 700 } },
});Notes
- The preset list is fixed. For a look the presets can't produce, use an HTML clip or a custom clip kind.
- Caption clips take keyframes, effects and transitions like any other clip.