Skip to content

feat: auto-align word timing from tapped line timing (HubertFA, en/zh/ja) - #264

Draft
adaliea wants to merge 3 commits into
t3code/debug-webgpu-initializationfrom
t3code/forced-alignment
Draft

adaliea wants to merge 3 commits into
t3code/debug-webgpu-initializationfrom
t3code/forced-alignment

Conversation

@adaliea

@adaliea adaliea commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Written by Claude (Anthropic's AI model, running as Claude Code on Opus 5.5) working with @Adalie. Claude did the research, benchmark, implementation and model hosting. Stacked on #262; please review that first.

What this does

You sync each line's start and end in Line mode (rough taps are fine), then open Auto-align in the Sync tab. It times every word from the separated vocals and switches to Word mode.

  • Model: HubertFA v0.0.7 (ONNX, Apache-2.0). It runs in a worker on WebGPU, or on WASM as a fallback.
  • Decoder: its forced-alignment (Viterbi) decoder is ported to TypeScript and matches the Python reference exactly (0.000 ms difference on the reference windows).
  • Popover: built like the vocal separation popover. It shows model download and alignment progress, Cancel, Retry/Dismiss, and steps for syncing lines and separating vocals.
  • Re-align: lines that already have word timing can be aligned again. Each part's text, explicit flag, syllable group and transliteration are kept; only the times change.
  • Timeline: the right-click menu has Auto-align / Re-align for selected lines.
  • Languages: English, Mandarin and Japanese, mixed freely within a line ("君と dance"). These are all the languages the model covers.
    • Character splitting: unspaced Chinese and Japanese is split into one timed part per character. Small kana, ー and っ stay with the kana before them, and each part gets its own timing from the model.
    • Mandarin: characters become toneless pinyin via pinyin-pro (MIT), with characters that have more than one reading resolved from context. It's loaded only when needed.
    • Japanese: kana map directly to the model's romaji morae. Kanji readings come from kuromoji (@sglkc/kuromoji, Apache-2.0, IPADIC). Its 18 MB dictionary is hosted next to the model, loaded only for Japanese lyrics, and cached the same way.
    • Han characters: they're read as kanji in Japanese projects or when any line has kana, and as Mandarin otherwise.
    • Transliterations: when a line has a transliteration that differs from the line itself, its syllables are what gets sung. The dictionaries only decide which characters each syllable belongs to: the two readings are matched syllable by syllable (an edit-distance alignment), with English words as fixed points. So 運命 sung as "sadame", or a kanji the tokenizer can't read, comes out right.
      • Lines whose transliteration is identical to the line (English lines), stale tracks, and lines without one use the dictionaries.
      • Per-part transliterations aren't used on their own, because imported lyrics often pair them with the wrong characters. In ENEMY (TWICE), for example, 完 is paired with "ka" and 璧 with "n". The whole line's text is used instead.
  • Fallbacks:
    • A line the model can't place gets the existing character split, and the run lists the reason. That covers words missing from the dictionary (about 0.5–0.8% of English words) and other scripts such as Korean.
    • Projects set to other languages are refused.
  • Hosting: the model (396 MiB) and dictionary (3.3 MB) are served next to the vocal separation model and cached with the Cache API.
  • Scripts:
    • scripts/build-hubertfa-onnx.sh builds the model. It removes two fp32→fp16→fp32 Cast round-trips from the graph, because WebGPU without shader-f16 can't run them; benchmark accuracy was unchanged.
    • scripts/upload-hubertfa.sh uploads through R2's S3 API, since the file is over wrangler's 300 MiB limit.

Benchmark

Setup:

  • Songs: 27 English songs.
    • 20 from JamendoLyrics (Creative Commons, research-grade word timings).
    • 7 commercial songs from Unison, with human-tapped word timings; that reference is noisier.
  • Vocals: HTDemucs, the same model the app uses.
  • Line taps: simulated from the ground truth. Starts land about 150 ms late ± 150 ms; ends land ± 250 ms.
  • Baseline: what Composer does today when expanding a line, which is to split it by character count.

Word start error

Char split (today) SOFA (English community model) HubertFA (this PR)
Jamendo median 270 ms 73 ms 50 ms
Jamendo words within 200 ms 38% 78% 90%
Unison median 202 ms 123 ms 83 ms
Unison words within 200 ms 50% 67% 83%

Lines needing no correction (every word within 200 ms)

Char split SOFA HubertFA
Jamendo 11% 37% 65%
Unison 14% 21% 44%
Words per line more than 0.3 s off (Jamendo / Unison) 3.0 / 2.2 1.0 / 1.5 0.4 / 0.6

Lines average about 7 words.

  • Every song improves: HubertFA beats both the char split and SOFA on all 27 songs. SOFA was worse than the char split on 2 commercial songs.
  • Perfect taps don't fix the char split: splitting by character count with perfect line taps still gives a 162 ms median on Jamendo.

What the settings tests showed

Helped:

  • Separated vocals are required. On the full mix SOFA does worse than the char split, and HubertFA loses about 30 ms of mean error.
  • Small margins around the line help. No margin cuts off the first word; a 1 s lead-in pulls in other lines' vocals. The app uses 0.3 s before the line (stopping where the previous line ends, since Line mode taps lines back to back) and 0.15 s past the end tap, so short last words aren't cut off. That end padding came from manual testing on real taps and isn't benchmarked.
  • Unknown words must not be dropped. Dropping them, as the reference tools do, lost 9.7% of words on raw Unison lyrics. The app normalises spelling and falls back per line instead.

Made no difference: HubertFA's padding ensemble, high-pass filtering, loudness normalisation, and the breath detector or post-processing.

Hurt:

  • aligning 2–3 lines at a time
  • retrying lines whose words hit the window edge
  • SOFA's "match" mode

Tap accuracy barely matters. With perfect taps HubertFA's mean error goes only from 109 to 98 ms.

Ported versions:

  • The fp32 model with the fp16 round-trips removed gives the same word onsets for 99.8% of words (Jamendo median 50.6 ms in both).
  • In the browser, one commercial song (37 lines, rough taps) scored a median of 87 ms, with 88% of words within 200 ms. Python on the same song scored 86 ms and 88.3%.

Speed: WASM on one thread takes about 11 s to load the model, then about 2 minutes per song. WebGPU should be much faster; it has been tried on an NVIDIA GPU without shader-f16, but there's no timing for it yet.

Mandarin and Japanese (real songs)

Setup:

  • Code under test: the shipped TypeScript pipeline run in Node: the text-to-phone step (G2P), the model via onnxruntime-web, the decoder and the windows.
  • Songs: 4 Unison songs with crowdsourced per-character timings: 平行宇宙 and 演员 (Mandarin), and 桜流し and 不可思議のカルテ (Japanese).
  • Line taps: simulated the same way as the English benchmark.
  • Caveat: the reference timings are low-vote and crowdsourced, so treat this as indicative.
Word start error Char split (today) Auto-align
Mandarin (802 characters): median / within 200 ms 441 ms / 20% 84 ms / 90%
Japanese (574 spans): median / within 200 ms 439 ms / 27% 103 ms / 76%

ENEMY (TWICE): 63 lines of mixed Japanese and English, with word timing from the file the user supplied.

Word start error Char split Auto-align
All 358 words: median / within 200 ms 236 ms / 43% 11 ms / 86%
Japanese lines (222 spans): median / within 200 ms 219 ms / 46% 7 ms / 95%

On this song the transliterations change no readings: kuromoji already reads every word the way the romanization does (完璧 → kanpeki), so the results with and without them are identical. They matter for lyrics with readings that differ from the dictionary's, which the unit tests cover.

No line fell back to the character split. Of every character pinyin-pro can read (20,902 tested), all but 51 rare ones map to a syllable in the model's dictionary. The interjections 嗯 and 呣 are mapped to the nearest syllable.

Known limits

  • Common mistakes:
    • repeated phrases within a line
    • "oh / la" lines
    • held last notes, where the last word's error is about 2× the others
    • backing-vocal echoes
  • Confidence flags aren't useful yet. The model's own per-word confidence is a weak predictor of errors, so this PR doesn't flag words. Flagging words where two runs disagree catches far more errors (the top 10% contain 56% of the big ones) and is a possible follow-up.
  • No syllable timing from the audio yet. Syllables get the word's time split by character count, though timing from the model's phone boundaries looked slightly better: 0.10 s vs 0.13 s mean error on 255 syllables.
  • Lyrics-specific kanji readings without a transliteration: a kanji sung with an unusual reading still gets the dictionary's reading unless the line has a transliteration. Adding one in the Languages tab fixes it.
  • English words missing from the dictionary: a few common lyric spellings are built in (woah, tryna, imma, finna, cuz). Other missing words still make their line fall back to the char split.
  • Cantonese: the published model has no Cantonese, so it isn't supported.

Testing

  • pnpm typecheck, biome lint, biome format, knip and pnpm build all pass.
  • pnpm test:unit passes: 4,304 tests, including the decoder (pinned to Python's output), G2P (real kuromoji and pinyin-pro, plus kana and pinyin coverage against the model's dictionaries), romaji and pinyin parsing, the rule for when a transliteration is used, matching transliterations to characters (including ENEMY's mispaired line), CJK segmentation, lexicon, window and re-align tests, and store tests (one undo step, unknown-word fallback, download consent, language and vocals gating, switch to Word mode).
  • In headless Chrome:
    • the full flow (Sync tab → Auto-align → model download → align 37 lines → Undo toast)
    • Japanese, Mandarin and mixed lines in the real browser worker, with kuromoji loading from the bucket
    • Korean falling back as intended
  • Not run locally: pnpm test:component. Playwright's Chromium download times out on this machine, so CI is the first run for the browser tests.

🤖 Generated with Claude Code

Tap each line's start and end in Line mode, then Auto-align in the Sync tab
times every word from the separated vocals. It runs HubertFA v0.0.7 (ONNX,
Apache-2.0) in a worker on WebGPU or WASM, with a TypeScript port of its
Viterbi decoder that matches the Python reference exactly.

- Auto-align popover in the Sync header, laid out like vocal separation:
  model download and alignment progress, cancel, retry, and steps for
  tapping lines and separating vocals. Switches to Word mode afterwards.
- Re-align lines that already have word timing, keeping each part's text,
  explicit flag, syllable group and transliteration.
- Timeline context menu: Auto-align / Re-align for selected lines.
- Lines with unknown words or no alignment path fall back to the
  character split and are reported; non-English lyrics are refused.
- Model and dictionary are fetched from the vocal model host and cached.
  scripts/build-hubertfa-onnx.sh removes the graph's fp16 Cast round-trips
  (WebGPU without shader-f16 can't run them); scripts/upload-hubertfa.sh
  uploads via R2's S3 API since the model is over wrangler's limit.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

React Doctor found 1 new issue in 1 file · 1 warning · score 92 / 100 (Great) · 0 fixed · vs master

1 warning

src/stores/alignment.ts

  • ⚠️ L240 await inside a loop async-await-in-loop

Reviewed by React Doctor for commit 05db14a. See inline comments for fixes.

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
composer 05db14a Commit Preview URL

Branch Preview URL
Oct 07 2026, 04:25 AM

@adaliea
adaliea added this pull request to stack #265 October 7, 2026 01:15
@adaliea
adaliea marked this pull request as draft October 7, 2026 01:15
HubertFA's model covers English, Mandarin and Japanese in one vocabulary, so
Auto-align now handles all three, mixed freely within a line ("君と dance").

- Unspaced Chinese/Japanese is split into one part per character (small
  kana, ー and っ stay with the kana before), and each part gets its own
  timing from the model rather than a character-count split.
- Mandarin: per-character toneless pinyin from pinyin-pro (polyphones read
  in context), looked up in HubertFA's pinyin dictionary.
- Japanese: kana map to the dictionary's romaji morae; kanji readings come
  from kuromoji (IPADIC), whose 18 MB dictionary is hosted next to the
  model, loaded only for Japanese lyrics, and cached like the model.
- Han characters read as kanji in Japanese projects or when any line has
  kana, otherwise as Mandarin. Lines in other scripts fall back to the
  character split; other project languages are refused.
- Model downloads still load when the browser can't cache them.

On real songs with simulated rough line taps, median word-start error went
from 441 to 84 ms (Mandarin) and 439 to 103 ms (Japanese).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@adaliea adaliea changed the title feat: auto-align word timing from tapped line timing (HubertFA) feat: auto-align word timing from tapped line timing (HubertFA, en/zh/ja) Oct 7, 2026
@socket-security

socket-security Bot commented Oct 7, 2026 •

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Added@​sglkc/​kuromoji@​1.1.0811001008090
Addedpinyin-pro@​3.29.410010010093100

View full report

When a line has a transliteration that differs from the line itself, its
syllables are what Auto-align sings. The dictionaries (kuromoji, pinyin-pro)
now only decide which characters each syllable belongs to: an edit-distance
alignment between the two readings, with English words as anchors, hands
each character the transliteration's syllables. So 運命 sung as "sadame"
and kanji the tokenizer can't read come out right. Lines whose
transliteration is just the line again (English lines), stale tracks and
lines without one keep using the dictionaries.

Per-part transliterations aren't trusted alone: imported lyrics often pair
them with the wrong characters (ENEMY: 完 → "ka", 璧 → "n") while the line
as a whole reads correctly, so the whole-line text is always used.

Also reads common lyric spellings the English dictionary lacks (woah,
tryna, imma, finna, cuz).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Comment thread src/stores/alignment.ts
signal.throwIfAborted();
const { line, taps, words, existing, previous } = plan;
const window = alignmentWindow(taps, { ...DEFAULT_WINDOW_OPTIONS, duration, previous });
const result = await aligner.align(

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

React Doctor · react-doctor/async-await-in-loop (warning)

This for…of loop waits before starting the next iteration. If iterations perform independent asynchronous work, consider bounded concurrency; await syntax alone does not establish a speedup.

Fix → Consider concurrent calls only for independent asynchronous work. Shared queues or synchronous work may not benefit. Preserve resource limits, transaction ordering, and failure/cancellation semantics; observe all callback promises.

Docs

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant