Repository navigation
Conversation
Tap each line's start and end in Line mode, then Auto-align in the Sync tab times every word from the separated vocals. It runs HubertFA v0.0.7 (ONNX, Apache-2.0) in a worker on WebGPU or WASM, with a TypeScript port of its Viterbi decoder that matches the Python reference exactly. - Auto-align popover in the Sync header, laid out like vocal separation: model download and alignment progress, cancel, retry, and steps for tapping lines and separating vocals. Switches to Word mode afterwards. - Re-align lines that already have word timing, keeping each part's text, explicit flag, syllable group and transliteration. - Timeline context menu: Auto-align / Re-align for selected lines. - Lines with unknown words or no alignment path fall back to the character split and are reported; non-English lyrics are refused. - Model and dictionary are fetched from the vocal model host and cached. scripts/build-hubertfa-onnx.sh removes the graph's fp16 Cast round-trips (WebGPU without shader-f16 can't run them); scripts/upload-hubertfa.sh uploads via R2's S3 API since the model is over wrangler's limit. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
React Doctor found 1 new issue in 1 file · 1 warning · score 92 / 100 (Great) · 0 fixed · vs Reviewed by React Doctor for commit |
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
composer | 05db14a | Commit Preview URL Branch Preview URL |
Oct 07 2026, 04:25 AM |
HubertFA's model covers English, Mandarin and Japanese in one vocabulary, so
Auto-align now handles all three, mixed freely within a line ("君と dance").
- Unspaced Chinese/Japanese is split into one part per character (small
kana, ー and っ stay with the kana before), and each part gets its own
timing from the model rather than a character-count split.
- Mandarin: per-character toneless pinyin from pinyin-pro (polyphones read
in context), looked up in HubertFA's pinyin dictionary.
- Japanese: kana map to the dictionary's romaji morae; kanji readings come
from kuromoji (IPADIC), whose 18 MB dictionary is hosted next to the
model, loaded only for Japanese lyrics, and cached like the model.
- Han characters read as kanji in Japanese projects or when any line has
kana, otherwise as Mandarin. Lines in other scripts fall back to the
character split; other project languages are refused.
- Model downloads still load when the browser can't cache them.
On real songs with simulated rough line taps, median word-start error went
from 441 to 84 ms (Mandarin) and 439 to 103 ms (Japanese).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
When a line has a transliteration that differs from the line itself, its syllables are what Auto-align sings. The dictionaries (kuromoji, pinyin-pro) now only decide which characters each syllable belongs to: an edit-distance alignment between the two readings, with English words as anchors, hands each character the transliteration's syllables. So 運命 sung as "sadame" and kanji the tokenizer can't read come out right. Lines whose transliteration is just the line again (English lines), stale tracks and lines without one keep using the dictionaries. Per-part transliterations aren't trusted alone: imported lyrics often pair them with the wrong characters (ENEMY: 完 → "ka", 璧 → "n") while the line as a whole reads correctly, so the whole-line text is always used. Also reads common lyric spellings the English dictionary lacks (woah, tryna, imma, finna, cuz). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
| signal.throwIfAborted(); | ||
| const { line, taps, words, existing, previous } = plan; | ||
| const window = alignmentWindow(taps, { ...DEFAULT_WINDOW_OPTIONS, duration, previous }); | ||
| const result = await aligner.align( |
There was a problem hiding this comment.
React Doctor · react-doctor/async-await-in-loop (warning)
This for…of loop waits before starting the next iteration. If iterations perform independent asynchronous work, consider bounded concurrency; await syntax alone does not establish a speedup.
Fix → Consider concurrent calls only for independent asynchronous work. Shared queues or synchronous work may not benefit. Preserve resource limits, transaction ordering, and failure/cancellation semantics; observe all callback promises.
What this does
You sync each line's start and end in Line mode (rough taps are fine), then open Auto-align in the Sync tab. It times every word from the separated vocals and switches to Word mode.
ーandっstay with the kana before them, and each part gets its own timing from the model.pinyin-pro(MIT), with characters that have more than one reading resolved from context. It's loaded only when needed.@sglkc/kuromoji, Apache-2.0, IPADIC). Its 18 MB dictionary is hosted next to the model, loaded only for Japanese lyrics, and cached the same way.scripts/build-hubertfa-onnx.shbuilds the model. It removes two fp32→fp16→fp32 Cast round-trips from the graph, because WebGPU withoutshader-f16can't run them; benchmark accuracy was unchanged.scripts/upload-hubertfa.shuploads through R2's S3 API, since the file is over wrangler's 300 MiB limit.Benchmark
Setup:
Word start error
Lines needing no correction (every word within 200 ms)
Lines average about 7 words.
What the settings tests showed
Helped:
Made no difference: HubertFA's padding ensemble, high-pass filtering, loudness normalisation, and the breath detector or post-processing.
Hurt:
Tap accuracy barely matters. With perfect taps HubertFA's mean error goes only from 109 to 98 ms.
Ported versions:
Speed: WASM on one thread takes about 11 s to load the model, then about 2 minutes per song. WebGPU should be much faster; it has been tried on an NVIDIA GPU without
shader-f16, but there's no timing for it yet.Mandarin and Japanese (real songs)
Setup:
ENEMY (TWICE): 63 lines of mixed Japanese and English, with word timing from the file the user supplied.
On this song the transliterations change no readings: kuromoji already reads every word the way the romanization does (完璧 → kanpeki), so the results with and without them are identical. They matter for lyrics with readings that differ from the dictionary's, which the unit tests cover.
No line fell back to the character split. Of every character
pinyin-procan read (20,902 tested), all but 51 rare ones map to a syllable in the model's dictionary. The interjections 嗯 and 呣 are mapped to the nearest syllable.Known limits
Testing
pnpm typecheck,biome lint,biome format,knipandpnpm buildall pass.pnpm test:unitpasses: 4,304 tests, including the decoder (pinned to Python's output), G2P (real kuromoji and pinyin-pro, plus kana and pinyin coverage against the model's dictionaries), romaji and pinyin parsing, the rule for when a transliteration is used, matching transliterations to characters (including ENEMY's mispaired line), CJK segmentation, lexicon, window and re-align tests, and store tests (one undo step, unknown-word fallback, download consent, language and vocals gating, switch to Word mode).pnpm test:component. Playwright's Chromium download times out on this machine, so CI is the first run for the browser tests.🤖 Generated with Claude Code