Repository navigation
Conversation
…aration
Backend selection (backend-selection.ts)
- Probe WebGPU with requestAdapter({ powerPreference: "high-performance" })
and use WASM when there is no adapter or only a software adapter
(e.g. SwiftShader on Linux Chrome without Vulkan), which previously
looked like a silent hang. Also fall back to WASM if the WebGPU runtime
or session creation fails. The chosen path and the reason are logged.
- Derive the CDN wasmPaths version from the installed onnxruntime-web
instead of a hardcoded 1.26.0 that had drifted from the bundled 1.30.0.
Failure handling (worker-host.ts)
- Reject in-flight init/process on worker error/messageerror and tear the
worker down; add an inactivity timeout to init() so a hung session
create surfaces in the UI. Non-Error throws from ORT get a readable
message.
Performance
- Cache FFT twiddles, bit-reversal table and Hann window; STFT/iSTFT is
bit-identical but ~3.4x faster.
- Pipeline CPU pre/post-processing with the in-flight inference call
(chunk-pipeline.ts) so the GPU is not idle between chunks.
- Log a per-run timing summary (CPU busy vs waiting on inference). An
opt-in localStorage["composer.profileSeparation"] = "1" flag attempts a
one-chunk GPU kernel profile (gpu-profile.ts).
worker.ts is split into focused modules (ort-runtime, backend-selection,
chunk-pipeline, gpu-profile, worker-log) to stay under the 300-line
file-size budget.
|
React Doctor found no new issues. 🎉 Reviewed by React Doctor for commit |
Deploying with
|
| Status | Name | Latest Commit | Preview URL | Updated (UTC) |
|---|---|---|---|---|
| ✅ Deployment successful! View logs |
composer | 65feeee | Commit Preview URL Branch Preview URL |
Oct 06 2026, 10:43 PM |
The fp16 model's WebGPU kernels need the `shader-f16` device feature. On adapters without it (e.g. RTX 4070 on Linux Chrome, which does not expose the feature) session creation fails, and the WASM fallback then runs fp16 at roughly 60x slower than fp32 on WebGPU (85 s for 10 s of audio), which looks like the model silently not working. Choose the backend before downloading the model and fail fast, with a message pointing at the fp32 setting, when fp16 is selected on a WebGPU device without shader-f16. fp16 on the WASM path is allowed but logs a warning. The precision setting description now states the requirement.
The pipeline starts run(i+1) before extracting chunk i, so cancelling or throwing (bad output tensor, extraction error) left a session.run still executing when the worker reported back. ORT sessions don't allow concurrent runs, so a following run on the same session could overlap it. Wrap the loop in try/finally that awaits the in-flight run (ignoring its result) and switches GPU profiling off, so the session is idle on exit. Adds tests with a fake session covering the cancel and error paths.
adaliea
added this pull request to stack #265
October 7, 2026 01:15
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
On Linux Chrome without Vulkan,
navigator.gpuexists andrequestAdapter()succeeds with a software adapter (SwiftShader). The vocal-separation model then runs on the CPU through WebGPU and effectively never finishes. There is no error, only a stalled progress bar and pegged CPU/RAM. Separately, even with a real GPU the pipeline left it mostly idle.Changes
Backend selection (
backend-selection.ts,ort-runtime.ts)powerPreference: "high-performance"(and set ORT'senv.webgpu.powerPreferenceto match).isFallbackAdapter) adapter. Also fall back if the WebGPU runtime fails to load orInferenceSession.createfails.using WebGPU (...)/using WASM (CPU): ...).wasmPathsversion was hardcoded to1.26.0while the bundled package is1.30.0. It is now derived from the installedonnxruntime-webviaimport.meta.env.VITE_ORT_VERSION(scripts/ort-version.ts,vite.config.ts,vitest.config.ts).Failure handling (
worker-host.ts)error/messageerrornow reject the in-flightinit()/process()and tear the worker down.init()fails after 120 s of worker inactivity (reset on download progress), so a hung session create reaches the UI.Errorthrows from ORT get a readable message instead ofundefined.fp16 model (
backend-selection.ts)shader-f16device feature. Some adapters do not expose it (an RTX 4070 on Linux Chrome 154 reportsUnsupported feature: shader-f16). Session creation then fails withProgram Transpose requires f16 but the device does not support it, and the WASM fallback ran fp16 about 60x slower than fp32 on WebGPU (85 s vs 1.4 s for 10 s of audio), which looked like fp16 silently not working.shader-f16fails immediately with a message pointing at the fp32 setting, so nothing is downloaded. fp16 on the WASM path still works but logs a warning (about 4x slower than fp32 on WASM).shader-f16. I had no such GPU available.Performance (
stft.ts,chunk-pipeline.ts)session.runis ever in flight.localStorage["composer.profileSeparation"] = "1"attempts a one-chunk GPU kernel profile. See notes below.worker.tsis split into focused modules to stay under the 300-line file-size budget.Verification
pnpm typecheck, Biome on changed files,pnpm jscpdpass.pnpm test:unit: everything passes except 4lrclibtests that hit the live LRCLib API, which are unrelated. The file-size budget test passes. I ran the unit tests locally withNODE_OPTIONS=--no-experimental-webstoragebecause of a localStorage issue with Node 25 in my environment. I did not run the Playwright component tests.Notes
timestamp-querysupport on that adapter but delivered no timestamps, so it currently only logs that fact. It is off by default and could be removed if you would rather not carry it.