Skip to content

fix: fall back to WASM when WebGPU is unusable and speed up vocal separation - #262

Open
adaliea wants to merge 4 commits into
masterfrom
t3code/debug-webgpu-initialization
Open

adaliea wants to merge 4 commits into
masterfrom
t3code/debug-webgpu-initialization

Conversation

@adaliea

@adaliea adaliea commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Problem

On Linux Chrome without Vulkan, navigator.gpu exists and requestAdapter() succeeds with a software adapter (SwiftShader). The vocal-separation model then runs on the CPU through WebGPU and effectively never finishes. There is no error, only a stalled progress bar and pegged CPU/RAM. Separately, even with a real GPU the pipeline left it mostly idle.

Changes

Backend selection (backend-selection.ts, ort-runtime.ts)

  • Request the adapter with powerPreference: "high-performance" (and set ORT's env.webgpu.powerPreference to match).
  • Use WASM when there is no WebGPU, no adapter, or only a software (isFallbackAdapter) adapter. Also fall back if the WebGPU runtime fails to load or InferenceSession.create fails.
  • The chosen backend and the reason are logged (using WebGPU (...) / using WASM (CPU): ...).
  • The CDN wasmPaths version was hardcoded to 1.26.0 while the bundled package is 1.30.0. It is now derived from the installed onnxruntime-web via import.meta.env.VITE_ORT_VERSION (scripts/ort-version.ts, vite.config.ts, vitest.config.ts).

Failure handling (worker-host.ts)

  • Worker error/messageerror now reject the in-flight init()/process() and tear the worker down.
  • init() fails after 120 s of worker inactivity (reset on download progress), so a hung session create reaches the UI.
  • Non-Error throws from ORT get a readable message instead of undefined.

fp16 model (backend-selection.ts)

  • The fp16 model's WebGPU kernels need the shader-f16 device feature. Some adapters do not expose it (an RTX 4070 on Linux Chrome 154 reports Unsupported feature: shader-f16). Session creation then fails with Program Transpose requires f16 but the device does not support it, and the WASM fallback ran fp16 about 60x slower than fp32 on WebGPU (85 s vs 1.4 s for 10 s of audio), which looked like fp16 silently not working.
  • The backend is now chosen before the download. Selecting fp16 on a WebGPU device without shader-f16 fails immediately with a message pointing at the fp32 setting, so nothing is downloaded. fp16 on the WASM path still works but logs a warning (about 4x slower than fp32 on WASM).
  • The precision setting description now states the requirement.
  • Not verified: fp16 on a device that does expose shader-f16. I had no such GPU available.

Performance (stft.ts, chunk-pipeline.ts)

  • Cache FFT twiddles, bit-reversal table and Hann window. STFT/iSTFT output is bit-identical to before (max diff 0) and about 3.4x faster (about 657 ms to 192 ms of FFT work per chunk in Node).
  • Pipeline the CPU pre/post-processing with the in-flight inference call so the GPU is not idle between chunks. Only one session.run is ever in flight.
  • Log a one-line timing summary per run (CPU busy vs waiting on inference).
  • Opt-in localStorage["composer.profileSeparation"] = "1" attempts a one-chunk GPU kernel profile. See notes below.

worker.ts is split into focused modules to stay under the 300-line file-size budget.

Verification

  • pnpm typecheck, Biome on changed files, pnpm jscpd pass.
  • pnpm test:unit: everything passes except 4 lrclib tests that hit the live LRCLib API, which are unrelated. The file-size budget test passes. I ran the unit tests locally with NODE_OPTIONS=--no-experimental-webstorage because of a localStorage issue with Node 25 in my environment. I did not run the Playwright component tests.
  • Sequential vs pipelined loops produce bit-identical vocals on the same input (compared in-browser, WASM path).
  • WASM fallback confirmed in a browser with no adapter, and with a SwiftShader software adapter (reported as a fallback adapter).
  • WebGPU on an RTX 4070 Laptop GPU (Chrome, Linux, Vulkan enabled): a 170 s file separates in about 17 s (29 chunks, about 0.6 s/chunk), and a 10 s file in 2.2 s with finite output.

Notes

  • A two-worker experiment (two sessions sharing the GPU) was only about 19% faster than one worker, so that is not included. The GPU already appears mostly busy.
  • The GPU kernel profile flag is best-effort: ORT reported timestamp-query support on that adapter but delivered no timestamps, so it currently only logs that fact. It is off by default and could be removed if you would rather not carry it.
  • ORT's standard timestamp mode ends a compute pass per kernel, so the profile is enabled for a single chunk only rather than the whole run.

…aration

Backend selection (backend-selection.ts)
- Probe WebGPU with requestAdapter({ powerPreference: "high-performance" })
  and use WASM when there is no adapter or only a software adapter
  (e.g. SwiftShader on Linux Chrome without Vulkan), which previously
  looked like a silent hang. Also fall back to WASM if the WebGPU runtime
  or session creation fails. The chosen path and the reason are logged.
- Derive the CDN wasmPaths version from the installed onnxruntime-web
  instead of a hardcoded 1.26.0 that had drifted from the bundled 1.30.0.

Failure handling (worker-host.ts)
- Reject in-flight init/process on worker error/messageerror and tear the
  worker down; add an inactivity timeout to init() so a hung session
  create surfaces in the UI. Non-Error throws from ORT get a readable
  message.

Performance
- Cache FFT twiddles, bit-reversal table and Hann window; STFT/iSTFT is
  bit-identical but ~3.4x faster.
- Pipeline CPU pre/post-processing with the in-flight inference call
  (chunk-pipeline.ts) so the GPU is not idle between chunks.
- Log a per-run timing summary (CPU busy vs waiting on inference). An
  opt-in localStorage["composer.profileSeparation"] = "1" flag attempts a
  one-chunk GPU kernel profile (gpu-profile.ts).

worker.ts is split into focused modules (ort-runtime, backend-selection,
chunk-pipeline, gpu-profile, worker-log) to stay under the 300-line
file-size budget.
@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

React Doctor found no new issues. 🎉

Reviewed by React Doctor for commit 65feeee.

Comment thread src/audio/separation/chunk-pipeline.ts Outdated
@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Oct 6, 2026 •

Copy link
Copy Markdown

Deploying with  Cloudflare Workers  Cloudflare Workers

The latest updates on your project. Learn more about integrating Git with Workers.

Status Name Latest Commit Preview URL Updated (UTC)
✅ Deployment successful!
View logs
composer 65feeee Commit Preview URL

Branch Preview URL
Oct 06 2026, 10:43 PM

The fp16 model's WebGPU kernels need the `shader-f16` device feature. On
adapters without it (e.g. RTX 4070 on Linux Chrome, which does not expose
the feature) session creation fails, and the WASM fallback then runs fp16
at roughly 60x slower than fp32 on WebGPU (85 s for 10 s of audio), which
looks like the model silently not working.

Choose the backend before downloading the model and fail fast, with a
message pointing at the fp32 setting, when fp16 is selected on a WebGPU
device without shader-f16. fp16 on the WASM path is allowed but logs a
warning. The precision setting description now states the requirement.
The pipeline starts run(i+1) before extracting chunk i, so cancelling or
throwing (bad output tensor, extraction error) left a session.run still
executing when the worker reported back. ORT sessions don't allow
concurrent runs, so a following run on the same session could overlap it.

Wrap the loop in try/finally that awaits the in-flight run (ignoring its
result) and switches GPU profiling off, so the session is idle on exit.
Adds tests with a fake session covering the cancel and error paths.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant