Repository navigation
fix(models): give back all the memory an idle unload claims to free - #2677
shivanshtalwar0 wants to merge 2 commits into
Conversation
/model/status reported idle while the backend still held ~3 GB of VRAM. unload_shared_model() cleared model_manager.model, but other holders kept the memory: - Cached OmniVoiceBackend adapters (_active_instance for /v1/audio/speech, _ENGINE_INSTANCES for self-test, explicit-engine /generate and worker assignments) stored the model on first generate. They now run on whatever model_manager holds, through acquire_resident_model(), which also touches the idle clock the cached path used to skip. - torch.compile's caches and reduce-overhead CUDA-graph pools are reset on unload, on the compiled-inference thread that captured them (cudagraph trees are thread-local; a reset from another thread asserts). FlashInfer's module context is cleared there too. - wav2vec2 aligners in _ALIGN_CACHE are released by WhisperX/MLX unload and on idle; WhisperX was clearing a per-instance dict nothing filled. - The pyannote pipeline is released on idle.
|
Thank you for contributing to VoiceStudio. Before this pull request can merge, everyone who contributed to it must sign the Contributor License Agreement 1.0 once. You keep your copyright; the agreement lets Yupcha Softwares Private Limited, the company that maintains VoiceStudio, ship your work in both the AGPL-3.0 app and commercial builds. Still to sign: @shivanshtalwar0, @shivanshtalwar00 To sign, post this as a new comment on its own line: Comment |
✨ Finishing Touches 💡 1🛠️ Fix failing CI checks 💡
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Closing: opened against the wrong repository by mistake. |
|
[High risk] Fixes memory leaks in model unload and idle release paths. Prevent idle cleanup from dropping busy aligners and diarization pipelines before merging.
|
| now = time.monotonic() if now is None else now | ||
| if _diar_pipeline is None or now - _diar_last_used < idle_s: | ||
| return False | ||
| logger.info("Idle timeout reached. Releasing the speaker diarization pipeline.") | ||
| return unload_diarization_pipeline() |
There was a problem hiding this comment.
release_idle_diarization_pipeline() drops the cached pipeline while a long dub is still using it, and release_idle_align_models() does the same during long alignment. If another transcription starts after the configured timeout, it loads a second copy while the first job retains the original, increasing memory use and risking an out-of-memory failure. Track active users of both caches and start their idle clocks when the last user finishes.
| if model is None: | ||
| return False | ||
| compiled = getattr(getattr(model, "llm", None), "_orig_mod", None) is not None | ||
| flashinfer = getattr(model, "_fi_graph_cache", None) is not None |
There was a problem hiding this comment.
Checking only _fi_graph_cache skips FlashInfer cleanup after a runtime fallback or RAM offload, because _unapply_flashinfer() removes that attribute without clearing _CTX. The module can still hold the last attention wrapper and GPU tensors, so the new unload path leaves that memory behind. Track outstanding FlashInfer state separately from the current patch, and clear it on the inference thread before flushing.
Summary
After the idle timeout,
GET /model/statusreports{"status":"idle","loaded":false}while the backend still holds the voice model's VRAM (~3 GB seen on an RTX 5090 Docker host, image:stable).unload_shared_model()clearsmodel_manager.model, but other holders keep the memory: the cached engine adapters, torch.compile / CUDA-graph state, FlashInfer's module context, the wav2vec2 aligners, and the pyannote pipeline.Changes
OmniVoiceBackendcached intts_backend._active_instance(/v1/audio/speech) or_ENGINE_INSTANCES(engine self-test, explicit-engine/generate, worker assignments) stored the model on its first generate and kept it. On a normal server nothing sweeps_ENGINE_INSTANCES. The adapter now runs on whatevermodel_managerholds, via a newacquire_resident_model(). Explicit-model per-call views keep their model as before. This covers every place an adapter is cached, including stream sockets and retired engines, without sweeping each one.get_model(), so steady/v1/audio/speechtraffic looked idle.idle_workerunloaded the model under it, the adapter kept generating on the orphan, and the next native generate loaded a second copy.torch._dynamo.reset()releases the code caches and the reduce-overhead CUDA-graph pools. Inductor's cudagraph trees are thread-local: on torch 2.8,reset_cudagraph_trees()from any other thread raisesAssertionErrorand leaves the cached backends in place. So the reset runs on the compiled-inference thread ([Bug] Voice clone: first render is perfect, second render onward has static noise and slow playback (Windows) #315), waits behind any render still on it, and is bounded at 5 s becauseidle_workerunloads on the event loop. An eager model doesn't reset Dynamo, so other engines keep their compile caches._CTX, which holds the last attention workspace, is cleared on that same thread on unload.WhisperXBackend.unload()cleared a per-instance_align_cachethat nothing ever filled, while the real cache,asr_backend._ALIGN_CACHE, kept every loaded aligner for the life of the process. Unload (WhisperX and MLX Whisper) andidle_workernow release the loaded aligners. "No aligner for this language" entries are kept so the next run doesn't probe again.idle_worker.docs/performance.mdanddocs/remote-workers.mdand the aligner notes in the WhisperX / MLX Whisper engine pages now cover these releases.Type
Testing
tests/test_idle_unload_releases_memory.py: 24 tests. 21 fail onmainand pass here. The other 3 guard unchanged behaviour (cold load through the manager, explicit-model views, eager unload leaves Dynamo alone).torch._dynamo.reset()asserts out and keeps the cached Inductor backends, whileunload_shared_model()resets and flushes oncompiled-inferwith zero backends left. Not yet re-measured on a CUDA host.pytest backend/tests/: 472 passed.pytest tests/: 10,131 passed. The 18 failures on the macOS dev machine fail identically onmain. 14 come from the MPS OmniVoice sidecar and pass withOMNIVOICE_DEVICE=cpu, which matches the Linux CI routing.scripts/check_commit_identities.py: clean.Checklist
package.json,pyproject.toml,backend/core/version.py, and lockfilestests/fixtures/omnivoice_data/still loads green on thesmoke-matrixCI job (macOS + Windows + Linux)