You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This adds support for barge-in from live voice LLMs such as Gemini Live and OpenAI GPT Realtime.
For barge-in to work, a the TTS engine must be able to notify the Assist pipeline when a barge-in has occurred. The pipeline must then discard any queued audio immediately - this is both in Home Assistant and in the remote satellite.
This PR is limited just to the ESPHome satellite, Wyoming satellite integration is not included at this time to keep this PR as small as possible and can be easily added later.
This change adds a small, optional interruption path to the streaming TTS API:
TTSAudioRequest provides an interruption callback to the TTS engine. The engine invokes that callback when the live model reports an interruption.
TTS engines opt in via a new supports_audio_interrupt property.
When TTS engine is using interrupt then the cache is disabled. This not only makes the code easier but also it is preferable for a Live LLM model anyway.
Because interruptible playback bypasses the normal audio-conversion path, the engine must advertise and honor the consumer's requested output options. ESPHome currently requests 16 kHz, mono, 16-bit PCM WAV.
ESPHome uses a capability-negotiated flush extension. A capable satellite receives VOICE_ASSISTANT_TTS_STREAM_START with {flush: 1} and must synchronously clear its playback buffer without ending the TTS response.
ESPHome firmware that does not advertise the flush capability retains the existing playback behavior.
Wyoming interruption is intentionally not included because its current AudioStop event ends an audio stream rather than providing safe mid-response flush semantics.
Consumers that do not ever interrupt playback (e.g. media players using /api/tts_proxy URLs, media source playback, VoIP calls, and announcements) see no change in behavor because the new API is optional.
Support for the TTS API is implemented in my Gemini Live integration. The corresponding ESPHome satellite implementation is maintained in esphome-aec.
Note, I have full barge-in running on the bench using ESPHome, but there are serveral PRs that are all needed to get it working. I am trying to keep the size of each PR as small as possible.
Type of change
Dependency upgrade
Bugfix (non-breaking change which fixes an issue)
New integration (thank you!)
New feature (which adds functionality to an existing integration)
Deprecation (breaking change to happen in the future)
Breaking change (fix/feature causing existing functionality to break)
Code quality improvements to existing code or addition of tests
Additional information
This PR fixes or closes issue: fixes #
This PR is related to issue:
Link to documentation pull request:
Link to developer documentation pull request:
Link to frontend pull request:
Checklist
I understand the code I am submitting and can explain how it works.
The code change is tested and works locally.
Local tests pass. Your PR cannot be merged unless tests pass
Hey there @home-assistant/core, mind taking a look at this pull request as it has been labeled with an integration (tts) you are listed as a code owner for? Thanks!
Code owner commands
Code owners of tts can trigger bot actions by commenting:
@home-assistant close Closes the pull request.
@home-assistant mark-draft Mark the pull request as draft.
@home-assistant ready-for-review Remove the draft status from the pull request.
@home-assistant rename Awesome new title Renames the pull request.
@home-assistant reopen Reopen the pull request.
@home-assistant unassign tts Removes the current integration label and assignees on the pull request, add the integration domain after the command.
@home-assistant update-branch Update the pull request branch with the base branch.
@home-assistant add-label needs-more-information Add a label (needs-more-information, problem in dependency, problem in custom component, problem in config, problem in device, feature-request) to the pull request.
@home-assistant remove-label needs-more-information Remove a label (needs-more-information, problem in dependency, problem in custom component, problem in config, problem in device, feature-request) on the pull request.
Hey there @synesthesiam, mind taking a look at this pull request as it has been labeled with an integration (wyoming) you are listed as a code owner for? Thanks!
Code owner commands
Code owners of wyoming can trigger bot actions by commenting:
@home-assistant close Closes the pull request.
@home-assistant mark-draft Mark the pull request as draft.
@home-assistant ready-for-review Remove the draft status from the pull request.
@home-assistant rename Awesome new title Renames the pull request.
@home-assistant reopen Reopen the pull request.
@home-assistant unassign wyoming Removes the current integration label and assignees on the pull request, add the integration domain after the command.
@home-assistant update-branch Update the pull request branch with the base branch.
@home-assistant add-label needs-more-information Add a label (needs-more-information, problem in dependency, problem in custom component, problem in config, problem in device, feature-request) to the pull request.
@home-assistant remove-label needs-more-information Remove a label (needs-more-information, problem in dependency, problem in custom component, problem in config, problem in device, feature-request) on the pull request.
Hey there @jesserockz, @kbx81, @bdraco, mind taking a look at this pull request as it has been labeled with an integration (esphome) you are listed as a code owner for? Thanks!
Code owner commands
Code owners of esphome can trigger bot actions by commenting:
@home-assistant close Closes the pull request.
@home-assistant mark-draft Mark the pull request as draft.
@home-assistant ready-for-review Remove the draft status from the pull request.
@home-assistant rename Awesome new title Renames the pull request.
@home-assistant reopen Reopen the pull request.
@home-assistant unassign esphome Removes the current integration label and assignees on the pull request, add the integration domain after the command.
@home-assistant update-branch Update the pull request branch with the base branch.
@home-assistant add-label needs-more-information Add a label (needs-more-information, problem in dependency, problem in custom component, problem in config, problem in device, feature-request) to the pull request.
@home-assistant remove-label needs-more-information Remove a label (needs-more-information, problem in dependency, problem in custom component, problem in config, problem in device, feature-request) on the pull request.
PR Review — Add support for Barge In to the Assist pipeline
The TTS-side API design is solid, but it is not yet clear that the satellite flush actually works, and the audio format requirement is not enforced.
What works well: the opt-in supports_audio_interrupt flag plus the _stream_claimed single-consumer rule keeps non-interruptible consumers (proxy URLs, media players, VoIP) on the existing cached path. Using aclosing on the live generators is correct. The ESPHome stream_wav change now drops both pending_chunk and partially buffered bytes, which resolves the earlier Copilot concern about stale parser buffers. The TTS tests cover the fallback, claim-conflict and cache-bypass cases well.
The interrupt path skips conversion but only checks the file extension. A 24 kHz live engine on ESPHome (which asks for 16 kHz) fails with a ValueError and plays nothing, and the "must be 16kHz" rule exists only in the PR description.
Stop→start is used as a flush. For Wyoming, a mid-stream AudioStop probably makes the satellite send Played, and HA then calls tts_response_finished() early. For ESPHome, TTS_STREAM_END probably plays out the buffer rather than discarding it (I could not verify this against the firmware). Please point to the device-side support.
The ESPHome interrupt callback has no guard, so a late call after the stream ends sends an unmatched TTS_STREAM_START.
The Wyoming interruptible path duplicates _stream_tts and uses a lock, an event, a generation counter and a task. It could be simplified to a single generation check.
A fallback consumer of a string message starts a second engine session.
When on_audio_interrupt is set, async_generate_tts_audio skips _async_convert_audio entirely. The only check is extension != final_extension. Sample rate, channels and sample width are never checked. If the engine does not list them in supported_options, they are also removed from the options before the engine sees the request (lines 1148-1170), so the engine cannot tell what the consumer needs.
Why it matters:
ESPHome asks for wav/16000 Hz/mono/16-bit (esphome/assist_satellite.py:587-591) and does not pass accept_native_sample_rate.
Gemini Live and OpenAI Realtime both output 24 kHz PCM natively. If an engine forwards that, stream_wav raises ValueError("Expected 16000 Hz, got 24000 Hz"). _stream_tts_audio only logs it, so the user hears nothing for that whole response.
The PR description says interruptible engines "must provide audio in the native 16kHz format". Neither the code nor the supports_audio_interrupt docstring states or enforces this.
Suggested fix, pick one:
Enforce the requirement up front. Raise a clear HomeAssistantError when a requested rate, channel count or width is not in the engine's supported_options.
Or have ESPHome resample the way Wyoming does, passing accept_native_sample_rate=True and converting locally.
Either way, add the format requirement to the supports_audio_interrupt docstring, since it is the only place engine authors will find the API contract.
if on_audio_interrupt is not None:
async with aclosing(data_gen):
if extension != final_extension:
raise HomeAssistantError(
"Interruptible TTS requires the requested audio format"
)
interrupt_playback sends AudioStop and then AudioStart in the middle of a response. In the Wyoming protocol, AudioStop means "end of this audio". It is not a flush command. A satellite usually finishes playing the audio it has buffered and then sends Played.
HA treats Played as the end of the whole TTS response (assist_satellite.py:720-725): it calls tts_response_finished() and sets _played_event_received. After a barge-in, that Played can arrive while the replacement audio is still streaming. The satellite entity would then go idle, and a continue-conversation turn could start, before the answer has finished.
The same concern applies to ESPHome (esphome/assist_satellite.py:759-770). As far as I know, the firmware handles TTS_STREAM_END by setting a stream-ended flag and playing out its ring buffer, not by discarding it. I could not confirm this in this repo, since the firmware lives elsewhere.
Please point to the satellite/firmware behaviour, or companion PR, that turns stop→start into a flush without ending the response. If nothing does yet, the device side needs a dedicated flush or clear event. Without one, the "discard queued audio on the remote satellite" goal in the PR description is not met, and the response state can be wrong.
async def interrupt_playback() -> None:
nonlocal chunk_converter, pending_audio
nonlocal start_time, timestamp, total_seconds
while interrupt_event.is_set():
if (
sample_rate is None
or sample_width is None
or sample_channels is None
):
interrupt_event.clear()
continue
async with write_lock:
interrupt_event.clear()
await client.write_event(AudioStop(timestamp=timestamp).event())
chunk_converter = AudioChunkConverter(
rate=_TTS_SAMPLE_RATE,
width=SAMPLE_WIDTH,
channels=SAMPLE_CHANNELS,
)
pending_audio = b""
timestamp = 0
total_seconds = 0.0
start_time = monotonic()
await client.write_event(
AudioStart(
rate=_TTS_SAMPLE_RATE,
width=SAMPLE_WIDTH,
channels=SAMPLE_CHANNELS,
timestamp=timestamp,
).event()
)
on_audio_interrupt always sends TTS_STREAM_END and then TTS_STREAM_START. Live engines usually call it from their own receive task, not from inside the generator. A late call can land after stream_wav has emitted is_last, or after the finally block has sent the final TTS_STREAM_END. The device then sees an extra TTS_STREAM_START with no end, and may wait for audio that never arrives.
Fix: set a closed flag in the finally block, before aclose(), and return early from the callback when it is set. The Wyoming path already handles this case by cancelling latest_interrupt_task in its finally block.
For a str input, _async_get_cached_result never checks or sets _stream_claimed. The interruptible path sets the claim, but the fallback path does not look at it. If a live consumer (ESPHome or Wyoming) and a URL consumer (/api/tts_proxy, media player) both read the same token, the engine is called twice.
With a Live LLM engine, a second session can produce a different spoken response, which costs money and gives inconsistent output. Streamed input is already protected by _stream_claimed.
The simplest fix is to apply the same single-consumer rule to str inputs. The other option is to document that the duplicate generation is intended.
)
)
@callback
def _async_get_cached_result(self, result: str | AsyncGenerator[str]) -> TTSCache:
"""Return the cached result for non-interruptible playback.
Creates the cached result from the raw input if it does not exist yet.
"""
if self._cached_result is not None:
return self._cached_result
if isinstance(result, str):
self._cached_result = self._manager.async_cache_message_in_memory(
engine=self.engine,
message=result,
use_file_cache=self.use_file_cache,
language=self.language,
options=self.options,
)
return self._cached_result
# A message input stream can only be consumed once, so caching it
# claims it exclusively against interruptible playback.
if self._stream_claimed:
raise HomeAssistantError(
"Interruptible TTS streams can only be consumed once"
)
self._stream_claimed = True
self._cached_result = self._manager.async_cache_message_stream_in_memory(
engine=self.engine,
message_stream=result,
language=self.language,
options=self.options,
)
Checklist
No hardcoded secrets
Contract/format validation at API boundary — warning #1
To rebase and address feedback, mention me:@esphbot rebase critical(fixes 🔴 only), @esphbot rebase important(fixes 🔴 + 🟡), or@esphbot rebase --fixfor all.(A bare@esphbot rebaseonly rebases onto the base branch.)
Risk: The interrupt path can raise new HomeAssistantErrors ('Interruptible TTS requires the requested audio format', 'can only be consumed once', 'engine ... is unavailable'). They escape the background task, so tts_response_finished() and async_set_assist_pipeline_state(False) never run. The task now starts at INTENT_PROGRESS, so _tts_streaming_task is not None, and RUN_END doesn't reset the pipeline state either. The satellite stays in the responding state, and the only trace is an unretrieved task exception.
except ValueError as err:
_LOGGER.error("Error streaming WAV: %s", err)
except asyncio.CancelledError:
return # Don't trigger state change
...
self.tts_response_finished()
self._entry_data.async_set_assist_pipeline_state(False)
Fix: Catch HomeAssistantError (or Exception) next to ValueError, log it, and fall through to the state-change code so the satellite always leaves the responding state.
Risk: For engines that don't support preferred_sample_rate, channels or bytes, those options are popped earlier and ignored here, and conversion is skipped. Only the container extension is checked. ESPHome does not pass accept_native_sample_rate and expects 16 kHz/mono/16-bit, so any engine that outputs another native rate fails in stream_wav with a logged ValueError and the device plays no response.
if on_audio_interrupt is not None:
async with aclosing(data_gen):
if extension != final_extension:
raise HomeAssistantError(...)
async for chunk in data_gen:
yield chunk
return
Fix: Unless accept_native_sample_rate is set, reject the interruptible path when an unsupported sample option was requested (or fall back to the cached/converted path), so the mismatch fails clearly at the source.
Risk: aclose() runs the engine's generator cleanup (network or websocket teardown). If that raises, TTS_STREAM_END is never sent, the device is left waiting in streaming mode, and the original exception is hidden.
finally:
if audio_stream is not None:
await audio_stream.aclose()
self.cli.send_voice_assistant_event(
VoiceAssistantEventType.VOICE_ASSISTANT_TTS_STREAM_END, {}
)
Fix: Send TTS_STREAM_END in an inner finally around aclose(), or wrap aclose() in try/except that logs, so the device is always told the stream ended.
Risk: If aclose() raises, or is interrupted, the _tts_timeout background task is never scheduled, so the satellite never returns to its idle or listening state after the TTS response.
finally:
await audio_stream.aclose()
if latest_interrupt_task is not None:
latest_interrupt_task.cancel()
await asyncio.gather(latest_interrupt_task, return_exceptions=True)
send_duration = monotonic() - start_time
Fix: Protect aclose() with try/finally (or log its exceptions) so the interrupt-task cleanup and the _tts_timeout scheduling always run.
Risk: If interrupt_playback() fails (for example, write_event on a dropped connection) and a second interrupt arrives before the main loop awaits it, the failed task is replaced without its exception ever being read. The gather(return_exceptions=True) in finally also silently drops any failure of the last task, so failed AudioStop/AudioStart writes can go unnoticed.
if latest_interrupt_task is None or latest_interrupt_task.done():
latest_interrupt_task = self.hass.async_create_task(
interrupt_playback()
)
...
await asyncio.gather(latest_interrupt_task, return_exceptions=True)
Fix: Before replacing a done task, check its exception() and re-raise or log it. After the gather in finally, log any non-CancelledError result.
Treat the WAV data section as open-ended after an interruption. data_bytes_remaining still comes from the original header, so subtracting the discarded buffer makes replacement audio longer than the original remainder get truncated; for example, replacing the final 100 bytes of an 800-byte stream with 1024 bytes emits only 100 replacement bytes. The new test masks this by declaring the data size as old_audio + replacement_audio, which a live stream cannot generally know before barge-in.
Cancel an early TTS task when the pipeline reports an error. The streaming input queue receives its end sentinel only on the normal conversation path, so a failure after INTENT_PROGRESS leaves this background task waiting indefinitely; RUN_END then also keeps the Assist pipeline state active because _tts_streaming_task is non-None.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Proposed change
This adds support for barge-in from live voice LLMs such as Gemini Live and OpenAI GPT Realtime.
For barge-in to work, a the TTS engine must be able to notify the Assist pipeline when a barge-in has occurred. The pipeline must then discard any queued audio immediately - this is both in Home Assistant and in the remote satellite.
This PR is limited just to the ESPHome satellite, Wyoming satellite integration is not included at this time to keep this PR as small as possible and can be easily added later.
This change adds a small, optional interruption path to the streaming TTS API:
Consumers that do not ever interrupt playback (e.g. media players using /api/tts_proxy URLs, media source playback, VoIP calls, and announcements) see no change in behavor because the new API is optional.
Support for the TTS API is implemented in my Gemini Live integration. The corresponding ESPHome satellite implementation is maintained in esphome-aec.
Note, I have full barge-in running on the bench using ESPHome, but there are serveral PRs that are all needed to get it working. I am trying to keep the size of each PR as small as possible.
Type of change
Additional information
Checklist
ruff format homeassistant tests)If user exposed functionality or configuration variables are added/changed:
If the code communicates with devices, web services, or third-party tools:
Updated and included derived files by running:
python3 -m script.hassfest.requirements_all.txt.Updated by running
python3 -m script.gen_requirements_all.To help with the load of incoming pull requests: