Measured on a Bluetooth headset in #127: pressing the trigger and having a live microphone are about 1.4 seconds apart. 670ms goes on the first engine.start, 270ms until the route change fires, then the rebuild and the first buffer. Opening the input makes the headset switch to its hands-free profile, and no amount of app-side work shortens that.
Since #127 nothing lies about it. The begin sound waits for real audio, and the notch no longer shows the dead-microphone amber during the wait. But nothing tells the user either: the shape reveals and looks exactly as it does when it is listening, so anyone who starts talking straight away loses their first words.
A brief state for it would be honest, shown only while a session has yet to receive a buffer, and in practice only ever seen on Bluetooth since a built-in microphone delivers one almost at once. It should not become another reveal animation: the shape is already up by then, and the session is a second from starting.
Decided by feel, so it wants trying rather than specifying: a dimmed voice visual, a quiet caption, or simply holding the reveal until audio arrives are all plausible and only one of them will look right.
Measured on a Bluetooth headset in #127: pressing the trigger and having a live microphone are about 1.4 seconds apart. 670ms goes on the first
engine.start, 270ms until the route change fires, then the rebuild and the first buffer. Opening the input makes the headset switch to its hands-free profile, and no amount of app-side work shortens that.Since #127 nothing lies about it. The begin sound waits for real audio, and the notch no longer shows the dead-microphone amber during the wait. But nothing tells the user either: the shape reveals and looks exactly as it does when it is listening, so anyone who starts talking straight away loses their first words.
A brief state for it would be honest, shown only while a session has yet to receive a buffer, and in practice only ever seen on Bluetooth since a built-in microphone delivers one almost at once. It should not become another reveal animation: the shape is already up by then, and the session is a second from starting.
Decided by feel, so it wants trying rather than specifying: a dimmed voice visual, a quiet caption, or simply holding the reveal until audio arrives are all plausible and only one of them will look right.