Conversation
… awaitQuiet Pure refactor ahead of audio mode: the session type is no longer text-specific, and Listen's settle loop becomes a helper a turn can reuse after its speech.
|
Demo: the same debugger-audio.mp4 |
`lk agent debugger start --audio` speaks each `say` with LiveKit Inference TTS into the agent's microphone input instead of sending text over RunInput, so the agent runs its full audio pipeline (STT, turn detection, TTS, or a realtime model like GPT Live, which only takes audio). Text mode stays the default and is unchanged; the mode is per session, kept across restart, and shown by start and status. In audio mode: - The mic input is a continuous real-time 48 kHz stream: queued speech, else silence, so VAD and endpointing behave as with a live microphone. - Agent audio plays out in real time with no speaker; a playback flush is acknowledged only once that audio would have finished, and a clear cancels it, so agent speaking time and interruptions match a real call. - A turn ends once the agent has transcribed the user, responded, and stayed quiet; a turn it hears but never answers is reported silent. - `say` reports what the agent heard (`heard` in --json) and shows the sent text alongside when the words differ (Sara vs Sarah, 4 vs four). - Project credentials are resolved like other lk commands and passed to the daemon under LK_SESSION_* so the agent keeps its own .env. The echo-agent fixtures gain Inference STT/TTS and the e2e runs both modes.
f8d3ce4 to
a39fa59
Compare
|
This is a cool idea. We may want to ensure we warn about it spending inference tokens for TTS. i wonder if it should really just be per-turn also, rather than a flag on start? why not just be able to use a |
Warn how? I mean, even when using text mode theres an inference bill to pay. I think the MVP here is fine until there are people that have a real usecase |
|
maybe just ensure the help text indicates that it uses livekit inference on your connected livekit cloud account |
…ected Cloud project
Done |
lk agent debuggerdrives an agent in text mode: STT and TTS are off and eachsaygoes overRunInput. That makes it cheap, but it cannot exercise anything that happens in audio: a misheard name ("Sara" transcribed as "Sarah"), spelled-out digits, endpointing, silence after the caller stops, and realtime models such as GPT Live that only take audio at all.This adds an audio mode.
lk agent debugger start --audiospeaks everysayin that session with LiveKit Inference TTS into the agent's microphone input, so the agent runs its full audio pipeline. Text mode stays the default and is unchanged.What audio mode does
sayshows what the agent heard. The user line is the agent's transcript. When its words differ from the text sent, the sent text is shown under it, and--jsoncarries both (text,heard):lkcommands (--project,LIVEKIT_*,livekit.toml, default project). They reach the daemon underLK_SESSION_*names, so the agent still reads its own.env. Text mode needs no credentials, as before.restartkeeps it;startandstatusreport it.Known limitation
With GPT Live's default delegation, the agent stays
listeningbetween its acknowledgement ("Okay, checking that.") and the delegated answer, and no session event marks the delegation as in progress. The 3s quiet window covers typical pauses; a longer one would end the turn early, and the answer would show up in the next turn. The robust fix belongs in the agents SDK: report the agent as busy while a delegation runs.Commits
textSessiontoagentSessionand extractListen's settle loop intoawaitQuiet.Testing
-race: mic input, playout-timed flush acks, clear cancelling an ack, waiting for the agent to hear the user, silent turns, handoff introductions, timeouts.TestSessionE2Enow runs the full start/say/history/stop lifecycle in text and audio mode. It passes locally against both the Python and Node echo agents (on a staging project). Both echo agents gained Inference STT and TTS for this.