Skip to content

feat(agent): add audio mode to lk agent debugger - #1000

Open
u9g wants to merge 3 commits into
mainfrom
jason/debugger-audio
Open

u9g wants to merge 3 commits into
mainfrom
jason/debugger-audio

Conversation

@u9g

@u9g u9g commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

lk agent debugger drives an agent in text mode: STT and TTS are off and each say goes over RunInput. That makes it cheap, but it cannot exercise anything that happens in audio: a misheard name ("Sara" transcribed as "Sarah"), spelled-out digits, endpointing, silence after the caller stops, and realtime models such as GPT Live that only take audio at all.

This adds an audio mode. lk agent debugger start --audio speaks every say in that session with LiveKit Inference TTS into the agent's microphone input, so the agent runs its full audio pipeline. Text mode stays the default and is unchanged.

What audio mode does

  • Mic input behaves like a live microphone. The daemon sends a continuous real-time 48 kHz stream to the agent: queued speech when there is some, silence otherwise, so VAD and endpointing see what they would on a call.
  • Agent audio plays out in real time, with no speaker. A playback flush is acknowledged only once the audio before it would have finished playing, and a clear cancels that. Acknowledging immediately would tell the agent it finished speaking long before a real caller heard it, which breaks speaking-state timing and interruptions.
  • A turn ends when the agent has heard the user and answered. The turn waits until the agent has transcribed the speech, responded, and stayed quiet for 3s. A turn the agent hears but never answers is reported as silent.
  • say shows what the agent heard. The user line is the agent's transcript. When its words differ from the text sent, the sent text is shown under it, and --json carries both (text, heard):
    ● You
      My order number is four eight one five.
      (sent: "My order number is 4 8 1 5.")
    
  • Credentials. The TTS uses project credentials resolved like other lk commands (--project, LIVEKIT_*, livekit.toml, default project). They reach the daemon under LK_SESSION_* names, so the agent still reads its own .env. Text mode needs no credentials, as before.
  • The mode is per session. restart keeps it; start and status report it.

Known limitation

With GPT Live's default delegation, the agent stays listening between its acknowledgement ("Okay, checking that.") and the delegated answer, and no session event marks the delegation as in progress. The 3s quiet window covers typical pauses; a longer one would end the turn early, and the answer would show up in the next turn. The robust fix belongs in the agents SDK: report the agent as busy while a delegation runs.

Commits

  1. Pure refactor: rename textSession to agentSession and extract Listen's settle loop into awaitQuiet.
  2. The audio mode itself.

Testing

  • Unit tests cover both modes against a fake agent, under -race: mic input, playout-timed flush acks, clear cancelling an ack, waiting for the agent to hear the user, silent turns, handoff introductions, timeouts.
  • TestSessionE2E now runs the full start/say/history/stop lifecycle in text and audio mode. It passes locally against both the Python and Node echo agents (on a staging project). Both echo agents gained Inference STT and TTS for this.

… awaitQuiet

Pure refactor ahead of audio mode: the session type is no longer text-specific,
and Listen's settle loop becomes a helper a turn can reuse after its speech.
@u9g
u9g requested a review from bcherry September 25, 2026 15:58
@u9g

u9g commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

Demo: the same say in a text session, then with start --audio, where the agent hears "four eight one five" and "Sarah Lynn" and the sent text is shown next to what it heard; --json carries both, and status reports the mode.

debugger-audio.mp4

`lk agent debugger start --audio` speaks each `say` with LiveKit Inference TTS
into the agent's microphone input instead of sending text over RunInput, so
the agent runs its full audio pipeline (STT, turn detection, TTS, or a
realtime model like GPT Live, which only takes audio). Text mode stays the
default and is unchanged; the mode is per session, kept across restart, and
shown by start and status.

In audio mode:
- The mic input is a continuous real-time 48 kHz stream: queued speech, else
  silence, so VAD and endpointing behave as with a live microphone.
- Agent audio plays out in real time with no speaker; a playback flush is
  acknowledged only once that audio would have finished, and a clear cancels
  it, so agent speaking time and interruptions match a real call.
- A turn ends once the agent has transcribed the user, responded, and stayed
  quiet; a turn it hears but never answers is reported silent.
- `say` reports what the agent heard (`heard` in --json) and shows the sent
  text alongside when the words differ (Sara vs Sarah, 4 vs four).
- Project credentials are resolved like other lk commands and passed to the
  daemon under LK_SESSION_* so the agent keeps its own .env.

The echo-agent fixtures gain Inference STT/TTS and the e2e runs both modes.
@u9g
u9g force-pushed the jason/debugger-audio branch from f8d3ce4 to a39fa59 Compare September 25, 2026 16:04
@bcherry

bcherry commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

This is a cool idea. We may want to ensure we warn about it spending inference tokens for TTS. i wonder if it should really just be per-turn also, rather than a flag on start? why not just be able to use a --voice flag on the say command? you could specify a voice descriptor (e.g. cartesia/sonic-3:xyz) or leave blank for a default. if you don't pass it, the say goes in text. also maybe a --file command would be useful, that plays a recorded audio clip. this could be handy for reproducing bugs or testing STT performance. maybe the file playback would be best if it wasn't pre-understood to be a "turn" and instead just dumps audio and lets the agent figure out turn boundaries (if any)

@u9g

u9g commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

This is a cool idea. We may want to ensure we warn about it spending inference tokens for TTS. i wonder if it should really just be per-turn also, rather than a flag on start? why not just be able to use a --voice flag on the say command? you could specify a voice descriptor (e.g. cartesia/sonic-3:xyz) or leave blank for a default. if you don't pass it, the say goes in text. also maybe a --file command would be useful, that plays a recorded audio clip. this could be handy for reproducing bugs or testing STT performance. maybe the file playback would be best if it wasn't pre-understood to be a "turn" and instead just dumps audio and lets the agent figure out turn boundaries (if any)

Warn how? I mean, even when using text mode theres an inference bill to pay. I think the MVP here is fine until there are people that have a real usecase

bcherry commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

maybe just ensure the help text indicates that it uses livekit inference on your connected livekit cloud account

@u9g

u9g commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

maybe just ensure the help text indicates that it uses livekit inference on your connected livekit cloud account

Done

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants