Voice AI Agents — a pipeline that listens, looks things up and speaks
Wire listening through speaking in a single CPU pod and measure the latency
고급 · Lessons 30 · Lab 10
Start the lab
Curriculum
Audio basics — PCM, frames, resampling
VAD and turn-taking — end-of-speech detection and barge-in
Streaming ASR — partial results, final results, WER
Streaming LLM responses and time to first token
RAG for spoken queries — rewriting, refusal thresholds, citation checks
Agent workflow — tools, confirmation, retries, handoff
Streaming TTS — sentence splitting and time to first audio
Latency budget — running the pipeline overlapped
Evaluation and observability — quality, failure rate, SLO, gates
Safety — injection and personal data arriving by voice
Reference docs