> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.kotoba.tech/overview/audio-formats/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server. # Audio formats The ASR, STS, and TTS APIs all speak JSON over WebSocket, and audio payloads are Base64-encoded inside event bodies. ## Input audio (ASR / STS) | Format | Description | | ---------- | --------------------------------------- | | `pcm16` | 16-bit signed little-endian PCM | | `float32` | 32-bit IEEE float PCM | | `twilio` | μ-law 8 kHz mono (Twilio Media Streams) | | `ogg/opus` | Ogg-wrapped Opus | Sample rate defaults to **24,000 Hz** and mono (`1` channel). You can change both via the session update event. ## Output audio (STS / TTS) | Format | Description | | --------- | --------------------------------------- | | `pcm16` | 16-bit signed little-endian PCM | | `float32` | 32-bit IEEE float PCM | | `twilio` | μ-law 8 kHz mono (Twilio Media Streams) | ## Chunk size Each `input_audio_buffer.append` event can carry up to **1 MiB** (approximately 10 seconds at the default sample rate). For most real-time applications, 20–40 ms chunks work well. > Supported input and output encodings.