> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.kotoba.tech/overview/capabilities/speech-to-text/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server. # Speech to Text Kotoba's Speech-to-Text (ASR) lets you turn audio into text in either of two complementary modes: * **Live (WebSocket)** — push PCM16 chunks as they arrive, read transcript deltas back the same connection. Built for microphones and any latency-sensitive pipeline where the first words matter before the utterance ends. * **Batch (REST)** — POST an audio file and poll until a job completes. Best for long files, offline workflows, and anything that doesn't need partial results. Both modes support English (`en`), Japanese (`ja`), Korean (`ko`), and Chinese (`zh`) input. ## Pick a transport | You want… | Use | | ---------------------------------------- | ----------------------------------- | | Mic in, text out, while audio is flowing | Live (WebSocket) | | File in, full transcript out | Batch (REST) | | Per-segment timestamps | Batch (REST) with `with_timestamps` | ## Where to go next * [Python SDK for ASR](/s2t/python-sdk) — `client.asr.transcribe(...)` for REST batch, `client.asr.transcribe_stream(...)` for live. * [API reference — Live (WebSocket)](/s2t/streaming) — the AsyncAPI spec for the streaming channel. * [API reference — Batch (REST)](/s2t/batch) — the OpenAPI spec for the async transcription job endpoints. > Stream live transcripts or transcribe finished audio files.