> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.kotoba.tech/overview/introduction/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server. # Kotoba Technologies APIs > **Note** > > Kotoba APIs are in **private alpha**, available to selected customers. > [Request access](mailto:kotoba_product@kotoba.tech). Kotoba Technologies provides two families of speech APIs: **Realtime APIs** — WebSocket-based services for streaming speech with sub-second latency: * **ASR** — Automatic Speech Recognition. Stream audio and receive transcription deltas live. * **STS** — Speech-to-Speech Translation. Translate spoken audio between languages in real time. * **TTS** — Text-to-Speech. Synthesize natural-sounding speech from text. Each realtime API speaks JSON over a single WebSocket connection and is described by an AsyncAPI specification. **Transcription API** — A REST API for batch and offline workflows: * **Transcription** — Submit an audio file and poll for the result. Simpler than WebSocket when low latency is not required. ## Quickstart Install the Python SDK (Python 3.10+): ```bash pip install kotoba-sdk ``` Batch transcription via the REST API: ```python import kotoba client = kotoba.KotobaClient() # reads KOTOBA_API_KEY + KOTOBA_*_URL from env result = client.asr.transcribe("clip.mp3", language="ja") print(result.text) ``` Streaming TTS: ```python import kotoba client = kotoba.KotobaClient() with client.tts.stream(language="ja") as session: session.start_response() session.append_text("こんにちは、世界。") session.commit() for event in session: if event.type == "audio_chunk": handle(event.audio) elif event.type == "done": break ``` Each capability has its own Python SDK page with end-to-end snippets: [Speech to Text](/s2t/python-sdk) · [Speech to Speech Translation](/s2st/python-sdk) · [Text to Speech](/t2s/python-sdk). ## What's next * [Authentication](/overview/authentication) — server-side and browser flows. * [Audio formats](/overview/audio-formats) — picking the right encoding. * Capabilities: [Speech to Text](/overview/capabilities/speech-to-text), [Speech to Speech Translation](/overview/capabilities/speech-to-speech), [Text to Speech](/overview/capabilities/text-to-speech). * [Async transcription (REST)](/s2t/batch) — submit a file for async transcription (under the ASR tab). > Low-latency speech APIs for real-time and async applications.