> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.kotoba.tech/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server.

# Kotoba Technologies APIs

> **Note**
>
> Kotoba APIs are in **private alpha**, available to selected customers.
> [Request access](mailto:kotoba_product@kotoba.tech).

Kotoba Technologies provides two families of speech APIs:

**Realtime APIs** — WebSocket-based services for streaming speech with
sub-second latency:

* **ASR** — Automatic Speech Recognition. Stream audio and receive
  transcription deltas live.
* **STS** — Speech-to-Speech Translation. Translate spoken audio between
  languages in real time.
* **TTS** — Text-to-Speech. Synthesize natural-sounding speech from
  text.

Each realtime API speaks JSON over a single WebSocket connection and is
described by an AsyncAPI specification.

**Transcription API** — A REST API for batch and offline workflows:

* **Transcription** — Submit an audio file and poll for the result.
  Simpler than WebSocket when low latency is not required.

## Quickstart

Install the Python SDK (Python 3.10+):

```bash
pip install kotoba-sdk
```

Batch transcription via the REST API:

```python
import kotoba

client = kotoba.KotobaClient()  # reads KOTOBA_API_KEY + KOTOBA_*_URL from env
result = client.asr.transcribe("clip.mp3", language="ja")
print(result.text)
```

Streaming TTS:

```python
import kotoba

client = kotoba.KotobaClient()
with client.tts.stream(language="ja") as session:
    session.start_response()
    session.append_text("こんにちは、世界。")
    session.commit()
    for event in session:
        if event.type == "audio_chunk":
            handle(event.audio)
        elif event.type == "done":
            break
```

Each capability has its own Python SDK page with end-to-end snippets:
[Speech to Text](/s2t/python-sdk) · [Speech to Speech Translation](/s2st/python-sdk) · [Text to Speech](/t2s/python-sdk).

## What's next

* [Authentication](/overview/authentication) — server-side and browser flows.
* [Audio formats](/overview/audio-formats) — picking the right encoding.
* Capabilities: [Speech to Text](/overview/capabilities/speech-to-text),
  [Speech to Speech Translation](/overview/capabilities/speech-to-speech),
  [Text to Speech](/overview/capabilities/text-to-speech).
* [Async transcription (REST)](/s2t/batch) — submit a file for async transcription (under the ASR tab).