> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.kotoba.tech/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server.

# Python SDK — s2t

#### [kotoba-tech/kotoba-python](https://github.com/kotoba-tech/kotoba-python)

Source code, releases, and changelog.

## Install

```bash
pip install kotoba-sdk
```

Requires Python 3.10 or later. The package is imported as `kotoba`.
For live-microphone examples, install with the optional `mic` extra
(pulls in `sounddevice`, which needs PortAudio on the system):

```bash
pip install 'kotoba-sdk[mic]'
```

## Configure

`KotobaClient()` reads its credentials and per-route URLs from environment
variables. For ASR you need:

| Variable              | Purpose                                          |
| --------------------- | ------------------------------------------------ |
| `KOTOBA_API_KEY`      | Bearer token sent on both REST and WS requests   |
| `KOTOBA_ASR_REST_URL` | REST API base URL, e.g. `https://.../v1`         |
| `KOTOBA_ASR_URL`      | WebSocket URL for live ASR, e.g. `wss://.../asr` |

Pass them in code instead if you'd rather not rely on the environment:

```python
import kotoba

client = kotoba.KotobaClient(
    api_key="kotoba-...",
    url="https://.../v1",            # REST base
    asr_ws_url="wss://.../asr",      # WS
)
```

See [Authentication](/overview/authentication) for the full handshake
details.

## Pick a transport

| Transport        | Method                                              | Best for                               |
| ---------------- | --------------------------------------------------- | -------------------------------------- |
| REST             | `client.asr.transcribe(...)` (POST + poll)          | Batch / file-based work, long files    |
| REST (low-level) | `client.asr.submit_job(...)` + `get_job(...)`       | Custom polling, decoupled submit/fetch |
| WebSocket        | `client.asr.stream(...)` / `transcribe_stream(...)` | Live mic, latency-sensitive pipelines  |

## REST batch (`transcribe`)

POSTs the file to `KOTOBA_ASR_REST_URL`, polls until the job is `done`,
and returns the final transcript:

```python
import kotoba

client = kotoba.KotobaClient()
result = client.asr.transcribe("clip.mp3", language="ja", with_timestamps=True)
print(result.text)
for seg in result.segments:
    print(f"{seg.start:6.2f} - {seg.end:6.2f}  {seg.text}")
```

`transcribe()` accepts anything `soundfile` can decode (WAV / FLAC / OGG
/ MP3 / …). When `with_timestamps=True`, `result.segments` is populated
with `Segment(text, start, end)` entries. Polling knobs:

```python
result = client.asr.transcribe(
    "clip.mp3",
    language="ja",
    with_timestamps=True,
    poll_interval=1.0,      # initial GET polling interval (s)
    poll_backoff=1.5,       # multiplied each poll
    max_poll_interval=10.0,
    timeout=1200.0,         # overall deadline for job completion (s)
)
```

`TranscriptionError` is raised on a server-reported failure, `TimeoutError`
if the deadline elapses.

## REST low-level (`submit_job` / `get_job`)

If you'd rather poll yourself — for example to drive a job queue or
surface progress in a UI — call the two REST endpoints directly:

```python
import time

job = client.asr.submit_job("clip.mp3", language="ja")
print("submitted:", job.id)

while True:
    status = client.asr.get_job(job.id)
    if status.state == "done":
        print(status.transcription.text)
        break
    if status.state == "error":
        raise RuntimeError(status.error_message)
    time.sleep(2)
```

`JobStatus.state` is one of `processing | done | error`.

## WebSocket streaming (`transcribe_stream`)

For the realtime / mic case — where transcript deltas should surface
*while* audio is still being captured — pass a generator of PCM16 LE
mono bytes to `transcribe_stream(...)`. The feeder and receiver run
concurrently, so the first delta can fire before your source is
exhausted:

```python
for delta in client.asr.transcribe_stream(mic_chunks(), language="en"):
    print(delta, end="", flush=True)
```

Optional knobs on both `stream(...)` and `transcribe_stream(...)`:

* `language` — `"en"`, `"ja"`, `"ko"`, or `"zh"`
* `sample_rate` — defaults to 24 kHz; the session resamples internally
* `keywords` — list of hotword biases, e.g. `["Kotobatech", "LLM"]`

## Async

Every ASR entry point has an async equivalent via `AsyncKotobaClient`:

```python
import asyncio
import kotoba

async def main() -> None:
    async with kotoba.AsyncKotobaClient() as client:
        result = await client.asr.transcribe("clip.mp3", language="ja")
        print(result.text)

asyncio.run(main())
```

The sync wrapper runs an `asyncio` loop on a background daemon thread —
underlying transport is identical, only the call style differs.

## What's in the box (ASR)

| Symbol                                       | What                                  |
| -------------------------------------------- | ------------------------------------- |
| `client.asr.transcribe(path, ...)`           | REST batch (with optional timestamps) |
| `client.asr.submit_job(...)` / `get_job(id)` | Low-level REST helpers                |
| `client.asr.stream(...)`                     | WebSocket session you drive manually  |
| `client.asr.transcribe_stream(iter, ...)`    | Generator in → transcript deltas out  |

See the API reference for the on-the-wire protocol that this SDK wraps:
[Live (WebSocket)](/s2t/streaming) · [Batch (REST)](/s2t/batch).