> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.kotoba.tech/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server.

# Python SDK — s2st

#### [kotoba-tech/kotoba-python](https://github.com/kotoba-tech/kotoba-python)

Source code, releases, and changelog.

## Install

```bash
pip install kotoba-sdk
```

Requires Python 3.10 or later. The package is imported as `kotoba`.
For live-microphone examples, install with the optional `mic` extra
(pulls in `sounddevice`, which needs PortAudio on the system):

```bash
pip install 'kotoba-sdk[mic]'
```

## Configure

`KotobaClient()` reads its credentials and per-route URLs from environment
variables. For S2ST you need:

| Variable                | Purpose                                                 |
| ----------------------- | ------------------------------------------------------- |
| `KOTOBA_API_KEY`        | Bearer token sent on WS requests                        |
| `KOTOBA_S2ST_EN_JA_URL` | WebSocket URL for English → Japanese speech translation |

Or pass them in code:

```python
import kotoba

client = kotoba.KotobaClient(
    api_key="kotoba-...",
    s2st_en_ja_ws_url="wss://.../sts",
)
```

To use other language pairs, register them at runtime:

```python
kotoba.register_endpoint("s2st", "en", "ko", "wss://.../sts-en-ko")
```

## One-shot translation

`translate(...)` consumes a finished audio file and returns the
translated audio plus the source-side transcript:

```python
import kotoba

client = kotoba.KotobaClient()
result = client.s2st.translate("clip.mp3", src="en", tgt="ja")
result.to_wav("translated.wav")
print("source transcript:", result.transcript_source)
```

## Streaming translation

Use `client.s2st.stream(...)` for live audio in / live audio out — both
transcript deltas and synthesized chunks surface as the server produces
them:

```python
with client.s2st.stream(src="en", tgt="ja") as session:
    for chunk in pcm16_chunks_from_mic():
        session.send_audio(chunk)
    session.commit()

    for event in session:
        if event.type == "partial_transcript":
            print(event.text, end="", flush=True)
        elif event.type == "audio_chunk":
            speaker.write(event.audio)        # float32 PCM @ 24 kHz
        elif event.type == "done":
            break
```

### Tuning latency with `delay`

Both `stream(...)` and `translate(...)` accept an optional `delay`
parameter — an integer in the range `0`–`20` that controls how many
tokens of context the server buffers before emitting translated audio.
Higher values give the model more lookahead (better translation
quality); lower values reduce latency. Omit it to keep the server
default.

```python
with client.s2st.stream(src="en", tgt="ja", delay=10) as session:
    ...

result = client.s2st.translate("clip.mp3", src="en", tgt="ja", delay=10)
```

## Async

Both entry points have async equivalents via `AsyncKotobaClient`:

```python
import asyncio
import kotoba

async def main() -> None:
    async with kotoba.AsyncKotobaClient() as client:
        result = await client.s2st.translate("clip.mp3", src="en", tgt="ja")
        result.to_wav("translated.wav")

asyncio.run(main())
```

For a runnable mic demo see `examples/s2st_mic_async.py` in the SDK repo
(needs the `mic` extra — `sounddevice` — and PortAudio).

## What's in the box (S2ST)

| Symbol                                  | What                                        |
| --------------------------------------- | ------------------------------------------- |
| `client.s2st.translate(path, src, tgt)` | One-shot file → translated WAV + transcript |
| `client.s2st.stream(src, tgt)`          | Bi-directional WebSocket session            |

See the [API reference](/s2st/streaming) for the on-the-wire protocol that this
SDK wraps.