> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.kotoba.tech/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server.

# Python SDK — t2s

#### [kotoba-tech/kotoba-python](https://github.com/kotoba-tech/kotoba-python)

Source code, releases, and changelog.

## Install

```bash
pip install kotoba-sdk
```

Requires Python 3.10 or later. The package is imported as `kotoba`.

## Configure

`KotobaClient()` reads its credentials and per-route URLs from environment
variables. For TTS you need:

| Variable            | Purpose                                              |
| ------------------- | ---------------------------------------------------- |
| `KOTOBA_API_KEY`    | Bearer token sent on WS requests                     |
| `KOTOBA_TTS_JA_URL` | WebSocket URL for Japanese TTS, e.g. `wss://.../tts` |

Or pass them in code:

```python
import kotoba

client = kotoba.KotobaClient(
    api_key="kotoba-...",
    tts_ja_ws_url="wss://.../tts",
)
```

To use other voices / languages, register the route:

```python
kotoba.register_endpoint("tts", None, "ko", "wss://.../tts-ko")
```

## One-shot synthesis

```python
import kotoba

client = kotoba.KotobaClient()
audio = client.tts.synthesize("こんにちは、世界。", language="ja")
audio.to_wav("hello.wav")
```

`audio.to_wav()` converts the underlying `float32` 24 kHz mono signal to
a playable 16-bit WAV.

Available Japanese speakers: `ja-man-m02-azawa` (male, default) and
`ja-woman-f04-me` (female). Pass `speaker_id=...` to override:

```python
audio = client.tts.synthesize(
    "こんにちは、世界。",
    language="ja",
    speaker_id="ja-woman-f04-me",
)
```

## Streaming synthesis

The full text is sent in a single frame; the server streams the
synthesized audio back chunk-by-chunk, so you can play (or pipe to a
speaker / WebRTC track) without waiting for the utterance to finish:

```python
with client.tts.stream(language="ja") as session:
    session.synthesize("こんにちは。本日はよろしくお願いします。")

    for event in session:
        if event.type == "audio_chunk":
            handle(event.audio)               # float32 PCM @ 24 kHz
        elif event.type == "done":
            break
```

`synthesize_stream(...)` flattens the loop when you only want PCM bytes:

```python
for pcm in client.tts.synthesize_stream("こんにちは、世界。", language="ja"):
    speaker.write(pcm)
```

## Async

```python
import asyncio
import kotoba

async def main() -> None:
    async with kotoba.AsyncKotobaClient() as client:
        async with client.tts.stream(language="ja") as session:
            await session.synthesize("こんにちは。")
            async for event in session:
                if event.type == "audio_chunk":
                    await play(event.audio)
                elif event.type == "done":
                    break

asyncio.run(main())
```

## What's in the box (TTS)

| Symbol                                     | What                                 |
| ------------------------------------------ | ------------------------------------ |
| `client.tts.synthesize(text, language)`    | One-shot synthesis → `AudioResult`   |
| `client.tts.stream(language)`              | Streaming session you drive manually |
| `client.tts.synthesize_stream(text, lang)` | Single text → audio-chunk iterator   |

See the [API reference](/t2s/streaming) for the on-the-wire protocol that this
SDK wraps.