> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.kotoba.tech/t2s/python-sdk/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server. # Python SDK — t2s #### [kotoba-tech/kotoba-python](https://github.com/kotoba-tech/kotoba-python) Source code, releases, and changelog. ## Install ```bash pip install kotoba-sdk ``` Requires Python 3.10 or later. The package is imported as `kotoba`. ## Configure `KotobaClient()` reads its credentials and per-route URLs from environment variables. For TTS you need: | Variable | Purpose | | ------------------- | ---------------------------------------------------- | | `KOTOBA_API_KEY` | Bearer token sent on WS requests | | `KOTOBA_TTS_JA_URL` | WebSocket URL for Japanese TTS, e.g. `wss://.../tts` | Or pass them in code: ```python import kotoba client = kotoba.KotobaClient( api_key="kotoba-...", tts_ja_ws_url="wss://.../tts", ) ``` To use other voices / languages, register the route: ```python kotoba.register_endpoint("tts", None, "ko", "wss://.../tts-ko") ``` ## One-shot synthesis ```python import kotoba client = kotoba.KotobaClient() audio = client.tts.synthesize("こんにちは、世界。", language="ja") audio.to_wav("hello.wav") ``` `audio.to_wav()` converts the underlying `float32` 24 kHz mono signal to a playable 16-bit WAV. Available Japanese speakers: `ja-man-m02-azawa` (male, default) and `ja-woman-f04-me` (female). Pass `speaker_id=...` to override: ```python audio = client.tts.synthesize( "こんにちは、世界。", language="ja", speaker_id="ja-woman-f04-me", ) ``` ## Streaming synthesis The full text is sent in a single frame; the server streams the synthesized audio back chunk-by-chunk, so you can play (or pipe to a speaker / WebRTC track) without waiting for the utterance to finish: ```python with client.tts.stream(language="ja") as session: session.synthesize("こんにちは。本日はよろしくお願いします。") for event in session: if event.type == "audio_chunk": handle(event.audio) # float32 PCM @ 24 kHz elif event.type == "done": break ``` `synthesize_stream(...)` flattens the loop when you only want PCM bytes: ```python for pcm in client.tts.synthesize_stream("こんにちは、世界。", language="ja"): speaker.write(pcm) ``` ## Async ```python import asyncio import kotoba async def main() -> None: async with kotoba.AsyncKotobaClient() as client: async with client.tts.stream(language="ja") as session: await session.synthesize("こんにちは。") async for event in session: if event.type == "audio_chunk": await play(event.audio) elif event.type == "done": break asyncio.run(main()) ``` ## What's in the box (TTS) | Symbol | What | | ------------------------------------------ | ------------------------------------ | | `client.tts.synthesize(text, language)` | One-shot synthesis → `AudioResult` | | `client.tts.stream(language)` | Streaming session you drive manually | | `client.tts.synthesize_stream(text, lang)` | Single text → audio-chunk iterator | See the [API reference](/t2s/streaming) for the on-the-wire protocol that this SDK wraps. > Text-to-speech from Python, with one-shot and streaming synthesis.