Skip to navigation

Python SDK — t2s

Text-to-speech from Python, with one-shot and streaming synthesis.

Install

pip install kotoba-sdk

Requires Python 3.10 or later. The package is imported as kotoba.

Configure

KotobaClient() reads its credentials and per-route URLs from environment variables. For TTS you need:

VariablePurpose
KOTOBA_API_KEYBearer token sent on WS requests
KOTOBA_TTS_JA_URLWebSocket URL for Japanese TTS, e.g. wss://.../tts

Or pass them in code:

import kotoba
client = kotoba.KotobaClient(
api_key="kotoba-...",
tts_ja_ws_url="wss://.../tts",
)

To use other voices / languages, register the route:

kotoba.register_endpoint("tts", None, "ko", "wss://.../tts-ko")

One-shot synthesis

import kotoba
client = kotoba.KotobaClient()
audio = client.tts.synthesize("こんにちは、世界。", language="ja")
audio.to_wav("hello.wav")

audio.to_wav() converts the underlying float32 24 kHz mono signal to a playable 16-bit WAV.

Available Japanese speakers: ja-man-m02-azawa (male, default) and ja-woman-f04-me (female). Pass speaker_id=... to override:

audio = client.tts.synthesize(
"こんにちは、世界。",
language="ja",
speaker_id="ja-woman-f04-me",
)

Streaming synthesis

The full text is sent in a single frame; the server streams the synthesized audio back chunk-by-chunk, so you can play (or pipe to a speaker / WebRTC track) without waiting for the utterance to finish:

with client.tts.stream(language="ja") as session:
session.synthesize("こんにちは。本日はよろしくお願いします。")
for event in session:
if event.type == "audio_chunk":
handle(event.audio) # float32 PCM @ 24 kHz
elif event.type == "done":
break

synthesize_stream(...) flattens the loop when you only want PCM bytes:

for pcm in client.tts.synthesize_stream("こんにちは、世界。", language="ja"):
speaker.write(pcm)

Async

import asyncio
import kotoba
async def main() -> None:
async with kotoba.AsyncKotobaClient() as client:
async with client.tts.stream(language="ja") as session:
await session.synthesize("こんにちは。")
async for event in session:
if event.type == "audio_chunk":
await play(event.audio)
elif event.type == "done":
break
asyncio.run(main())

What’s in the box (TTS)

SymbolWhat
client.tts.synthesize(text, language)One-shot synthesis → AudioResult
client.tts.stream(language)Streaming session you drive manually
client.tts.synthesize_stream(text, lang)Single text → audio-chunk iterator

See the API reference for the on-the-wire protocol that this SDK wraps.