> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.kotoba.tech/s2st/python-sdk/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server. # Python SDK — s2st #### [kotoba-tech/kotoba-python](https://github.com/kotoba-tech/kotoba-python) Source code, releases, and changelog. ## Install ```bash pip install kotoba-sdk ``` Requires Python 3.10 or later. The package is imported as `kotoba`. For live-microphone examples, install with the optional `mic` extra (pulls in `sounddevice`, which needs PortAudio on the system): ```bash pip install 'kotoba-sdk[mic]' ``` ## Configure `KotobaClient()` reads its credentials and per-route URLs from environment variables. For S2ST you need: | Variable | Purpose | | ----------------------- | ------------------------------------------------------- | | `KOTOBA_API_KEY` | Bearer token sent on WS requests | | `KOTOBA_S2ST_EN_JA_URL` | WebSocket URL for English → Japanese speech translation | Or pass them in code: ```python import kotoba client = kotoba.KotobaClient( api_key="kotoba-...", s2st_en_ja_ws_url="wss://.../sts", ) ``` To use other language pairs, register them at runtime: ```python kotoba.register_endpoint("s2st", "en", "ko", "wss://.../sts-en-ko") ``` ## One-shot translation `translate(...)` consumes a finished audio file and returns the translated audio plus the source-side transcript: ```python import kotoba client = kotoba.KotobaClient() result = client.s2st.translate("clip.mp3", src="en", tgt="ja") result.to_wav("translated.wav") print("source transcript:", result.transcript_source) ``` ## Streaming translation Use `client.s2st.stream(...)` for live audio in / live audio out — both transcript deltas and synthesized chunks surface as the server produces them: ```python with client.s2st.stream(src="en", tgt="ja") as session: for chunk in pcm16_chunks_from_mic(): session.send_audio(chunk) session.commit() for event in session: if event.type == "partial_transcript": print(event.text, end="", flush=True) elif event.type == "audio_chunk": speaker.write(event.audio) # float32 PCM @ 24 kHz elif event.type == "done": break ``` ### Tuning latency with `delay` Both `stream(...)` and `translate(...)` accept an optional `delay` parameter — an integer in the range `0`–`20` that controls how many tokens of context the server buffers before emitting translated audio. Higher values give the model more lookahead (better translation quality); lower values reduce latency. Omit it to keep the server default. ```python with client.s2st.stream(src="en", tgt="ja", delay=10) as session: ... result = client.s2st.translate("clip.mp3", src="en", tgt="ja", delay=10) ``` ## Async Both entry points have async equivalents via `AsyncKotobaClient`: ```python import asyncio import kotoba async def main() -> None: async with kotoba.AsyncKotobaClient() as client: result = await client.s2st.translate("clip.mp3", src="en", tgt="ja") result.to_wav("translated.wav") asyncio.run(main()) ``` For a runnable mic demo see `examples/s2st_mic_async.py` in the SDK repo (needs the `mic` extra — `sounddevice` — and PortAudio). ## What's in the box (S2ST) | Symbol | What | | --------------------------------------- | ------------------------------------------- | | `client.s2st.translate(path, src, tgt)` | One-shot file → translated WAV + transcript | | `client.s2st.stream(src, tgt)` | Bi-directional WebSocket session | See the [API reference](/s2st/streaming) for the on-the-wire protocol that this SDK wraps. > Speech-to-speech translation from Python, WebSocket streaming.