> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.kotoba.tech/s2t/python-sdk/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.kotoba.tech/_mcp/server. # Python SDK — s2t #### [kotoba-tech/kotoba-python](https://github.com/kotoba-tech/kotoba-python) Source code, releases, and changelog. ## Install ```bash pip install kotoba-sdk ``` Requires Python 3.10 or later. The package is imported as `kotoba`. For live-microphone examples, install with the optional `mic` extra (pulls in `sounddevice`, which needs PortAudio on the system): ```bash pip install 'kotoba-sdk[mic]' ``` ## Configure `KotobaClient()` reads its credentials and per-route URLs from environment variables. For ASR you need: | Variable | Purpose | | --------------------- | ------------------------------------------------ | | `KOTOBA_API_KEY` | Bearer token sent on both REST and WS requests | | `KOTOBA_ASR_REST_URL` | REST API base URL, e.g. `https://.../v1` | | `KOTOBA_ASR_URL` | WebSocket URL for live ASR, e.g. `wss://.../asr` | Pass them in code instead if you'd rather not rely on the environment: ```python import kotoba client = kotoba.KotobaClient( api_key="kotoba-...", url="https://.../v1", # REST base asr_ws_url="wss://.../asr", # WS ) ``` See [Authentication](/overview/authentication) for the full handshake details. ## Pick a transport | Transport | Method | Best for | | ---------------- | --------------------------------------------------- | -------------------------------------- | | REST | `client.asr.transcribe(...)` (POST + poll) | Batch / file-based work, long files | | REST (low-level) | `client.asr.submit_job(...)` + `get_job(...)` | Custom polling, decoupled submit/fetch | | WebSocket | `client.asr.stream(...)` / `transcribe_stream(...)` | Live mic, latency-sensitive pipelines | ## REST batch (`transcribe`) POSTs the file to `KOTOBA_ASR_REST_URL`, polls until the job is `done`, and returns the final transcript: ```python import kotoba client = kotoba.KotobaClient() result = client.asr.transcribe("clip.mp3", language="ja", with_timestamps=True) print(result.text) for seg in result.segments: print(f"{seg.start:6.2f} - {seg.end:6.2f} {seg.text}") ``` `transcribe()` accepts anything `soundfile` can decode (WAV / FLAC / OGG / MP3 / …). When `with_timestamps=True`, `result.segments` is populated with `Segment(text, start, end)` entries. Polling knobs: ```python result = client.asr.transcribe( "clip.mp3", language="ja", with_timestamps=True, poll_interval=1.0, # initial GET polling interval (s) poll_backoff=1.5, # multiplied each poll max_poll_interval=10.0, timeout=1200.0, # overall deadline for job completion (s) ) ``` `TranscriptionError` is raised on a server-reported failure, `TimeoutError` if the deadline elapses. ## REST low-level (`submit_job` / `get_job`) If you'd rather poll yourself — for example to drive a job queue or surface progress in a UI — call the two REST endpoints directly: ```python import time job = client.asr.submit_job("clip.mp3", language="ja") print("submitted:", job.id) while True: status = client.asr.get_job(job.id) if status.state == "done": print(status.transcription.text) break if status.state == "error": raise RuntimeError(status.error_message) time.sleep(2) ``` `JobStatus.state` is one of `processing | done | error`. ## WebSocket streaming (`transcribe_stream`) For the realtime / mic case — where transcript deltas should surface *while* audio is still being captured — pass a generator of PCM16 LE mono bytes to `transcribe_stream(...)`. The feeder and receiver run concurrently, so the first delta can fire before your source is exhausted: ```python for delta in client.asr.transcribe_stream(mic_chunks(), language="en"): print(delta, end="", flush=True) ``` Optional knobs on both `stream(...)` and `transcribe_stream(...)`: * `language` — `"en"`, `"ja"`, `"ko"`, or `"zh"` * `sample_rate` — defaults to 24 kHz; the session resamples internally * `keywords` — list of hotword biases, e.g. `["Kotobatech", "LLM"]` ## Async Every ASR entry point has an async equivalent via `AsyncKotobaClient`: ```python import asyncio import kotoba async def main() -> None: async with kotoba.AsyncKotobaClient() as client: result = await client.asr.transcribe("clip.mp3", language="ja") print(result.text) asyncio.run(main()) ``` The sync wrapper runs an `asyncio` loop on a background daemon thread — underlying transport is identical, only the call style differs. ## What's in the box (ASR) | Symbol | What | | -------------------------------------------- | ------------------------------------- | | `client.asr.transcribe(path, ...)` | REST batch (with optional timestamps) | | `client.asr.submit_job(...)` / `get_job(id)` | Low-level REST helpers | | `client.asr.stream(...)` | WebSocket session you drive manually | | `client.asr.transcribe_stream(iter, ...)` | Generator in → transcript deltas out | See the API reference for the on-the-wire protocol that this SDK wraps: [Live (WebSocket)](/s2t/streaming) · [Batch (REST)](/s2t/batch). > Speech-to-text from Python, REST batch + WebSocket streaming.