Live and file transcription on open models. POC environment.
https://speech.inferix.aiAuthorization: Bearer txs_.... WebSocket clients that cannot set headers can add ?api_key=txs_.... The Azure-style api-key header also works.GET /v1/models lists models with lines in use and status.
| Model ID | Languages | Text behind speech |
|---|---|---|
| nemotron-3.5-asr-streaming | 40 languages | ~0.3 s (default) |
| nemotron-speech-streaming-en | English | ~0.3 s |
| parakeet-unified-en | English | ~1.1 s |
| voxtral-mini-4b-realtime | 13 languages | ~0.8 s |
| kyutai-stt-2.6b-en | English | ~2.5 s by design |
| kyutai-stt-1b-en-fr | English, French | ~1 s |
For clients that already use OpenAI or Azure OpenAI realtime transcription. Change the URL and the key; the events are the same.
wss://speech.inferix.ai/v1/realtime?model=nemotron-3.5-asr-streaming
transcription_session.update (beta shape) or session.update (GA shape). Supported input formats: pcm16 24 kHz, g711_ulaw, g711_alaw, or GA audio/pcm with a rate, audio/pcmu, audio/pcma.input_audio_buffer.append. The model finds sentence ends itself, so commit is optional.conversation.item.input_audio_transcription.delta while people talk and ...completed per segment.gpt-4o-transcribe map to our default model.import asyncio, base64, json, websockets
async def main():
url = "wss://speech.inferix.ai/v1/realtime?model=nemotron-3.5-asr-streaming"
async with websockets.connect(url, additional_headers={"Authorization": "Bearer txs_..."}) as ws:
await ws.send(json.dumps({"type": "transcription_session.update",
"session": {"input_audio_format": "pcm16"}}))
# send 24 kHz mono PCM16 chunks:
# await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": base64.b64encode(chunk).decode()}))
async for msg in ws:
ev = json.loads(msg)
if ev["type"].endswith("transcription.completed"):
print(ev["transcript"])
asyncio.run(main())
wss://speech.inferix.ai/v1/stream?model=nemotron-3.5-asr-streaming&sample_rate=16000&encoding=pcm_s16le&channels=1
encoding: pcm_s16le, mulaw or alaw. sample_rate: 8000 to 48000.channels=2 and send interleaved stereo. Every event carries channel 0 or 1.{"type":"stop"} to finish. You get the last text, then done with usage.{"type": "ready", "call_id": "call_...", "model": "...", "channels": 1}
{"type": "partial", "channel": 0, "text": "I can see there's an outstanding"}
{"type": "final", "channel": 0, "text": "I can see there's an outstanding balance on your account."}
{"type": "done", "call_id": "call_...", "usage": {"audio_seconds": 70.2, "first_text_ms": 640, "final_ms": 310}}
{"type": "error", "code": "model_busy", "message": "..."}
OpenAI-compatible upload. Most formats work: wav, mp3, m4a, ogg/opus, flac.
curl https://speech.inferix.ai/v1/audio/transcriptions \
-H "Authorization: Bearer txs_..." \
-F model=nemotron-3.5-asr-streaming \
-F file=@call.wav \
-F response_format=json # or text, verbose_json
| Code | Meaning |
|---|---|
| unauthorised / invalid_api_key | Missing, revoked or wrong key |
| unknown_model | The model ID is not in /v1/models |
| model_busy | All lines for that model are in use. Retry with backoff |
| key_limit | Your key is at its calls-at-once limit |
| model_unavailable | The model server is not reachable right now |
| call_too_long | The call passed the POC length limit |