Open the app

TensorX Speech API

Live and file transcription on open models. POC environment.

Basics

Models

GET /v1/models lists models with lines in use and status.

Model IDLanguagesText behind speech
nemotron-3.5-asr-streaming40 languages~0.3 s (default)
nemotron-speech-streaming-enEnglish~0.3 s
parakeet-unified-enEnglish~1.1 s
voxtral-mini-4b-realtime13 languages~0.8 s
kyutai-stt-2.6b-enEnglish~2.5 s by design
kyutai-stt-1b-en-frEnglish, French~1 s

Live: OpenAI Realtime-compatible

For clients that already use OpenAI or Azure OpenAI realtime transcription. Change the URL and the key; the events are the same.

wss://speech.inferix.ai/v1/realtime?model=nemotron-3.5-asr-streaming
import asyncio, base64, json, websockets

async def main():
    url = "wss://speech.inferix.ai/v1/realtime?model=nemotron-3.5-asr-streaming"
    async with websockets.connect(url, additional_headers={"Authorization": "Bearer txs_..."}) as ws:
        await ws.send(json.dumps({"type": "transcription_session.update",
                                  "session": {"input_audio_format": "pcm16"}}))
        # send 24 kHz mono PCM16 chunks:
        # await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": base64.b64encode(chunk).decode()}))
        async for msg in ws:
            ev = json.loads(msg)
            if ev["type"].endswith("transcription.completed"):
                print(ev["transcript"])

asyncio.run(main())

Live: TensorX stream (simplest)

wss://speech.inferix.ai/v1/stream?model=nemotron-3.5-asr-streaming&sample_rate=16000&encoding=pcm_s16le&channels=1
{"type": "ready", "call_id": "call_...", "model": "...", "channels": 1}
{"type": "partial", "channel": 0, "text": "I can see there's an outstanding"}
{"type": "final",   "channel": 0, "text": "I can see there's an outstanding balance on your account."}
{"type": "done",    "call_id": "call_...", "usage": {"audio_seconds": 70.2, "first_text_ms": 640, "final_ms": 310}}
{"type": "error",   "code": "model_busy", "message": "..."}

Files

OpenAI-compatible upload. Most formats work: wav, mp3, m4a, ogg/opus, flac.

curl https://speech.inferix.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer txs_..." \
  -F model=nemotron-3.5-asr-streaming \
  -F file=@call.wav \
  -F response_format=json        # or text, verbose_json

Errors

CodeMeaning
unauthorised / invalid_api_keyMissing, revoked or wrong key
unknown_modelThe model ID is not in /v1/models
model_busyAll lines for that model are in use. Retry with backoff
key_limitYour key is at its calls-at-once limit
model_unavailableThe model server is not reachable right now
call_too_longThe call passed the POC length limit