Hold the button, speak, let go. Your speech is transcribed, HARI answers, and the reply is spoken back while it is still being written. One WebSocket, everything on SPUR's sovereign Canadian GPUs.
Signed-in users connect with their account session. API users: create a key at /account. Usage is metered like the REST API: STT per second, TTS per character, HARI per token.
wss://ai.spuric.com/v1/realtime?key=YOUR_KEY (or Authorization: Bearer header)
-> {"type":"session.update","voice":"af_heart","speak":true,"language":"en",
"instructions":"You are the front desk assistant for Acme Dental."}
-> {"type":"input_audio.append","audio":"<base64 webm/ogg/wav/pcm16>","format":"webm"}
-> {"type":"input_audio.commit"} or {"type":"input_text","text":"hello"}
<- {"type":"input_audio.transcript","text":"...","seconds":2.4}
<- {"type":"response.text.delta","delta":"Hello"} (streamed)
<- {"type":"response.audio.delta","seq":0,"format":"wav","sample_rate":24000,
"text":"Hello there.","audio":"<base64 wav>"} (per sentence, in order)
<- {"type":"response.done","text":"...","usage":{...},"cost_uusd":123}
Barge-in: send a new input while a response is streaming and the old one is cancelled.
Python: import websockets, json, base64 -> async with websockets.connect(url) as ws: ...