Why Orca Streaming Text-to-Speech?
Natural-sounding confirmation at 29 MB peak memory.
128 ms
First-token-to-speech latency (2.6× faster than ElevenLabs Streaming at 335 ms)
29 MB
Peak memory (11× less than the lightest on-device alternative)
2.3×
Less CPU than the most compute-efficient neural on-device TTS
Orca Streaming Text-to-Speech reads each dialer response aloud, such as "Calling Sarah Chen on mobile," "I found Sarah Chen and Sarah Khan. Which one?", or "There is no work number for Sarah Chen. Should I try mobile?". This way the user never has to look at the screen while driving, walking, or operating equipment. Most high-quality TTS engines require hundreds of megabytes of RAM. Orca uses 29 MB peak memory, 10–50× less than any natural-sounding on-device alternative, which fits easily inside a car head unit, a Bluetooth headset firmware image, or a pair of smart glasses. First-token latency is 128 ms, fast enough that spoken confirmations feel conversational.
ElevenLabs TTS Streaming335 ms