On-device inference + cloud LLM
AI Devices
Devices built around a 0.5–16 TOPS NPU. Everything up to the model runs locally — wake word, far-field audio, speech recognition, vision — and at the top of that band a small quantised model runs on the device too, so the product still answers with the network down. The full conversational model sits in the cloud. Desktop, in-car, companion and speaker form factors share the same inference and audio stack.
The part that decides the project
Latency, and where you spend it. A voice device is judged on how long it sits silent after you stop talking — under roughly 700 ms it feels like conversation, past 1.5 s it feels broken. With the model in the cloud, most of that budget is already spent before the network is touched: wake-word detection, endpointing, noise suppression and speech recognition all run on the device, and every millisecond they waste comes out of the model’s allowance. Which is why the NPU, the audio front end and the streaming design are decided together, months before anyone can tune them.
What you need to decide before we can quote
- How much runs on-device — wake word, full speech recognition, or a small local model
- Response budget: how long the device may take before the model is called
- Which cloud model sits behind it, and in which region
- What still works with no network — and what you promise the user
What this direction covers
- AI desktop robots
- In-car AI assistants
- Companion robots
- AI speakers
- LLM voice modules
Silicon it runs on
Linux SoC with NPU (0.5–16 TOPS) · WiFi SoC · MCU
CapabilitiesRequest a Quote
Start a ai devices project
Send the market, target price band and certifications. You get a platform recommendation, a cost envelope and the certification path — from an engineer, within one business day.