Selective
Auditory
Attention
One decision per utterance: only addressee speech reaches your STT, LLM, and TTS. No wake word required.
Allow camera
and microphone.
Look for the permission prompt in your browser. Both feeds stay on this device.
Choose.
Model
Voice
Warming up.
- Initializing vision model
- Initializing audio model
- Fusing multimodal embeddings
- Calibrating to your room
- Warming up classifier
NATIVE · OPENAI SERVER VAD
SILENT
Faces
—
Voice activity
—
Conv. state
—
00000
Tokens saved