
VoiceMem brings streaming memory to voice agents
VoiceMem is an Apache-2.0 memory layer for voice agents that separates factual recall from emotional and persona context. Its streaming architecture retrieves a small set of relevant memories while users are still speaking, targeting personalization without added conversational delay.
VoiceMem makes a strong case that voice-agent memory needs both context discipline and emotional continuity, though its benchmark claims still need independent reproduction.
- –Dual-brain storage separates durable facts from evolving feelings, preferences, and relationships
- –Streaming retrieval can hide memory latency inside the user’s speaking turn
- –Top-five retrieval and roughly 430 memory tokens point toward leaner, more controllable agent contexts
- –Multimodal ingestion covers speech, speakers, sound events, and multi-party conversations
- –The modular, open-source design should make it easier to swap memory engines and voice models
DISCOVERED
1h ago
2026-08-30
PUBLISHED
1h ago
2026-08-30
RELEVANCE
AUTHOR
AI Search