VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Abstract
VoiceMem introduces a dual-brain streaming memory architecture for speech language models that improves retrieval accuracy, emotional personalization, and real-time efficiency.
Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.
Community
Memory foundation for real-time voice interaction.
Project page: https://github.com/xzf-thu/VoiceMem
Code: https://github.com/xzf-thu/VoiceMem
Model: https://huggingface.co/zhifeixie/VoiceMem_MF_Qwen3_6_35B_A3B_Qlora
Dataset: https://huggingface.co/datasets/zhifeixie/VoiceMem-ChatMem400k
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue (2026)
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation (2026)
- LightMem-Ego: Your AI Memory for Everyday Life (2026)
- TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents (2026)
- Light-Omni: Reflex over Reasoning in Agentic Video Understanding with Long-Term Memory (2026)
- JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents (2026)
- HERO: Human-profile Enhanced Retrieval Optimization Framework for Long-term Agent Memory (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.26005 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper