Voice Agent Evaluation Benchmarks
Public benchmarks for voice agents, speech tools, computer action, ASR, and meeting understanding, with a structured landscape.
Viewer • Updated • 13 • 130 • 2Note Start here: structured, source-linked capability map across public voice-agent benchmarks.
krutrim-ai-labs/VoiceAgentBench
Viewer • Updated • 5.39k • 4.15k • 9Note Speech-based tool tasks, multi-turn interaction, and safety.
MathLLMs/VoiceAssistant-Eval
Viewer • Updated • 10.5k • 6.3k • 12Note Listening, speaking, viewing, multi-turn behavior, and safety.
sierra-research/mu-bench
Viewer • Updated • 4.27k • 366 • 8Note Multilingual customer-service ASR across five locales.
arcada-labs/audio-agent-bench-suite
Updated • 9Note Multi-turn spoken-agent suite; public files were limited at the 2026-09-01 review.
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
Paper • 2510.07978 • PublishedNote VoiceAgentBench paper.
VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
Paper • 2509.22651 • Published • 24Note VoiceAssistant-Eval paper.
τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains
Paper • 2603.13686 • PublishedNote tau-Voice full-duplex benchmark paper.
hlt-lab/voicebench
Viewer • Updated • 20.6k • 7.25k • 16Note Spoken instruction-following, knowledge, reasoning, safety, accent, and limited multi-turn evaluation.
Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
Paper • 2604.22821 • PublishedNote Audio2Tool benchmark paper for evaluating spoken tool use and conversational repair.
VoiceBench: Benchmarking LLM-Based Voice Assistants
Paper • 2410.17196 • Published • 1Note VoiceBench paper covering spoken instruction following, safety, reasoning, and robustness.
RVtech/Audio2Tool
Viewer • Updated • 30.7k • 4.25k • 2Note Spoken tool-call benchmark with eight tiers spanning multi-intent, correction, multi-turn, and overlapping speech.
EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
Paper • 2605.13841 • Published • 78Note EVA-Bench paper on end-to-end task accuracy, speech fidelity, conversation quality, and robustness.
ServiceNow-AI/eva-bench
Viewer • Updated • 213 • 114 • 25Note 213 end-to-end enterprise voice-agent scenarios with accuracy, experience, perturbation, and reliability evaluation.