HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 6 days ago • 73
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published 28 days ago • 38
$π$-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows Paper • 2605.14678 • Published May 19 • 108
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning Paper • 2601.18631 • Published Jan 26 • 48
RxSafeBench: Identifying Medication Safety Issues of Large Language Models in Simulated Consultation Paper • 2511.04328 • Published Nov 6, 2025 • 1
X-Boundary: Establishing Exact Safety Boundary to Shield LLMs from Multi-Turn Jailbreaks without Compromising Usability Paper • 2502.09990 • Published Feb 14, 2025 • 2
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published about 1 month ago • 41
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published about 1 month ago • 20
Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models Paper • 2607.03751 • Published about 1 month ago • 20
Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process Paper • 2607.03748 • Published about 1 month ago • 41
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning Paper • 2601.18631 • Published Jan 26 • 48
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Paper • 2501.05444 • Published Jan 9, 2025 • 3
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark Paper • 2501.05444 • Published Jan 9, 2025 • 3
FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow Paper • 2505.17399 • Published May 23, 2025 • 14