Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 26 days ago • 58
Do Text Edits Generalize to Visual Generation? Benchmarking Cross-Modal Knowledge Editing in UMMs Paper • 2606.00477 • Published May 30 • 1
BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models Paper • 2605.05758 • Published May 7 • 6
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation Paper • 2605.25874 • Published May 25 • 108
BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models Paper • 2605.05758 • Published May 7 • 6
Can Prompts Rewind Time for LLMs? Evaluating the Effectiveness of Prompted Knowledge Cutoffs Paper • 2510.02340 • Published Sep 26, 2025