VGI-Bench: Probing Visual Intelligence in Video Generation Models Paper • 2608.19583 • Published 14 days ago • 179
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 23 days ago • 340
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published 26 days ago • 170
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published 28 days ago • 291
PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say Paper • 2606.00152 • Published Aug 6 • 4
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published Aug 3 • 142
electricsheepafrica/africa-rwanda-season-c-fertilizer-and-pesticide-use-e4c59f8a Viewer • Updated 28 days ago • 2.31k • 43 • 1
Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs Paper • 2607.03936 • Published Jul 4 • 5
MBench: A Comprehensive Benchmark on Memory Capability for Video World Models Paper • 2606.00793 • Published Jun 8 • 12