Keep Drafting Parallel
AI & ML interests
Efficient AI
Recent Activity
Papers
DFlash: Block Diffusion for Flash Speculative Decoding
ParoQuant: Pairwise Rotation Quantization for Efficient Reasoning LLM Inference
models 59
z-lab/Muse-Glimmer-30B-DFlash2
Text Generation • 3B • Updated • 1.17k • 12
z-lab/Qwen3.8-27B-DFlash2
Text Generation • 2B • Updated • 21.1k • 164
z-lab/Muse-Glimmer-30B-DFlash2-GGUF
Text Generation • 3B • Updated • 1.72k • 6
z-lab/Qwen3.8-27B-DFlash2-GGUF
Text Generation • 2B • Updated • 15.9k • 68
z-lab/Qwen3.8-27B-PARO
Image-Text-to-Text • 6B • Updated • 216 • 5
z-lab/Alpamayo-R1-10B
Robotics • 11B • Updated • 155 • 3
z-lab/Alpamayo-1.5-10B
Robotics • 11B • Updated • 5.5k • 5
z-lab/Alpamayo-R1-10B-PARO
Robotics • 4B • Updated • 18 • 3
z-lab/Alpamayo-R1-10B-DFlash
Robotics • 0.5B • Updated • 155 • 4
z-lab/Alpamayo-1.5-10B-PARO
Robotics • 4B • Updated • 81 • 4
datasets 8
z-lab/glm52-cc
Viewer • Updated • 2.26k • 39 • 3
z-lab/long-code-test
Viewer • Updated • 1k • 68 • 3
z-lab/kimi-k26-regen
Viewer • Updated • 1M • 8 • 3
z-lab/humaneval-long
Viewer • Updated • 1k • 43 • 1
z-lab/gsm8k-filtered
Viewer • Updated • 1.31k • 65 • 1
z-lab/mt-bench-filtered
Viewer • Updated • 79 • 29 • 1
z-lab/mbpp-sanitized-filtered
Viewer • Updated • 256 • 40
z-lab/humaneval-filtered
Viewer • Updated • 137 • 63 • 1