Bekko Embedding: Parameter-Efficient Multilingual Retrieval with Ultra-Compact Encoders Paper • 2607.25180 • Published 15 days ago • 2
view article Article GLInt: Geometry-Matched Hard Negatives for Late-Interaction Retrieval chungimungi • 4 days ago • 11
Compact Language Models via Pruning and Knowledge Distillation Paper • 2407.14679 • Published Jul 19, 2024 • 43
DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search Paper • 2607.27178 • Published 14 days ago • 6
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 7 days ago • 23
view article Article After the party comes the free lunch: regularizing ColBERT models to enhance pooling capabilities and reduce index footprint lightonai • Jul 6 • 15
view article Article mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models lightonai • 13 days ago • 35
jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation Paper • 2607.18152 • Published 23 days ago • 4
ProRank: Prompt Warmup via Reinforcement Learning for Small Language Models Reranking Paper • 2506.03487 • Published Jun 4, 2025 • 7
view article Article NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval nvidia • 27 days ago • 58
view article Article Beyond LoRA: Can you beat the most popular fine-tuning technique? +2 BenjaminB, sayakpaul, hubnemo, kashif • Jun 18 • 91
Training Sparse Mixture Of Experts Text Embedding Models Paper • 2502.07972 • Published Feb 11, 2025 • 12
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Paper • 2501.14818 • Published Jan 20, 2025 • 10
view article Article Party is over: regularizing ColBERT models to fix efficient ANN methods lightonai • Jun 16 • 23
Rank-DistiLLM: Closing the Effectiveness Gap Between Cross-Encoders and LLMs for Passage Re-Ranking Paper • 2405.07920 • Published May 13, 2024 • 4
F2LLM-v2: Inclusive, Performant, and Efficient Embeddings for a Multilingual World Paper • 2603.19223 • Published Mar 19 • 39
Is Position Bias in Dense Retrievers Built In-or Learned from Data? Paper • 2605.26578 • Published May 26 • 21