Hard Negatives, Hard Lessons: Relabeling False Negatives with LLMs
Hard Negatives, Hard Lessons (Findings of EMNLP 2025) prunes the BGE retrieval training collection from 15 datasets to 7, reducing it from approximately 1.6 million training pairs to 680,000. Despite this reduction by a factor of about 2.35, E5-base gains approximately 1 point in average nDCG@10 across 14 BEIR datasets.
The remaining data still contains labeling errors. Some passages answer the query but are treated as hard negatives during training. The paper proposes RLHN (ReLabeling Hard Negatives), which uses two LLM stages to identify these passages and relabel them as positives. Compared with training on the pruned data with its original labels, relabeling improves E5-base and Qwen2.5-7B by approximately 0.7 and 1.4 points on BEIR, respectively.