NUS · CS4248 Natural Language Processing · Aug – Dec 2025
Margin-Triggered Reranking for Extractive QA
A group project on extractive question answering: given a passage and a question, find the span of words in the passage that answers it. We fine-tuned RoBERTa-base on SQuAD v1.1 and got 84.28 exact match and 90.93 F1. The five of us are close friends; the others had heavy course loads or an internship that semester, so one teammate and I did the coding and experiments, and the other three wrote the report.
01Where the mistakes were
The model scores every possible answer span, but normally you only ever use the top one. I had it return its top candidates instead and checked where the right answer landed. With the top five, the correct span was there for 95.1% of questions. If you could always pick it, exact match would go from 84.28 to 95.11, and F1 from 90.93 to 97.08. So most of the model's mistakes were about ranking: it found the answer but put another one first.
The gap between its top two scores turned out to be the useful signal. When the right answer was ranked first, the median gap was about 0.60. When the right answer was second, the gap dropped to about 0.11. A small gap meant the model itself wasn't sure.
02The reranker
So we only rerank when that gap is below 0.05. In those cases a small sentence-embedding model (all-MiniLM-L6-v2) scores both candidates against the question, and that score is blended evenly with the original one. Confident answers are left alone, so they cost nothing extra.
It changed 55 of the 10,570 answers: 14 went from wrong to right and 1 from right to wrong, which took exact match from 84.28 to 84.40 and F1 from 90.93 to 91.04. That's a small gain from a single run, so I wouldn't call it significant. We also tried the obvious alternatives: reranking every question, reranking the top three or five instead of the top two, and a heavier cross-encoder. The first two made results worse, and the cross-encoder cost more without doing consistently better. The repo keeps all of those runs.