Published on
11 min readmentions

Stop Averaging Search Evals

Report p25/p50 and debug your queries

Looking for TL;DR? Check key takeaways

Any serious team tracks latency as p50, p90, p99. Reporting just "our average latency is 40ms" is naive, because it hides the p99 150ms tail that's hurting real users. Yet we report search relevance as a single average NDCG all the time, and somehow that's normal. Relevance has a tail too, but the average hides it.

Average can't represent the distribution

I ran a plain word-BM25 baseline over 12 NanoBEIR datasets (599 queries). Instead of averaging each metric, I calculated percentiles over the per-query scores. The mean NDCG@10 is 52.5, which feels respectable, as if every query comes back roughly half-right. In reality, they range from perfect to nothing: one in five queries retrieves nothing relevant in the top 10, and the average buries every one.

(Note: Every score here is ×100, the way IR papers report NDCG. A mean of 52.5 reads better than 0.525, and a +2.2 gain reads as a real change where +0.022 reads as rounding.)

Let's look at the distribution.

Observations:

  • 119/599 queries score exactly 0 and 129/599 score exactly 100 on NDCG@10. That's 41% of all queries sitting at the two extreme ends.
  • Only 43 of 599 queries land within 5 of the mean (52.5), so it describes a query that barely exists.
  • Switch to recall@10 and it gets worse: 63% of queries at the two ends, 13 near the mean.
  • The mean's error flips sign between metrics. It overstates recall@10's median (56.3 vs 50.0) and understates NDCG@10's (52.5 vs 55.4), so you can't correct for it in your head.

Understanding impact of your changes

Take a real improvement: plain word BM25, then apply a Porter stemmer and a stopword list (word_en). It moves mean NDCG@10 by +1.9, a bump you'd squint at.

(Setup: BM25 with k1=1.5, b=0.75 and Lucene IDF. word lowercases and splits on word characters. word_en also drops NLTK's English stopword list and applies NLTK's Porter stemmer.)

The first thing I've started reading is p25 and p50 along with mean. It comes out of the per-query scores I already have with a 1 line change. Toggle the statistic:

Observations:

  • The mean never moves most. On every metric, p25 or p50 moves further (NDCG@10 p25 +4.5, recall@10 p50 +10.0, recall@100 p25 +12.6) while the mean moves +1.9, +2.2 and +2.8.
  • recall@100's median is pinned at 100 before and after, so it reads +0.0 while p25 moves +12.6.
  • recall@10 is the other way round, the median jumps +10.0 while p25 moves +3.6. Neither percentile wins every time, which is why I report both.

p50/p25/mean are great for communicating. But if you wanna drill down and properly compare two methods, you should look into exact queries that are improved/degraded.

That's why I like to bucket the exact deltas per query for my metrics and use that to find queries to focus on:

Observations:

  • 294/599 (49%) queries don't move. About half the set sits in the no-change bar, so nearly half the deltas behind +1.9 are zeros.
  • Win/tie/loss is 172/294/133. Of the 305 queries that actually moved, 56% got better, which is why the mean lands so near zero.
  • Weight each query by how far it moved and it's 3359 points gained against 2217 lost. 60% of the movement was upward, better than the 56% the raw count suggests, because winners move further than losers do (+19.5 each against −16.7).
  • The stemmer and stopword list rescued 25 queries from 0 (Which colors can one use to fill out a check in the US? goes 0 to 100) and sent 9 to 0 (who sang here i go again goes 100 to 0). That moves the zero-rate, the share of queries with nothing relevant in the top 10, from 19.9% to 17.2%.
  • Both ends hit the full 100. Which colors can one use to fill out a check in the US? gains 100, who sang here i go again loses 100, and 5 queries lost more than 50 points against 11 that gained more than 50.

I like to ask an LLM to build a UI for me to explore the dataset. It just needs basic sorting by metrics, filtering, and search (yeah. use search to improve search :D). Below is a cut-down version with three queries from this run, click one to see what each analyzer sent to the index and what came back:

Loading…

Some interesting examples:

  • who sang here i go again went from a perfect 100 to 0 on NDCG@10. "Here I Go Again" is a song title made of stopwords, so the stopword list deletes most of it and we only ever query the index for sang go.
  • who appoints the members of the given branch in the united states also went 100 to 0. The query becomes appoint member given branch unit state. The relevant doc ("...appointed by the President of the United States...") was rank 1 before and falls out of the top 10 after.
  • Even recall@100, the most forgivingit only asks whether a relevant doc landed anywhere in the top 100, so a doc can slip 90 places and still count metric here, falls to 0 on 4 queries. One is the ClimateFEVER claim No state generates as much solar power as California, and the stemmer cuts "generates" down to generPorter gives that same stem to general, generally, generic, generous and generation, so lots of unrelated documents pop up.
  • 9/599 queries fell from a real NDCG@10 score to exactly 0 like this.
  • What's most surprising is that stemmer and stopwords are considered standard but it's clear even they hurt some queries. There's no free lunch. You must look into your data/distribution.

The same thing happens one level up. Split the +1.9 by dataset:

Observations:

  • Should prostitution be legal? drops 73 → 25 and Should the voting age be lowered? drops 68 → 30. Both come from Touche2020, a set of debate posts, which loses −6.6 overall (14 queries better, 33 worse). 10 of the other 11 datasets gain and SCIDOCS stays flat.
  • Both halves of the change hurt Touche2020. The stemmer alone costs −1.3, the stopword list alone −3.3, and the two together −6.6.
  • Should recreational marijuana be legal? loses everything to the stemmer alone (61 → 26). Stemming makes legal also match legalized, legalization and legalizing, and with word_en a post like "Recreational Marijuana should not be legalized. It is harmful and dangerous..." lands in the top 10. It's on topic, but the judges didn't mark it relevant.
  • MSMARCO gains +0.5 on its mean, yet 2 more of its queries now score 0 (14 → 16).
Three more ways to slice the same run

Splitting the mean into what was gained and what was given back:

metricgainedlostnet (the mean)
NDCG@10+5.6−3.7+1.9
recall@10+4.2−2.0+2.2
recall@100+4.0−1.2+2.8

Every mean here is a small difference between two larger numbers. NDCG@10 gives back two thirds of what it wins, recall@100 under a third. Same-looking gains, different trades.

A saturated percentile (one stuck at 0 or 100) can't show a gain. NDCG@10's p90 is already at 100, so a better method moves it +0.0. The 22% of queries already scoring 100 have no room left to improve, and they dilute the mean.

Subtracting one run's percentiles from the other's answers a different question: did the whole distribution shift?

Observations:

  • The mean smears the gain across every query (+2.2). The median reads it where the queries can actually move (+10.0, 4.6×).
  • p10, p75 and p90 all sit at +0.0 because they're saturated, p10 at 0 and the other two at 100.

The pattern holds across metrics:

metricmean Δwhere it shows up
recall@10+2.2median jumps 50 → 60, a +10.0 move (4.6×), while p75 and p90 stay pinned at 100
NDCG@10+1.9p25 / p50 / p75 all move ~2× the mean (+3.7 to +4.5)
recall@100+2.8median already 100, so the gain lands in p25: 50 → 63 (+12.6, 4.5×)

The gain lands wherever the queries still have room to move, and that percentile shifts 2-5× as far as the mean.

The relevance tail

Latency and relevance are the same problem pointed at opposite tails. For latency, higher is worse, so you watch the high end: p90, p99. For relevance, higher is better, so the tail that hurts is the low end. That's the relevance tail, the weaker half of your queries. p25 is the middle of that half, so it's the number I watch.

Report p25, not p10. p10 pins at 0 once more than 10% of queries fail, and 20% do here, so it can't show the tail moving. p25 is 21.9 on NDCG@10 and still moves when the method changes.

I timed this BM25 baseline at 3.2ms p50, 7.2ms p90. Nobody would report only the 4.4ms mean, because the p90 is the request a real user is waiting on. A p25 relevance score is the score a quarter of your queries get or worse. We just forget to report it.

So what should we actually track?

The queries are the real thing. Everything else in this post is a way of finding them faster.

A search engineer should be able to tell you what kinds of queries their system fails on. In this run it was song titles made of stopwords, and queries whose one distinctive word gets stemmed into a common one. That's the real result of an eval. The distribution, the percentiles and the mean are summaries of it.

  • The distribution shows you where to look. A spike at 0 means go read the zeros.
  • Win/tie/loss shows you what a change moved, so it's what I reach for when comparing two of my own runs. It stops at what moved, so you still end up reading the queries it flags.
  • p50, p25 and the mean are for communicating. Three numbers fit in a message and compare across runs.

Your golden set is probably small. Mine is 599 queries and most teams have fewer. That's small enough to treat like a test suite. A query that used to work and now returns nothing is a regression, and you want to see it by name.

Beyond search: any eval with graded per-item scores, an LLM judge's rubric, a similarity score, a per-query metric, hides its tail in the mean the same way.*Binary pass/fail evals are the exception: there the mean already tells you everything.

Mean alone can't capture a distribution. We realised this for latency decades ago, and search relevance deserves the same. Then go read the queries.

Key Takeaways

  • A good average can hide queries that are hurting real users. Plain BM25 averages 52.5 NDCG@10, yet 119 of 599 queries find nothing relevant in the top 10.
  • It hides what changed. Adding a stemmer and stopwords moves mean NDCG@10 by a forgettable +1.9, while 172 queries get better, 133 get worse and one dataset (Touche2020) drops 6.6.
  • Report p50 and p25 next to the mean. Higher is better for relevance, so the tail that hurts is the low end. p25 is the relevance equivalent of p90/p99 latency.
  • Win/tie/loss says what the mean can't. That +1.9 was 172 queries up, 133 down and 294 unmoved. Weighted by how far each one moved, 60% of the movement was upward.
  • Track your golden queries like test cases. Mine is 599 queries and most sets are smaller. A query that used to work and now returns nothing is a regression you should see by name, not as a mean that moved by 1.9.

I believe that the industry and academia should start reporting p50/p25 along with the mean.

✓ link copied
← Back to the blog