/// METADATA
DATE PUBLISHED
2026.08.23
EST. READING TIME
10 min read
VIEWS
41
TAGS
elasticsearchperformance

The Forgotten Routing Key: Looking Past the One-Line Fix

You got home late one evening, changed, and picked up your phone to look for a piece of clothing you had been wanting. But the seller’s page didn’t quite explain the material, the size table, whether it was cotton or polyester, whether it would fit you.

In a physical store you would just ask. Online, you can still send the seller your question — but with some luck, someone has already asked it, and the answer is waiting in the question box under the product.

That little box is doing the job of the salesperson we lost when shopping moved online, and it does that job for every shopper at once. It is not one salesperson and one customer; it is one product and everyone who is about to buy it, arriving through the same narrow window. Most days the crowd is thin enough that we forget it is a crowd.

Then November arrives. Black Friday. The little box stays exactly the same size, but the crowd trying to look through it does not. Every year we tighten what sits behind it. Every year the crowd outgrows the November before it.

In this article, we’ll walk through what sat behind that question box, and why all those years of tightening never gave us a finish line. Join me as we examine the query the crowd actually hits.

Some of what follows took two Novembers to play out, and the part that mattered most we only understood this year. So we kept tightening. Before every November we took another pass at the questions page. Last year we went deeper than usual. One product’s questions already lived together on one shelf; what was missing was the pointer, so we taught the store to walk straight there. Then, instead of deepening that improvement, we reached for something more complex: a circuit that drops the topic counts and lets the list live if the cluster starts to drown. It helped. It was not enough.

This year’s first load test was not a drill either. Night jobs were running, and cache covered about twenty percent of the traffic. We pushed anyway. The system fell over. And that confused us. Last year’s routing fix was supposed to point everything home. Meanwhile, a big rewrite already sat on the table. It would move the topic counts out of Elasticsearch and into a table of our own. It felt close. It felt like the real work.

Before committing to that, we went back to the page itself. We had looked last year too, though not deep enough, we suppose. This time we opened the page, listed every query it fires, and read each Elasticsearch query one by one. And there it was: the chips are a second query, a different one, and nothing pointed it home.

The questions page: topic chips with counts on the left, the question list on the right

The list query is simple in spirit. Give us this product’s questions. Last year we taught that path the product id as a routing key. The store could walk to the right shelf.

The numbers on the chips are a different query, an aggregation. Count the questions per topic, still for this one product. We filtered by product. We did not route. So the store asked every shelf, and every shelf counted, and then we added the counts up. That was the heavy one. That was the one we had not watched.

A routing key tells the store which shelf holds this product. Without it, every shelf pays for one product’s questions.

That’s The Forgotten Routing Key. The cheap instruction was already in the house. We had not put it on the query that hurt.

Scattered

We run Elasticsearch 8.13.4. There is no coordinating-node pool. Three nodes are master-only. Twelve nodes are data. Search does not land on a master. The search client points at the data endpoints. The data node that accepts the HTTP request plays the coordinating role for that one search. Then it fans the work to the shards that might hold a hit.

The questions index has twelve primary shards. One replica. About 136 million questions, spread evenly, about eleven million on each primary. Without a routing key, “might” means all twelve. A filter on product id still wakes every shard. Each shard scans its own slice. The receiving data node merges what comes back.

One product. Twelve shards. Every data node that holds a slice still pays.

Last year’s stick

Last year we changed where a question sits. The indexer no longer lets Elasticsearch hash the document id. It sends _routing set to the product id. Same product, same shard. If the product id on a row changes, the old document is deleted with the old routing and written with the new one.

The list query does the same on read. When the request has a product id, the search sends that value as routing. The data node that accepted the HTTP request does not fan to twelve shards. It talks to the one shard that holds that product. The product-id filter is still in the query. The routing is what skips the other eleven.

PUT /questions/_doc/{id}?routing={productId}

GET /questions/_search?routing={productId}

The list could walk home. We thought the page could too.

The count still broadcasts

The topic chips are not a lookup. They are a terms aggregation. size: 0. No hits come back. A search can ask an inverted index where this product lives. A terms agg cannot return topic counts from that index alone. Each shard that receives the request walks documents and increments a bucket per topic.

for shard in shards_that_got_the_request:
  for doc in shard:
    if doc.productId == productId:
      counts[doc.topic] += 1

Without _routing, the outer loop is twelve. The inner loop is about eleven million documents on each shard. Almost none match. They still pay the for. With _routing, the outer loop is one. The inner loop is only that product’s questions on one shard.

Elasticsearch says the same thing in two steps. Collect on each shard. Then merge. To get more accurate results, the terms aggregation in 8.13 fetches more than the top size terms from each shard — it fetches the top shard_size terms — so even a merge of almost-empty results takes more bytes over the wire and more waiting in memory on the coordinating node. The for is the collect. Our query still had the product filter. It did not have _routing. So we paid twelve collects, then a merge of twelve almost-empty results.

Last year we put _routing on the list. The aggregation still only carried a term on product id inside the query body. That term is a filter. Elasticsearch does not use it as a routing key. Omit _routing on the aggregation, and the receiving data node fans that for to all twelve shards. Eleven of them hold none of the product. They still scan. Send the wrong routing value, and you land on a shard that does not hold the product. The counts come back empty, or short. The documents did not move. The query did.

So the page still paid twelve fors for every load of topic counts. Last year’s stick helped the list. It did not help the counts. This year’s work was the aggregation.

GET /questions/_search
{
  "size": 0,
  "query": { "term": { "productId": "{productId}" } },
  "aggs": { "topics": { "terms": { "field": "topic" } } }
}

The documents were home. The count query was not.

Routed, work follows the product. Unrouted, work follows the whole index.

We did not remove the for. A terms agg still walks documents. That’s how Elasticsearch works. We stopped the walk from running on twelve shards. One collect, on the shard that actually holds the product. That is the whole gain.

The count sticks

This year we put _routing on the aggregation. Same product id the indexer already used. Same shard the list already walked. The for still runs. It runs once.

GET /questions/_search?routing={productId}
{
  "size": 0,
  "query": { "term": { "productId": "{productId}" } },
  "aggs": { "topics": { "terms": { "field": "topic" } } }
}

One collect. Eleven shards idle.

Last year’s ceiling was Elasticsearch CPU. In every earlier load test, its CPU utilization climbed toward one hundred percent, and that is what broke us. It is also why we kept night jobs off in those runs: night jobs drive indexing inside Elasticsearch, and indexing is pure CPU work. So with night jobs off, indexing off, and cache covering about forty percent of the traffic, we topped out around 600 thousand requests per minute. This year we ran the harder configuration on purpose — night jobs on, indexing on, cache off — which makes any comparison to that old ceiling conservative.

Cache off, the same page held 2.1 million requests per minute. With the night jobs in the mix, about 3 million. For scale: a quiet day peaks around 200 thousand — this test ran at more than ten times that.

Elasticsearch did not fall. The cluster still had room left. What throttled us in the end was Java CPU on the applications. That is a whole other problem of ours

Looking past the one-line fix

The Forgotten Routing Key, restated in one breath: a term in the query body filters, and only _routing points. Two jobs sat on one page. We pointed the list and left the count waking twelve shards. In other words, one heavy query ran on every shard instead of one, and the CPU cost grew with the catalog.

And here is the part that still stings, because it is the AI-era version of this mistake. The rewrite did not feel like the real work despite the tools. It felt like the real work because of them. Code generation got so fast that the old cost of a big change — dual-write, backfill, cutover, weeks of careful planning — collapsed into an afternoon of scaffolding. So instead of sitting down and debugging the simple thing right in front of us, we shrugged: whatever, refactors are cheap now, we will just do the big one. We aimed at the big path and never moved. The cheapest collect on the board sat there, one parameter away, while we daydreamed about architecture.

In conclusion, the line was already in the house. The rewrite was not required. The urge to write more code is ours to manage. We still own the system, and the walk above is why. When a bigger system feels close, the query we already fear is still the first place to look.