Research/ ToolRet05 / 05
ToolRet logoIndependent undergraduate retrieval research

ToolRet.

When is reranking worth its cost?

An empirical study of conditional tool reranking across a 37,292-tool catalog, with untouched confirmation queries and explicit quality tolerances.

TOOLRET / CONFIRMATION
23.4%

fewer cross-encoder calls

CROSS-ENCODER CALLS · 1,500 QUERIES

Always rerank1,500
Conditional router1,149

351 fewer calls on untouched confirmation queries.

23.4% fewer calls1,500 QUERIES
THE IDEA

What ToolRet does.

This empirical study asks whether a learned router can skip cross-encoder reranking while preserving retrieval quality within a tolerance fixed before confirmation. It searches a 37,292-tool catalog and separates development, calibration and untouched confirmation queries by relevant-tool components.

01

Hybrid first-stage retrieval

BM25 and MiniLM dense search each retrieve 100 tools; equal-weight reciprocal-rank fusion combines their rankings.

02

A cheap conditional decision

A seven-feature ridge router uses first-stage disagreement, overlap, coherence, query length and score margin to decide whether to rerank the first 20 candidates.

03

Frozen confirmation design

300 development, 300 calibration and 1,500 confirmation queries have disjoint relevant-tool components; fitted parameters and criteria were frozen before confirmation.

04

Evidence you can inspect

The public report, source-stratified intervals, random controls, failure cases, real conditional execution, raw timing and 317-check independent audit are retained.

BUILT WITH

The technology.

FROM INPUT TO OUTCOME

How it works.

Always reranking adds compute and latency, and can even worsen rankings for some query sources. A cost-aware decision must be evaluated on fresh queries, against simple gates and matched-budget controls, rather than selected after seeing the holdout.

01 / 04

Retrieve

BM25 and all-MiniLM-L6-v2 search the same complete tool catalog.

ENGINEERING DECISIONS

Why it’s built this way.

01

Protect the holdout

Relevant-tool components separate supervised exposure. The primary policy and 0.01 nDCG@10 loss tolerance were retained even when a simpler gate had a stronger observed tradeoff.

02

Compare allocation, not just volume

Global and source-count-matched random controls test whether router allocation adds value beyond simply reducing the number of reranker calls.

03

Measure real conditional execution

Every confirmation decision was replayed with actual conditional inference. Warm sequential CPU timing includes retrieval, query encoding, exact search, fusion, routing and reranking.

RECORDED CHECKPOINTS

What the evidence shows.

On 1,500 untouched confirmation queries, the primary policy made 1,149 cross-encoder calls rather than 1,500: 351 skipped calls, or 23.4% fewer. nDCG@10 was 0.5273 versus 0.5241 for always reranking, and both paired interval lower bounds met the declared -0.01 tolerance. Quality superiority was not established.

23.4%Fewer cross-encoder calls
1,500Untouched confirmation queries
37,292Tools in the retrieval catalog

CROSS-ENCODER CALLS · 1,500 QUERIES

Always rerank1,500
Conditional router1,149

351 fewer calls on untouched confirmation queries.

SCOPE & LIMITS

This is independent undergraduate empirical research, not a peer-reviewed publication or a newly invented gating algorithm. The simple disagreement gate had a stronger observed quality/compute tradeoff. The benchmark cohort and warm CPU setting are not production traffic or downstream agent execution.

Read the full research PDF
OPEN THE WORK

Code, documents & proof.

Read the original study and inspect the code, frozen evaluation, and audit.

RESEARCH STUDYToolRetMichael Baffour Awuah · 2026

The original ToolRet study

The original confirmation report is available as a PDF, unchanged from the research repository.