Hybrid first-stage retrieval
BM25 and MiniLM dense search each retrieve 100 tools; equal-weight reciprocal-rank fusion combines their rankings.
GitHub
Independent undergraduate retrieval researchAn empirical study of conditional tool reranking across a 37,292-tool catalog, with untouched confirmation queries and explicit quality tolerances.
fewer cross-encoder calls
CROSS-ENCODER CALLS · 1,500 QUERIES
351 fewer calls on untouched confirmation queries.
This empirical study asks whether a learned router can skip cross-encoder reranking while preserving retrieval quality within a tolerance fixed before confirmation. It searches a 37,292-tool catalog and separates development, calibration and untouched confirmation queries by relevant-tool components.
BM25 and MiniLM dense search each retrieve 100 tools; equal-weight reciprocal-rank fusion combines their rankings.
A seven-feature ridge router uses first-stage disagreement, overlap, coherence, query length and score margin to decide whether to rerank the first 20 candidates.
300 development, 300 calibration and 1,500 confirmation queries have disjoint relevant-tool components; fitted parameters and criteria were frozen before confirmation.
The public report, source-stratified intervals, random controls, failure cases, real conditional execution, raw timing and 317-check independent audit are retained.
Always reranking adds compute and latency, and can even worsen rankings for some query sources. A cost-aware decision must be evaluated on fresh queries, against simple gates and matched-budget controls, rather than selected after seeing the holdout.
BM25 and all-MiniLM-L6-v2 search the same complete tool catalog.
Relevant-tool components separate supervised exposure. The primary policy and 0.01 nDCG@10 loss tolerance were retained even when a simpler gate had a stronger observed tradeoff.
Global and source-count-matched random controls test whether router allocation adds value beyond simply reducing the number of reranker calls.
Every confirmation decision was replayed with actual conditional inference. Warm sequential CPU timing includes retrieval, query encoding, exact search, fusion, routing and reranking.
On 1,500 untouched confirmation queries, the primary policy made 1,149 cross-encoder calls rather than 1,500: 351 skipped calls, or 23.4% fewer. nDCG@10 was 0.5273 versus 0.5241 for always reranking, and both paired interval lower bounds met the declared -0.01 tolerance. Quality superiority was not established.
CROSS-ENCODER CALLS · 1,500 QUERIES
351 fewer calls on untouched confirmation queries.
This is independent undergraduate empirical research, not a peer-reviewed publication or a newly invented gating algorithm. The simple disagreement gate had a stronger observed quality/compute tradeoff. The benchmark cohort and warm CPU setting are not production traffic or downstream agent execution.
Read the original study and inspect the code, frozen evaluation, and audit.
The original confirmation report is available as a PDF, unchanged from the research repository.