For research and information-retrieval evaluation only. Not a medical device. Not for diagnosis, treatment, or patient care.

Results

Long-tail / rarity analysis

RQ7: does RAG quality degrade as disease evidence gets sparser? Recall@10 by rarity stratum and retriever, on the v1 task set.

Read this as descriptive, not statistically tested. Stratum sizes are small and uneven — high=12, medium=4, sparse=19 gold-bearing tasks. Per claims_registry.md C-020, this is still preliminary: the medium stratum in particular (n=4) is too small for a bootstrap confidence interval. The pattern below is directionally consistent with the project’s central hypothesis, but is explicitly not yet a confirmed finding — no per-stratum significance test has been run, unlike the pooled retriever comparisons on the benchmark page.

Recall@10 by rarity stratum and retriever

Rarity stratumn tasksBM25DenseHybrid
high120.6250.8540.701
medium40.6670.7500.875
sparse190.6490.7190.724

Dense retrieval’s Recall@10 declines with rarity — high=0.854 → medium=0.750 → sparse=0.719 — the same direction as the v0 preliminary finding, now on roughly double the sample (claims_registry.md C-020). This held up under more data rather than disappearing, which is modestly reassuring for RQ7, but is still not a bootstrap-confirmed effect.

Figure

Recall by rarity stratum across retrievers

Why rarity strata, not raw prevalence

The rarity stratum is a reproducible corpus-coverage proxy — document count in the frozen corpus, split into terciles — not an external prevalence claim. A disease with few indexed papers may still be well understood clinically through channels this corpus doesn’t capture. Findings here generalize only as far as this project’s disease selection does, which was chosen for a spread across the proxy and ontology-ID availability, not randomly sampled from all known rare diseases.