How We Test, Measure and Cite
The benchmark setup, the hardware, the versions and the rules every page on this site is held to.
RAG Explained Better holds every substantive page to one quality bar: current primary sources checked before it publishes, every number sourced or omitted, volatile claims dated in prose, and wrong claims corrected on the page. This page is that bar — how we test, measure and cite — not a tutorial on evaluating your own RAG system (that lives under how to evaluate a RAG system).
How does this site research a topic before it publishes?
Every substantive page is researched against primary sources before it publishes — vendor documentation, maintained repositories, papers, and the maintainers of the benchmarks it cites — read for the current state of the topic, not written from a stale summary. The bar is coverage plus gain: a page has to address what a reader working on that topic would expect it to answer, and then add at least one fact, measurement, or worked example the existing write-ups do not have. A page that only restates what is already published has no reason to exist and does not ship.
If you want to see the method applied end to end, build a RAG pipeline is a runnable walk-through, and chunking evaluation shows the measurement side worked out in full.
How are numbers and benchmarks cited on this site?
A pinned library version, an embedding price, a memory formula, a latency figure or a benchmark score is only published when it traces to a named origin in the same sentence — vendor documentation, a maintained repository, a paper, or the benchmark’s maintainer — and is never invented to fill a gap. If no public number exists, the page writes “not published” or drops the claim. Primary sources are preferred over press write-ups of the same result. Decorative bibliographies that do not attach to the claim in prose are not used; extractors and readers both need the source next to the number. Where a figure could drift, the page says to verify it before you rely on it.
The shorter sibling statement of this rule sits on how sources are chosen. Public RAG benchmarks and what they actually test are catalogued at public RAG benchmarks — cited when a page needs them, never fabricated to decorate a comparison.
When does a page get an as-of date?
Every volatile claim — model versions, library APIs, leaderboard ranks, pricing, deprecations — carries its as-of date in the prose itself, not only in a meta field. Benchmarks move and a result that was current at publish may not be six months later; dating the claim is how a reader weights it. This site does not promise a fixed quarterly re-test of every URL. Freshness is enforced by dating what can move, and by correcting the page when a dated claim is shown to be wrong — the route is below.
How are measurement pages held to a harder bar?
Pages that report a score this site computed — a chunking experiment, a worked precision or recall example, a tool survey run on one pipeline — must state the pinned library and model versions, the inputs the number depends on, and, when hardware matters to the claim, the hardware class used, so a reader can reproduce or reject the result. Pages that only cite someone else’s published benchmark still name the maintainer and the as-of; they do not invent a site-run setup that was never run. A measurement page without those pins is incomplete.
Proof surfaces on the site today: how to evaluate a chunking strategy for a repeatable experiment, and RAG evaluation tools compared for a same-pipeline survey. The cluster hub is how to evaluate a RAG system.
Found an error? How to get it corrected.
Claims here are meant to be checkable, which means some will be wrong and should be fixed. If a number has drifted, a source has moved, or a benchmark no longer holds, reach the author through any profile in who writes this — GitHub is the fastest — or via Mrunmay Phanse. Corrections to a factual claim are made on the page itself, with a visible note when the change is material, not argued only in a comment thread and not quietly rewritten as if the wrong number never existed. A site that asks you to trust its numbers has to be willing to change them.
How does this site research topics?
Every substantive page is researched against primary sources before it publishes — vendor documentation, maintained repositories, papers, and the maintainers of the benchmarks it cites — read for the current state of the topic. The bar is coverage plus gain: a page must answer what a reader working on that topic expects, and add at least one fact, measurement or worked example the existing write-ups do not have.
Do you invent benchmark numbers?
No. A pinned version, price, latency, memory formula or benchmark score is only published when it traces to a named origin in the same sentence — vendor docs, a repository, a paper or a leaderboard maintainer. If no public number exists, the page says “not published” or omits the claim. Inventing a figure to fill a gap is treated as a hard failure.
Why do pages say “as of” a date?
Model versions, library APIs, leaderboard ranks, pricing and deprecations move. Every volatile claim carries its as-of date in the prose so a reader can weight it. The site does not promise a fixed quarterly re-test of every URL; freshness is enforced by dating the claim and correcting the page when a dated claim is shown to be wrong.
How do I report an error?
Reach the author through any profile linked from /about — GitHub is fastest — or via /author/mrunmay-phanse/. Corrections to a factual claim are made on the page itself, with a visible note when the change is material. Silent stealth-edits of wrong numbers are not the policy.