GraphRAG Retrieval-Quality Evaluation — 2026-08-09
Status: [MEASURED — live endpoint]
Gold set: gold_v1.jsonl · Backend: endpoint · Generated: 2026-08-09T22:58:43.177851+00:00
Summary
| Metric | Value |
|---|---|
| n_questions | 9 |
| n_scored | 7 |
| n_negative | 2 |
| precision@5 | 0.0286 |
| recall@5 | 0.4286 |
| precision@10 | 0.0143 |
| recall@10 | 0.4286 |
| mrr | 0.0476 |
| graph_coverage | 0.7143 |
| facet_recall | 0.2143 |
| negative_pass_rate | 1.0 |
| latency_p50_ms | 7722.46 |
| latency_p95_ms | 8369.81 |
Per-question results
[
{
"id": "gq-001",
"archetype": "journey_traversal",
"latency_ms": 7840.58,
"n_retrieved": 10,
"precision@5": 0.2,
"recall@5": 1.0,
"precision@10": 0.1,
"recall@10": 1.0,
"mrr": 0.3333,
"graph_coverage": 0.0,
"facet_recall": 1.0
},
{
"id": "gq-002",
"archetype": "journey_traversal",
"latency_ms": 7904.69,
"n_retrieved": 6,
"precision@5": 0.0,
"recall@5": 0.0,
"precision@10": 0.0,
"recall@10": 0.0,
"mrr": 0.0,
"graph_coverage": 0.0,
"facet_recall": 0.5
},
{
"id": "gq-101",
"archetype": "causal_path",
"latency_ms": 7722.46,
"n_retrieved": 5,
"precision@5": 0.0,
"recall@5": 0.0,
"precision@10": 0.0,
"recall@10": 0.0,
"mrr": 0.0,
"graph_coverage": 1.0,
"facet_recall": 0.0
},
{
"id": "gq-102",
"archetype": "causal_path",
"latency_ms": 8134.97,
"n_retrieved": 5,
"precision@5": 0.0,
"recall@5": 0.0,
"precision@10": 0.0,
"recall@10": 0.0,
"mrr": 0.0,
"graph_coverage": 1.0,
"facet_recall": 0.0
},
{
"id": "gq-201",
"archetype": "intent_explanation",
"latency_ms": 8369.81,
"n_retrieved": 5,
"precision@5": 0.0,
"recall@5": 0.0,
"precision@10": 0.0,
"recall@10": 0.0,
"mrr": 0.0,
"graph_coverage": 1.0,
"facet_recall": 0.0
},
{
"id": "gq-301",
"archetype": "aggregation",
"latency_ms": 7545.12,
"n_retrieved": 5,
"precision@5": 0.0,
"recall@5": 1.0,
"precision@10": 0.0,
"recall@10": 1.0,
"mrr": 0.0,
"graph_coverage": 1.0,
"facet_recall": 0.0
},
{
"id": "gq-302",
"archetype": "aggregation",
"latency_ms": 7371.26,
"n_retrieved": 5,
"precision@5": 0.0,
"recall@5": 1.0,
"precision@10": 0.0,
"recall@10": 1.0,
"mrr": 0.0,
"graph_coverage": 1.0,
"facet_recall": 0.0
},
{
"id": "gq-901",
"archetype": "negative",
"latency_ms": 7136.92,
"n_retrieved": 5,
"negative_pass": true
},
{
"id": "gq-902",
"archetype": "negative",
"latency_ms": 7372.17,
"n_retrieved": 5,
"negative_pass": true
}
]
How to read these means (TRUTH.md 4.4 — a flattering number ships with its deflating context)
The summary means are NOT all comparable, and two of them are inflated by
construction. metrics.py returns a free 1.0 when an expectation set is
empty: recall@k when expected_node_ids is empty, and graph_coverage
when expected_edges is empty. Aggregation-archetype questions legitimately
expect a count rather than nodes, and questions about a graph that holds no
edge of the relevant type legitimately expect no edges — so those rows score
1.0 without the retriever having retrieved anything.
Read precision@k, mrr and facet_recall as the load-bearing numbers, and
always read the per-question table below before quoting any mean. Do not cite
a summary recall or graph_coverage figure without stating how many rows
carried an empty expectation set.
Claim protocol
A retrieval-quality improvement claim must cite this artifact AND a prior baseline artifact produced with the SAME gold set version. Gold set and retriever must not change within the same claim window.