GraphRAG Retrieval-Quality Evaluation — 2026-08-09

Status: [MEASURED — live endpoint] Gold set: gold_v1.jsonl · Backend: endpoint · Generated: 2026-08-09T22:58:43.177851+00:00

Summary

Metric Value
n_questions 9
n_scored 7
n_negative 2
precision@5 0.0286
recall@5 0.4286
precision@10 0.0143
recall@10 0.4286
mrr 0.0476
graph_coverage 0.7143
facet_recall 0.2143
negative_pass_rate 1.0
latency_p50_ms 7722.46
latency_p95_ms 8369.81

Per-question results

[
  {
    "id": "gq-001",
    "archetype": "journey_traversal",
    "latency_ms": 7840.58,
    "n_retrieved": 10,
    "precision@5": 0.2,
    "recall@5": 1.0,
    "precision@10": 0.1,
    "recall@10": 1.0,
    "mrr": 0.3333,
    "graph_coverage": 0.0,
    "facet_recall": 1.0
  },
  {
    "id": "gq-002",
    "archetype": "journey_traversal",
    "latency_ms": 7904.69,
    "n_retrieved": 6,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 0.0,
    "facet_recall": 0.5
  },
  {
    "id": "gq-101",
    "archetype": "causal_path",
    "latency_ms": 7722.46,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0
  },
  {
    "id": "gq-102",
    "archetype": "causal_path",
    "latency_ms": 8134.97,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0
  },
  {
    "id": "gq-201",
    "archetype": "intent_explanation",
    "latency_ms": 8369.81,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0
  },
  {
    "id": "gq-301",
    "archetype": "aggregation",
    "latency_ms": 7545.12,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 1.0,
    "precision@10": 0.0,
    "recall@10": 1.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0
  },
  {
    "id": "gq-302",
    "archetype": "aggregation",
    "latency_ms": 7371.26,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 1.0,
    "precision@10": 0.0,
    "recall@10": 1.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0
  },
  {
    "id": "gq-901",
    "archetype": "negative",
    "latency_ms": 7136.92,
    "n_retrieved": 5,
    "negative_pass": true
  },
  {
    "id": "gq-902",
    "archetype": "negative",
    "latency_ms": 7372.17,
    "n_retrieved": 5,
    "negative_pass": true
  }
]

How to read these means (TRUTH.md 4.4 — a flattering number ships with its deflating context)

The summary means are NOT all comparable, and two of them are inflated by construction. metrics.py returns a free 1.0 when an expectation set is empty: recall@k when expected_node_ids is empty, and graph_coverage when expected_edges is empty. Aggregation-archetype questions legitimately expect a count rather than nodes, and questions about a graph that holds no edge of the relevant type legitimately expect no edges — so those rows score 1.0 without the retriever having retrieved anything.

Read precision@k, mrr and facet_recall as the load-bearing numbers, and always read the per-question table below before quoting any mean. Do not cite a summary recall or graph_coverage figure without stating how many rows carried an empty expectation set.

Claim protocol

A retrieval-quality improvement claim must cite this artifact AND a prior baseline artifact produced with the SAME gold set version. Gold set and retriever must not change within the same claim window.

← All docsView source on GitHub →