GraphRAG Retrieval-Quality Evaluation — 2026-08-10

Status: [MEASURED — live endpoint] Gold set: gold_v1.jsonl · Backend: endpoint · Generated: 2026-08-10T12:37:47.235985+00:00

Summary

Metric Value
n_questions 9
n_scored 7
n_negative 2
precision@5 0.0
recall@5 0.2857
precision@10 0.0143
recall@10 0.4286
mrr 0.0238
graph_coverage 0.7143
facet_recall 0.2143
faithfulness 0.249
faithfulness_scored 7
faithfulness_failed 0
negative_pass_rate 1.0
latency_p50_ms 4525.41
latency_p95_ms 4672.03

Per-question results

[
  {
    "id": "gq-001",
    "archetype": "journey_traversal",
    "latency_ms": 4525.41,
    "n_retrieved": 10,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.1,
    "recall@10": 1.0,
    "mrr": 0.1667,
    "graph_coverage": 0.0,
    "facet_recall": 1.0,
    "faithfulness": 0.4
  },
  {
    "id": "gq-002",
    "archetype": "journey_traversal",
    "latency_ms": 4408.54,
    "n_retrieved": 6,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 0.0,
    "facet_recall": 0.5,
    "faithfulness": 0.35294117647058826
  },
  {
    "id": "gq-101",
    "archetype": "causal_path",
    "latency_ms": 4410.35,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0,
    "faithfulness": 0.18181818181818182
  },
  {
    "id": "gq-102",
    "archetype": "causal_path",
    "latency_ms": 4538.12,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0,
    "faithfulness": 0.0
  },
  {
    "id": "gq-201",
    "archetype": "intent_explanation",
    "latency_ms": 4572.03,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 0.0,
    "precision@10": 0.0,
    "recall@10": 0.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0,
    "faithfulness": 0.0
  },
  {
    "id": "gq-301",
    "archetype": "aggregation",
    "latency_ms": 4424.13,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 1.0,
    "precision@10": 0.0,
    "recall@10": 1.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0,
    "faithfulness": 0.4444444444444444
  },
  {
    "id": "gq-302",
    "archetype": "aggregation",
    "latency_ms": 4550.67,
    "n_retrieved": 5,
    "precision@5": 0.0,
    "recall@5": 1.0,
    "precision@10": 0.0,
    "recall@10": 1.0,
    "mrr": 0.0,
    "graph_coverage": 1.0,
    "facet_recall": 0.0,
    "faithfulness": 0.36363636363636365
  },
  {
    "id": "gq-901",
    "archetype": "negative",
    "latency_ms": 4397.71,
    "n_retrieved": 5,
    "negative_pass": true
  },
  {
    "id": "gq-902",
    "archetype": "negative",
    "latency_ms": 4672.03,
    "n_retrieved": 5,
    "negative_pass": true
  }
]

How to read these means (TRUTH.md 4.4 — a flattering number ships with its deflating context)

The summary means are NOT all comparable, and two of them are inflated by construction. metrics.py returns a free 1.0 when an expectation set is empty: recall@k when expected_node_ids is empty, and graph_coverage when expected_edges is empty. Aggregation-archetype questions legitimately expect a count rather than nodes, and questions about a graph that holds no edge of the relevant type legitimately expect no edges — so those rows score 1.0 without the retriever having retrieved anything.

Read precision@k, mrr and facet_recall as the load-bearing numbers, and always read the per-question table below before quoting any mean. Do not cite a summary recall or graph_coverage figure without stating how many rows carried an empty expectation set.

Claim protocol

A retrieval-quality improvement claim must cite this artifact AND a prior baseline artifact produced with the SAME gold set version. Gold set and retriever must not change within the same claim window.

← All docsView source on GitHub →