GC ACTIVATION — Edge Latency Benchmark (Workstream B)

Date: 2026-08-21 · Harness: miz-oki-adk-agents/boss/edge_latency_benchmark.py (edge-bench-v1) · Clock: time.perf_counter_ns (monotonic — latency is measured around session.run, never read from model output)

First REAL numbers for the serving path. Scope stated exactly: this bounds the model's compute cost under onnxruntime CPU on the benchmark host; the browser WebGPU path remains UNMEASURED, so every "sub-10ms" browser figure stays a labeled DESIGN TARGET (labels applied in this workstream's diff).

{
  "benchmark_version": "edge-bench-v1",
  "clock": "time.perf_counter_ns (monotonic)",
  "execution_provider": "CPUExecutionProvider",
  "onnxruntime_version": "1.29.0",
  "model": {
    "features": 16,
    "hidden": 32,
    "outputs": 2,
    "bytes": 2700
  },
  "runs": 2000,
  "warmup": 200,
  "latency_ms": {
    "p50": 0.0108,
    "p95": 0.0203,
    "p99": 0.0452,
    "mean": 0.0188,
    "min": 0.0104,
    "max": 4.1361
  },
  "scope": "serving-path CPU measurement on the benchmark host ONLY; the browser WebGPU path is UNMEASURED and its sub-10ms figure stays a design target until a browser benchmark exists"
}

Reproduce: python3 miz-oki-adk-agents/boss/edge_latency_benchmark.py --runs 2000 --warmup 200 (needs onnx, onnxruntime).

← All docsView source on GitHub →