GC ACTIVATION — Edge Latency Benchmark (Workstream B)
Date: 2026-08-21 · Harness: miz-oki-adk-agents/boss/edge_latency_benchmark.py (edge-bench-v1) · Clock: time.perf_counter_ns (monotonic — latency is measured around session.run, never read from model output)
First REAL numbers for the serving path. Scope stated exactly: this bounds the model's compute cost under onnxruntime CPU on the benchmark host; the browser WebGPU path remains UNMEASURED, so every "sub-10ms" browser figure stays a labeled DESIGN TARGET (labels applied in this workstream's diff).
{
"benchmark_version": "edge-bench-v1",
"clock": "time.perf_counter_ns (monotonic)",
"execution_provider": "CPUExecutionProvider",
"onnxruntime_version": "1.29.0",
"model": {
"features": 16,
"hidden": 32,
"outputs": 2,
"bytes": 2700
},
"runs": 2000,
"warmup": 200,
"latency_ms": {
"p50": 0.0108,
"p95": 0.0203,
"p99": 0.0452,
"mean": 0.0188,
"min": 0.0104,
"max": 4.1361
},
"scope": "serving-path CPU measurement on the benchmark host ONLY; the browser WebGPU path is UNMEASURED and its sub-10ms figure stays a design target until a browser benchmark exists"
}
Reproduce: python3 miz-oki-adk-agents/boss/edge_latency_benchmark.py --runs 2000 --warmup 200 (needs onnx, onnxruntime).