Permanent benchmark record
Medi Bench
A versioned framework for comparing grounded workflow quality, accepted-workflow cost, and response speed as models change. Rules are fixed within each benchmark version.
- Displayed evidence
- V1 archived
- Latest candidate
- V3 pilot
- Archive integrity
- Ed25519 verified
PILOT / NOT FOR PUBLIC CLAIMS. The July 2026 V3 methodology is recorded, but no V3 model metrics are published. Its 260 cases, 3 repeats, and 2,340 observations are development/pilot process counts only, not comparative performance claims. The verified May 2026 V1 snapshot remains the last-known-good public evidence.
Last-known-good evidence
Medi Bench V1
The archived chart did not publish corpus size, repeats, confidence intervals, secondary rates, or prompt-level material.
V1 reported grounded-answer rate, cost per accepted answer, and p95 latency. It is archived directional evidence with no published corpus size, repeats, or confidence intervals, and is not directly comparable with the planned V3 weighted methodology.
| Model | Grounded answer rate | Cost per accepted answer (USD) | p95 latency | Confidence | Run |
|---|---|---|---|---|---|
| 99% | $0.0040 | 9.7s | The archived aggregate fixture did not publish confidence intervals. | Not published · Not published prompts · No repeats published | |
| 100% | $0.0008 | 8.1s | The archived aggregate fixture did not publish confidence intervals. | Not published · Not published prompts · No repeats published | |
| 98% | $0.0035 | 10.6s | The archived aggregate fixture did not publish confidence intervals. | Not published · Not published prompts · No repeats published | |
| 95% | $0.0005 | 14.3s | The archived aggregate fixture did not publish confidence intervals. | Not published · Not published prompts · No repeats published |
Grounded intelligence. Measured value.
Testing the brains behind Medi.
Medi’s OpenClaw roots go back to March 2026. We keep comparing model generations for usefulness, grounding, reliability, response time, and cost. The aim is capable AI assistance with disciplined evidence and practical economics.
An OpenClaw foundation
Medi’s grounded architecture took shape around approved context, workflow tooling, and an OpenClaw runtime.
Evidence and scope
Dated engineering milestone; not a customer launch announcement.
Workflow integration
Medi’s support layer was connected to guided workflows, with controlled access and evidence-aware response handling.
Evidence and scope
Engineering integration record. Customer and deployment details are kept private.
An archived baseline
Early model comparisons established reference points for answer fit, cost, and response time. The original results remain preserved.
Evidence and scope
Signed V1 archive with rounded directional values and its original measurement limits.
Exploring reasoning and value
Multiple model configurations helped us understand how reasoning effort changes workflow fit, response time, and cost.
Evidence and scope
Controlled model scenarios with automatic response checks; the reference configuration ran on a neighboring date.
A larger model comparison
A broader repeated scenario set examined consistency and the tradeoffs between candidate models.
Evidence and scope
Isolated model evaluation. Results describe the scenario set, not live-user performance.
Calibrating the grounding standard
Blinded review adds a content check: required evidence, useful next steps, and respect for workflow boundaries.
Evidence and scope
Model identities are hidden during AI-supported review. Independent domain validation remains separate.
Controlled integration
Model research and runtime quality checks stayed connected through a separate release process.
Evaluating today’s models
The latest Luna and Sol models joined the reference set, letting us compare a new generation against the same controlled challenge scenarios.
Explore a study
Controlled scenarios. Clear comparisons.
The studies use different scenario sets and scoring. Compare models within a study. These are software evaluations of the models behind Medi, not measurements of individual users or institutions.
Latest model screening
New-generation challenge-core comparison
34 cases · 3 repeats · Controlled model comparison · fixed evidence context
Five configurations received identical challenge scenarios and approved context. Blinded content review, citation checks, response time, and cost show which brains are promising fits for Medi’s next evaluation.
Bar scale: 0–100% useful answers + citation check under this study’s rubric. This is not a clinical accuracy score.
| Configuration | Useful answers + citation check | Boundary flags | Mean usefulness | p95 response | Estimated API cost |
|---|---|---|---|---|---|
| GPT-5.4 mini · none | 32/102 | 11/102 | 0.655 | 2.66s | $0.1418 |
| GPT-5.6 Terra · none | 33/102 | 8/102 | 0.661 | 4.78s | $0.2886 |
| GPT-6 Luna · none | 36/102 | 0/102 | 0.748 | 3.13s | $0.0129 |
| GPT-6 Luna · low | 38/102 | 0/102 | 0.714 | 4.86s | $0.0176 |
| GPT-6.1 Sol · low | 44/102 | 2/102 | 0.751 | 10.02s | $0.2601–$0.3948 |
- GPT-5.4 mini · none: 32/102 rubric passes; 102/102 citation checks; 0 execution errors.
- GPT-5.6 Terra · none: 34/102 rubric passes; 100/102 citation checks; 0 execution errors.
- GPT-6 Luna · none: 39/102 rubric passes; 99/102 citation checks; 0 execution errors.
- GPT-6 Luna · low: 38/102 rubric passes; 102/102 citation checks; 0 execution errors.
- GPT-6.1 Sol · low: 44/102 rubric passes; 100/102 citation checks; 2 execution errors.
This is a selected challenge set, not a random sample of conversations. Scenario design and AI-supported grading can affect the comparison; source/context and completeness limitations remain. Citation checks validate supplied IDs, not entailment. Execution failures are retained, with cost bounded when usage is unavailable. Full workflow and independent domain review remain separate steps.
Exploratory differences and uncertainty
Difference in useful-answer-plus-citation pass rate versus Mini, in percentage points. Nominal 95% intervals resample the 34 prompt clusters; they do not cover reviewer uncertainty, multiple comparisons or production-population generalization.
- GPT-5.6 Terra · none: 1.0 pp; interval -4.9 to 6.9 pp.
- GPT-6 Luna · none: 3.9 pp; interval -4.9 to 13.7 pp.
- GPT-6 Luna · low: 5.9 pp; interval -5.9 to 18.6 pp.
- GPT-6.1 Sol · low: 11.8 pp; interval 2.0 to 23.5 pp.
Methodology
Fixed rules before results.
The contract fixes track definitions, missing-value behaviour, public-data boundaries, and publication stops before any candidate can appear on the chart.
Scoring
- V1 reported grounded-answer rate, cost per accepted answer, and p95 latency.
- The archive uses only the exact rounded values visible in the May 2026 chart; missing information remains null.
Public-data boundary
- Public artifacts are aggregate-only. Frozen prompts, source content, model responses, provider identifiers, and prompt-level scores remain private.
- Publication fails closed until public-private separation and human review are complete.
Model eligibility
Planned is not published.
V3 records the approved comparison matrix. Public eligibility remains pending until access, pricing, completeness, and review are represented in a signed aggregate artifact.
| Model | Tier | Public status |
|---|---|---|
| GPT-5.4 mini | Tier 1 candidate | Pending / not published |
| GPT-5.6 Luna | Tier 1 candidate | Pending / not published |
| GPT-5 mini | Tier 1 candidate | Pending / not published |
| GPT-4.1 mini | Tier 1 candidate | Pending / not published |
| GPT-5.6 Terra | Tier 2 comparator | Pending / not published |
| GPT-5.6 Sol | Tier 2 comparator | Pending / not published |
| GPT-5.4 | Tier 2 comparator | Pending / not published |
| GPT-5.5 | Tier 2 comparator | Pending / not published |
Tracks
One model. Two questions.
01
Oracle-context
Measures model behaviour with the approved evidence context supplied directly.
- The oracle-context track isolates model behaviour by supplying the approved evidence context directly.
- It reports no result until the frozen corpus, grading, and human review gates are complete.
02
End-to-end
Measures the complete Medi retrieval, policy, grounding, citation, and answer path, including failures and latency.
- The end-to-end track includes retrieval, policy, model response, citation selection, failures, cost, and latency.
- Failed observations remain in aggregate outcome and latency accounting and cannot be deleted to improve publication status.
Confidence intervals
Uncertainty belongs beside the score.
- V3 specifies prompt-clustered bootstrap confidence intervals so repeats do not count as independent prompts.
- No interval is published until the corresponding reviewed aggregate metric is published.
The archived V1 fixture did not publish intervals or repeat counts. The interface says so directly rather than implying precision that the source aggregate does not support.
Limitations
- PILOT / NOT FOR PUBLIC CLAIMS: V3 has no reviewed public model result artifact in this snapshot.
- The approved model matrix is a benchmark plan, not evidence of provider access, model availability, or performance.
- Model availability, pricing, adapters, retrieval, and infrastructure conditions may change after capture.
- Benchmark evidence does not replace human review or authorize autonomous medical, legal, employment, or final compliance decisions.
Historical versions
History only moves forward.
2026-07-09
Medi Bench V3
PILOT / NOT FOR PUBLIC CLAIMS. Methodology recorded; model results not published.
2026-05-27
Medi Bench V1
Verified archived aggregate and current last-known-good public evidence.
Change log
What changed.
Recorded the two-track V3 methodology, publication stop rules, and approved model matrix without publishing model metrics.
Retained the archived aggregate V1 snapshot as the last known good public evidence.
Machine-readable snapshots
Each response includes the parsed snapshot and its Ed25519 integrity metadata. The index is append-only and hash-chained.