Medi

Permanent benchmark record

Medi Bench

A versioned framework for comparing grounded workflow quality, accepted-workflow cost, and response speed as models change. Rules are fixed within each benchmark version.

Displayed evidence
V1 archived
Latest candidate
V3 pilot
Archive integrity
Ed25519 verified

PILOT / NOT FOR PUBLIC CLAIMS. The July 2026 V3 methodology is recorded, but no V3 model metrics are published. Its 260 cases, 3 repeats, and 2,340 observations are development/pilot process counts only, not comparative performance claims. The verified May 2026 V1 snapshot remains the last-known-good public evidence.

Last-known-good evidence

Medi Bench V1

The archived chart did not publish corpus size, repeats, confidence intervals, secondary rates, or prompt-level material.

Medi end-to-end benchmark comparisonGrounded answer rate on the vertical axis and Cost per accepted answer (USD) on the horizontal axis. Larger bubbles indicate faster p95 latency. Use arrow keys to move between models.Cost per accepted answer (USD)Grounded answer rate

V1 reported grounded-answer rate, cost per accepted answer, and p95 latency. It is archived directional evidence with no published corpus size, repeats, or confidence intervals, and is not directly comparable with the planned V3 weighted methodology.

Equivalent table for the end-to-end chart
ModelGrounded answer rateCost per accepted answer (USD)p95 latencyConfidenceRun
99%$0.00409.7sThe archived aggregate fixture did not publish confidence intervals.Not published · Not published prompts · No repeats published
100%$0.00088.1sThe archived aggregate fixture did not publish confidence intervals.Not published · Not published prompts · No repeats published
98%$0.003510.6sThe archived aggregate fixture did not publish confidence intervals.Not published · Not published prompts · No repeats published
95%$0.000514.3sThe archived aggregate fixture did not publish confidence intervals.Not published · Not published prompts · No repeats published

Grounded intelligence. Measured value.

Testing the brains behind Medi.

Medi’s OpenClaw roots go back to March 2026. We keep comparing model generations for usefulness, grounding, reliability, response time, and cost. The aim is capable AI assistance with disciplined evidence and practical economics.

  1. An OpenClaw foundation

    Medi’s grounded architecture took shape around approved context, workflow tooling, and an OpenClaw runtime.

    Evidence and scope

    Dated engineering milestone; not a customer launch announcement.

  2. Workflow integration

    Medi’s support layer was connected to guided workflows, with controlled access and evidence-aware response handling.

    Evidence and scope

    Engineering integration record. Customer and deployment details are kept private.

  3. An archived baseline

    Early model comparisons established reference points for answer fit, cost, and response time. The original results remain preserved.

    Evidence and scope

    Signed V1 archive with rounded directional values and its original measurement limits.

  4. Exploring reasoning and value

    Multiple model configurations helped us understand how reasoning effort changes workflow fit, response time, and cost.

    Evidence and scope

    Controlled model scenarios with automatic response checks; the reference configuration ran on a neighboring date.

  5. A larger model comparison

    A broader repeated scenario set examined consistency and the tradeoffs between candidate models.

    Evidence and scope

    Isolated model evaluation. Results describe the scenario set, not live-user performance.

  6. Calibrating the grounding standard

    Blinded review adds a content check: required evidence, useful next steps, and respect for workflow boundaries.

    Evidence and scope

    Model identities are hidden during AI-supported review. Independent domain validation remains separate.

  7. Controlled integration

    Model research and runtime quality checks stayed connected through a separate release process.

  8. Evaluating today’s models

    The latest Luna and Sol models joined the reference set, letting us compare a new generation against the same controlled challenge scenarios.

Explore a study

Controlled scenarios. Clear comparisons.

The studies use different scenario sets and scoring. Compare models within a study. These are software evaluations of the models behind Medi, not measurements of individual users or institutions.

Latest model screening

New-generation challenge-core comparison

34 cases · 3 repeats · Controlled model comparison · fixed evidence context

Five configurations received identical challenge scenarios and approved context. Blinded content review, citation checks, response time, and cost show which brains are promising fits for Medi’s next evaluation.

Bar scale: 0–100% useful answers + citation check under this study’s rubric. This is not a clinical accuracy score.

Exact counts and response cost
ConfigurationUseful answers + citation checkBoundary flagsMean usefulnessp95 responseEstimated API cost
GPT-5.4 mini · none32/10211/1020.6552.66s$0.1418
GPT-5.6 Terra · none33/1028/1020.6614.78s$0.2886
GPT-6 Luna · none36/1020/1020.7483.13s$0.0129
GPT-6 Luna · low38/1020/1020.7144.86s$0.0176
GPT-6.1 Sol · low44/1022/1020.75110.02s$0.2601–$0.3948
  • GPT-5.4 mini · none: 32/102 rubric passes; 102/102 citation checks; 0 execution errors.
  • GPT-5.6 Terra · none: 34/102 rubric passes; 100/102 citation checks; 0 execution errors.
  • GPT-6 Luna · none: 39/102 rubric passes; 99/102 citation checks; 0 execution errors.
  • GPT-6 Luna · low: 38/102 rubric passes; 102/102 citation checks; 0 execution errors.
  • GPT-6.1 Sol · low: 44/102 rubric passes; 100/102 citation checks; 2 execution errors.

This is a selected challenge set, not a random sample of conversations. Scenario design and AI-supported grading can affect the comparison; source/context and completeness limitations remain. Citation checks validate supplied IDs, not entailment. Execution failures are retained, with cost bounded when usage is unavailable. Full workflow and independent domain review remain separate steps.

Exploratory differences and uncertainty

Difference in useful-answer-plus-citation pass rate versus Mini, in percentage points. Nominal 95% intervals resample the 34 prompt clusters; they do not cover reviewer uncertainty, multiple comparisons or production-population generalization.

  • GPT-5.6 Terra · none: 1.0 pp; interval -4.9 to 6.9 pp.
  • GPT-6 Luna · none: 3.9 pp; interval -4.9 to 13.7 pp.
  • GPT-6 Luna · low: 5.9 pp; interval -5.9 to 18.6 pp.
  • GPT-6.1 Sol · low: 11.8 pp; interval 2.0 to 23.5 pp.

Methodology

Fixed rules before results.

The contract fixes track definitions, missing-value behaviour, public-data boundaries, and publication stops before any candidate can appear on the chart.

Scoring

  • V1 reported grounded-answer rate, cost per accepted answer, and p95 latency.
  • The archive uses only the exact rounded values visible in the May 2026 chart; missing information remains null.

Public-data boundary

  • Public artifacts are aggregate-only. Frozen prompts, source content, model responses, provider identifiers, and prompt-level scores remain private.
  • Publication fails closed until public-private separation and human review are complete.

Model eligibility

Planned is not published.

V3 records the approved comparison matrix. Public eligibility remains pending until access, pricing, completeness, and review are represented in a signed aggregate artifact.

July 2026 V3 model matrix
ModelTierPublic status
GPT-5.4 miniTier 1 candidatePending / not published
GPT-5.6 LunaTier 1 candidatePending / not published
GPT-5 miniTier 1 candidatePending / not published
GPT-4.1 miniTier 1 candidatePending / not published
GPT-5.6 TerraTier 2 comparatorPending / not published
GPT-5.6 SolTier 2 comparatorPending / not published
GPT-5.4Tier 2 comparatorPending / not published
GPT-5.5Tier 2 comparatorPending / not published

Tracks

One model. Two questions.

01

Oracle-context

Measures model behaviour with the approved evidence context supplied directly.

  • The oracle-context track isolates model behaviour by supplying the approved evidence context directly.
  • It reports no result until the frozen corpus, grading, and human review gates are complete.

02

End-to-end

Measures the complete Medi retrieval, policy, grounding, citation, and answer path, including failures and latency.

  • The end-to-end track includes retrieval, policy, model response, citation selection, failures, cost, and latency.
  • Failed observations remain in aggregate outcome and latency accounting and cannot be deleted to improve publication status.

Confidence intervals

Uncertainty belongs beside the score.

  • V3 specifies prompt-clustered bootstrap confidence intervals so repeats do not count as independent prompts.
  • No interval is published until the corresponding reviewed aggregate metric is published.

The archived V1 fixture did not publish intervals or repeat counts. The interface says so directly rather than implying precision that the source aggregate does not support.

Limitations

  • PILOT / NOT FOR PUBLIC CLAIMS: V3 has no reviewed public model result artifact in this snapshot.
  • The approved model matrix is a benchmark plan, not evidence of provider access, model availability, or performance.
  • Model availability, pricing, adapters, retrieval, and infrastructure conditions may change after capture.
  • Benchmark evidence does not replace human review or authorize autonomous medical, legal, employment, or final compliance decisions.

Historical versions

History only moves forward.

  1. 2026-07-09

    Medi Bench V3

    PILOT / NOT FOR PUBLIC CLAIMS. Methodology recorded; model results not published.

  2. 2026-05-27

    Medi Bench V1

    Verified archived aggregate and current last-known-good public evidence.

Change log

What changed.

V3

Recorded the two-track V3 methodology, publication stop rules, and approved model matrix without publishing model metrics.

V1

Retained the archived aggregate V1 snapshot as the last known good public evidence.

Machine-readable snapshots

Each response includes the parsed snapshot and its Ed25519 integrity metadata. The index is append-only and hash-chained.

Trust by design

Trust, in plain sight.

Visit Trust Centre →

Canada-hosted

Customer data stays in Canada by default.

Human oversight

Consequential and uncertain outcomes route to people.

Audit-ready evidence

Evidence and reviewer actions stay traceable.

Protected access

RBAC and MFA protect privileged access.