Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Introduction

This is the reader’s guide to mcp-conformance, an independent, trace-based conformance toolkit for the Model Context Protocol (MCP). It judges whether an MCP implementation conforms to a specific spec revision — 2025-11-25 by default, 2026-07-28 behind a feature flag, or both at once in a single pass — and, just as importantly, it is built so that an ecosystem with its own authoritative test suite has reason to trust the verdict.

Status: pre-release. The crates are published on crates.io at 0.x; the public API and the verdicts it produces may still change between minor releases, and the changelog says so explicitly when they do. This book tracks main.

The one idea

The validator’s input is a trace: an ordered, serialized record of a protocol interaction. Its engine is a pure function — &[TraceEvent] -> Report — with no network, no clock, and no I/O of its own. Capturing the trace is somebody else’s job; judging it is the validator’s only job. That split is what makes a verdict deterministic, replayable, language- and SDK-agnostic, and — because it is reproducible from a committed file — auditable rather than a “trust us.”

The credibility mechanism is the agreement check: on every CI run the reference server is driven by the official conformance suite, the same sessions are captured as traces, and this toolkit’s verdicts are diffed against the official runner’s. Agreement is the default; an unexplained divergence fails the build. A validator whose verdicts are checked against the recognized authority on every commit is calibrated, not merely asserted to be correct.

The crates

CrateWhat it is
mcp-conformance-coreThe spec as data: the requirement registry, the trace schema, and RFC 8785 canonical JSON. Serde-only; no protocol SDK.
mcp-trace-validatorThe deterministic judgment engine and its CLI (human / JSON / JUnit reports, documented exit codes).
mcp-everything-serverA reference server on the official rmcp SDK that passes the suite’s server scenarios, with a session tap that records traces for the agreement check.
mcp-reference-hostA reference host (client) that passes the suite’s client scenarios and captures host-side traces.

How this book is organized

Each chapter is a curated entry point. The planning documents — charter, ecosystem register, architecture, conformance strategy, engineering standards, security model, roadmap, and decision records — remain the single source of truth, and every chapter links to the authoritative document behind it.

  • Architecture — why judge traces instead of live sessions, and the design trade-offs that follow.
  • The trace format — the JSON Lines schema, with a worked example embedded from the README so it cannot drift.
  • Two revisions at once — how one trace is judged against 2025-11-25 and 2026-07-28 together, and what a migration looks like when it is measured instead of estimated.
  • The trace corpus — the golden good/violation fixtures and their provenance, embedded from corpus/README.md.
  • Conformance results — the published tier-gap measurements of real SDKs against the official suite.

Architecture

This chapter orients you in the design. The authoritative treatment — ten decisions, each with its cost — is the standalone design note, written for an upstream audience:

📄 Design note: trace-validation architecture and trade-offs

The internal workspace layout and dependency rules live in 02-architecture.md.

The foundational split: capture, then judge

The validator judges a recorded trace, not a live session. Its engine is a pure function with no network, clock, or I/O. Whoever owns the socket — the server’s session tap, the host’s capture wrapper, or any external proxy — produces the trace, assigning the total-order seq at capture. The judge only judges.

This buys determinism, replayability, SDK-agnostic input, and auditable verdicts. It costs one hard thing: capture fidelity becomes a first-class engineering problem — interleaving, partial writes, and SSE framing all have to be recorded faithfully by the component that holds the bytes.

The decisions that follow, in brief

  • The judge never links the SDK it judges. mcp-trace-validator does not depend on rmcp. A validator built on the SDK under test would inherit that SDK’s interpretation of the spec as ground truth — exactly the thing being examined. The cost is re-expressing the protocol as our own data.
  • The spec as data. Every normative clause becomes a registry record with its RFC 2119 level, actor, a verbatim source quote, an optional capability gate, and either one or more mechanical checks or a written reason it cannot be judged from a trace. There is no “not yet looked at” state. A scheduled job re-verifies every quote against the live spec text so the registry cannot silently drift.
  • Determinism by canonicalization. All JSON is canonicalized (RFC 8785) before any comparison, so a report is byte-identical across platforms and runs. Correctness here is a fixpoint property, tested as one — a real bug hid until a fuzzer first generated -0.0.
  • A pass has to be earned. A requirement gated on a capability that was never negotiated is reported as not-applicable, and one whose subject matter the session never carried is reported not-observed — never as a vacuous pass. Every check counts the subjects it considered, so the report can tell “complied with” from “never came up”. Inflating a pass rate with checks that had nothing to look at is how a conformance tool loses credibility.
  • Calibration against the authority. The agreement check (see the Introduction) diffs our verdicts against the official runner’s on every CI run; an unexplained divergence fails the build, and a stale recorded divergence fails it too.
  • Falsifiability. Every check is killed by at least one committed injected-violation trace; a check with no killer trace is unfalsifiable. The corpus is both the regression suite and the evidence each check does what it claims.
  • Verdicts are a contract. A change that makes a previously-passing trace fail is a breaking change even when no function signature moved; it is named in the changelog and the version is bumped accordingly. API-signature compatibility is additionally checked mechanically against the published crates.io baseline.

The trace format

A trace is JSON Lines: one event per line. Each event carries a capture-assigned seq (the only ordering authority — never inferred later), a direction (client-to-server / server-to-client), a transport, and a kind:

  • message events hold the JSON-RPC payload verbatim;
  • http events record the conformance-relevant headers, a response’s status, and a client request’s method — Streamable HTTP binds different obligations to POST, GET, and DELETE, so a clause addressed to one of them is judged only where the recording says which it was; and
  • lifecycle events mark transport open/close.

The full schema, including the redaction rules that keep credential-bearing headers out of a trace by construction, is in 02-architecture.md and 05-security-model.md.

One worked example

The example below is embedded verbatim from the README, where a test (readme_examples.rs) pins it to the validator’s real output — so what you read here cannot drift from what the tool actually produces. It is a session that reuses a request ID, and the verdict that catches it:

{"seq":0,"direction":"client-to-server","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"my-host","version":"1.0.0"}}}}
{"seq":1,"direction":"server-to-client","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-11-25","capabilities":{"tools":{}},"serverInfo":{"name":"my-server","version":"1.0.0"}}}}
{"seq":2,"direction":"client-to-server","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","method":"notifications/initialized"}}
{"seq":3,"direction":"client-to-server","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","id":1,"method":"tools/list"}}

and the validator answers with the violated clause, verbatim from the spec via the registry, addressed to the offending event:

  FAIL  BASE-003 (MUST NOT)
        seq 3: request "tools/list" reuses id 1, already used by the same party at seq 0
totals: 17 pass, 1 fail, 0 warn, 87 excluded, 0 unsupported, 6 not applicable, 31 not observed
verdict: fail

The totals line distinguishes the verdict’s components, and two of them are not passes. The not-applicable rows are capability-gated requirements this session never negotiated (the resources and prompts clauses); the not-observed rows are clauses whose subject matter never appeared at all — nothing was paginated, no binary content was sent, no error was returned. A four-line session earns 17 passes, not a hundred. See Architecture for why that distinction is load-bearing.

Two revisions at once

MCP is a dated specification. 2025-11-25 is what implementations ship against today; 2026-07-28 removes the initialize handshake outright (SEP-2575), makes every request carry its own _meta envelope, and adds caching hints, multi-round-trip requests, subscriptions, and a discovery probe. That is not a version bump anyone should take on faith.

So this toolkit judges both, and can judge one trace against both in a single pass. A migration stops being an inventory of changes somebody read in a changelog and becomes a measurement.

What “both” means here

2025-11-252026-07-28
Registry entries142272
Judged by a named check55125
Carrying a documented exclusion87147
Shipped by defaultyesbehind the draft-2026-07-28 feature

The two registries are extracted per revision rather than sharing entries: each clause is quoted from its own revision’s published text, and a clause restated with different words gets its own ID. 2025-11-25’s BASE-003 forbids reusing a request ID within a session; 2026-07-28’s BASE-045 forbids reusing one that is still in flight, so reuse after a response is now legal. Those are two clauses, not one clause with a footnote — and pointing the new one at the old one’s check would have reported conforming traces as violations.

The cost of that method is that no clause is in force at both revisions, so a side-by-side report never shows one clause passing at one revision and failing at the other. What it shows instead is presence, which is what a migration actually consists of.

A trace judged against both

This is a conforming 2025-11-25 handshake — the client offers a version, the server answers, the client confirms:

{"seq":0,"direction":"client-to-server","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","id":1,"method":"initialize","params":{"protocolVersion":"2025-11-25","capabilities":{},"clientInfo":{"name":"corpus-client","version":"0.1.0"}}}}
{"seq":1,"direction":"server-to-client","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","id":1,"result":{"protocolVersion":"2025-11-25","capabilities":{},"serverInfo":{"name":"corpus-server","version":"0.1.0"}}}}
{"seq":2,"direction":"client-to-server","transport":"stdio","kind":"message","payload":{"jsonrpc":"2.0","method":"notifications/initialized"}}

Judged against both revisions at once:

$ mcp-trace-validator validate --revision 2025-11-25 --revision 2026-07-28 handshake.jsonl

Four rows out of the report say the whole thing:

  LIFE-001   (MUST)  2025-11-25=pass  2026-07-28=absent  *differs
  DISC-001   (MUST)  2025-11-25=absent  2026-07-28=not-observed  *differs
  BASE-030   (MUST)  2025-11-25=absent  2026-07-28=fail  *differs
  BASE-045   (MUST NOT)  2025-11-25=absent  2026-07-28=pass  *differs
  • LIFE-001 — “the client MUST initiate the lifecycle” — is a clause the newer revision removes. It reads pass then absent.
  • DISC-001 — the server/discover probe that replaces the handshake — is a clause the newer revision adds, and this session carried nothing it binds to, so it reads not-observed rather than counting as a pass.
  • BASE-030 — every request must carry an io.modelcontextprotocol/clientCapabilities _meta member — is a clause the newer revision adds that this session breaks. It reads fail, and it is the migration work this trace has actually discovered.
  • BASE-045 — the narrowed request-ID rule — reads pass, because reusing an ID after its response is exactly what the new text permits.

And the verdict splits:

per revision:
  2025-11-25: 17 pass, 0 fail, 0 warn, 87 excluded, 0 unsupported, 14 not applicable, 24 not observed — verdict pass
  2026-07-28: 14 pass, 8 fail, 0 warn, 147 excluded, 0 unsupported, 0 not applicable, 103 not observed — verdict fail
overall verdict: fail

A conforming session today, and eight concrete failures tomorrow. Both summary lines account for every clause in their revision — the counts add up to 142 and 272 — because a line that quietly omits an outcome is a line that overstates what was measured.

The example above is executed by a test (book_examples.rs) against the real validator on every cargo test, so this page cannot drift from what the tool actually prints.

Judged against the wrong revision

Naming a revision is a choice, and the default is 2025-11-25. Point the validator at a 2026-07-28 recording without saying so and every clause the two revisions disagree about becomes a finding: the session opens with tools/list rather than initialize, so LIFE-001 fails; it reuses request ids after their responses, so BASE-003 fails. Both are correct answers to the question that was asked, and neither is a defect in the implementation.

A trace is not silent about which revision it belongs to — the handshake states it, a stateless session states it in every request’s _meta, and the HTTP transport states it in a header. This session says 2026-07-28 twice over:

{"seq":0,"direction":"client-to-server","transport":"streamable-http","kind":"http","method":"POST","headers":{"accept":"application/json, text/event-stream","mcp-protocol-version":"2026-07-28"}}
{"seq":1,"direction":"client-to-server","transport":"streamable-http","kind":"message","payload":{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{"_meta":{"io.modelcontextprotocol/protocolVersion":"2026-07-28","io.modelcontextprotocol/clientCapabilities":{}}}}}
{"seq":2,"direction":"server-to-client","transport":"streamable-http","kind":"message","payload":{"jsonrpc":"2.0","id":1,"result":{"tools":[],"resultType":"complete"}}}

so validating it without naming a revision says so, above the rows and again under the verdict:

  NOTE  this session declares protocol revision 2026-07-28, not 2025-11-25.
        Every outcome here judges it against rules it was not playing by;
        re-run with `--revision 2026-07-28` to judge it against its own.

It is deliberately quiet. A session that proposed one revision and negotiated another has touched both, so judging it against either draws no note; a version a server refused states nothing, because the session never ran under it; and a recording that declares no revision at all gets no note, since there is nothing to disagree with.

The four ways a clause is not a pass

A side-by-side report is where these stop being pedantry, because all four appear at once and they mean different things:

ReadsMeans
absentThe clause does not exist at that revision.
excludedIt exists, but no recorded trace can judge it — the reason is stated per entry.
not-applicableIt exists and is judgeable, but it is gated on a capability this session never negotiated.
not-observedIt exists, is judgeable, is not gated — and the session simply carried nothing it binds to.

None of them is a pass, and the report never quietly promotes one. That rule is the whole reason the numbers here are smaller than a conformance tool’s numbers usually are.

What is measured at 2026-07-28 today

  • The reference server serves the stateless surface with --protocol-version 2026-07-28, and the official suite’s 2026-07-28 scenarios score 41 passing / 0 failing against it.
  • Five committed captures — a conforming session over each transport, a probe session of deliberately malformed requests, and the official runner’s two — evidence 114 of the 125 judgeable clauses between them. What no capture reaches is named, one clause at a time, in the corpus chapter.
  • cargo xtask draft-readiness re-runs that measurement and ratchets it against a committed baseline: any change in either direction fails the build, so progress toward the next revision is recorded in the commit that earned it rather than estimated.

That ratchet has already paid for itself in the direction nobody plans for. Through suite 0.2.0-alpha.9 the runner scored the 2025-11-25 server and the 2026-07-28 server identically — its scenarios exercise features, not revisions, so it could not tell them apart. This project’s registry could, and flagged the 2025-11-25 server for returning cacheable results without the ttlMs hint the new revision requires (CACH-001). Six weeks later alpha.11 added a schema check and found the same clause, independently, from the other side. The registry’s finding is recorded as superseded rather than deleted, because the supersession is itself the result: a negative finding about an instrument expires when the instrument improves.

Full detail is in the ecosystem register (row 1.5i) and the conformance strategy.

Trace corpus

Fixtures for the golden-corpus tests (crates/mcp-trace-validator/tests/golden.rs):

  • good/ — sessions that must validate with verdict pass.
  • violations/ — single-issue sessions, each falsifying at least one check; named after the requirement whose check they exist to kill.
  • draft/{good,violations}/ — the same two kinds for revision 2026-07-28, judged against that revision’s registry behind the draft-2026-07-28 feature.
  • draft/captured/ — sessions recorded off the wire from implementations this repository did not write. No verdict is asserted for them; their whole report is byte-pinned instead. See Captured versus authored below.
  • golden/ — the byte-pinned expected report for every trace, with golden/draft/ holding the 2026-07-28 corpus’s. Regenerate only via cargo xtask bless and review the diff like code.
  • golden/exclusions/ — one ledger per revision, holding the clauses that revision’s registry declines to judge from a trace, and why.

The revisions keep separate golden directories because both name traces after requirement IDs drawn from their own registry: base-045-… is one clause at 2025-11-25 and a different one at 2026-07-28, so a shared directory would let one revision’s golden answer for the other’s trace.

What a golden pins

A report states two kinds of fact, and they are pinned separately (ADR-0013).

golden/<stem>.json holds what the trace decided: the revision, the whole totals, and every pass / fail / warn / not-observed / not-applicable / unsupported row, in registry order.

golden/exclusions/<revision>.json holds what the registry decided: the excluded rows, whose outcome and prose come from the registry alone and are identical for every trace judged against it — engine::build_row never consults the session for them. One distinct set per revision, and before the split every golden restated its revision’s whole set: at the time that was 59% of the corpus’s lines.

2025-11-252026-07-28
Goldens5780
Excluded rows the ledger holds87147
Distinct not-observed sets across them2969

Every cell is verified against the corpus by cargo xtask ci (xtask::coverage::corpus). It is checked rather than trusted because it had already drifted twice: this paragraph is prose, and prose does not recount itself.

Nothing stopped being pinned. Splicing a golden back together with its ledger reproduces the whole report, and assert_reconstructs_the_full_report proves it does on every trace, on every run — so a judged row that went missing, a ledger that drifted from its registry, or rows that interleave in any order but the registry’s all fail the suite. totals.excluded stays in every golden as the per-trace tie back to the ledger: a clause entering or leaving the excluded set moves every golden for that revision, one line each.

Only excluded collapses. not-observed and not-applicable are per-trace evidence, and the table’s third row is the measure of it: near one distinct set per golden, in both revisions. Collapsing those would delete exactly what makes the captures differentiate each other.

Violation naming contract

A violation trace is named area-nnn-<slug>.jsonl and must produce a Fail/Warn row with findings for exactly requirement AREA-NNN — the golden harness enforces the attribution by name (violation_traces_fail_and_match_goldens, and draft::draft_violation_traces_fail_and_match_goldens for draft/), so a defect re-routed to a different requirement cannot re-bless silently. A stem that does not begin with a requirement ID fails the suite loudly.

Captured versus authored

Everything in good/, violations/ and draft/{good,violations}/ is authored: hand-written to exercise one clause. That is what makes them precise, and it is also their limit — an authored fixture can only ever confirm its author’s reading of the specification. A check that is wrong in the same way the author is wrong passes its unit tests and its corpus alike, and neither notices.

draft/captured/ exists to close that gap. Its traces are real recordings of traffic between implementations neither of which was written to satisfy these checks, so they can disagree with us — and when they do, the disagreement is information rather than a test failure. Their goldens are pinned without any pass/fail expectation, because a real implementation is whatever it actually is; what the golden protects is which requirements fire, on which events, with what detail. A check that starts misfiring on real traffic moves it — which is precisely the half of the report that stays in the per-trace file.

Provenance ledger

Every trace’s origin, in one reviewable place that survives history rewrites (the invariant test in golden.rs fails if a trace is added without a ledger row). All current traces share one provenance: hand-authored for this repository as synthetic sessions (no third-party traffic, no recorded production data), written against the 2025-11-25 spec text fetched live from modelcontextprotocol.io on 2026-06-09 (re-verified clause-by-clause against the live text on 2026-06-11) and validated against the embedded registry at the commit that introduced them.

Traces under draft/captured/ are the exception and record their own origin below: the implementations on both ends, the tool that recorded them, and when.

good/

TraceExercises
http-session.jsonlStreamable HTTP session: session-ID assignment and echo, MCP-Protocol-Version headers, Accept/Content-Type discipline, ping (TRAN-011/013/017/018/025/029/039/040 pass paths). It carries all three client request forms, which is what lets the two Accept clauses be told apart: POSTs offering both media types (TRAN-025), a standalone-stream GET offering text/event-stream (TRAN-039), and a session-terminating DELETE carrying Accept: */*. That last exchange is a regression guard — the clause defining session termination places no Accept obligation on it, but until 2026-08-20 both clauses were applied to every client request whatever its method, so this exact teardown drew two MUST-level failures from a client that had violated nothing.
stdio-feature-session.jsonlEvery feature area conformant in one session: tools (incl. outputSchema + structuredContent, and a call returning an embedded resource against a declared resources capability — TOOL-009’s pass path), resources (read/blob/subscribe/updated), prompts (text/image/audio/embedded), logging, completion, pagination cursor flow, and a well-formed params._meta key (BASE-019/020’s pass path)
stdio-full-session.jsonlHandshake plus ping, tools/list, tools/call over stdio
stdio-minimal-init.jsonlSmallest conformant session: the three-message handshake

violations/

Each file injects exactly the violation its name states; golden/ shows every row the trace decided, including any intrinsic secondary findings the injected defect causes (a malformed notification also fails lifecycle accounting, for example).

TraceFalsifies
base-001-request-id-boolean.jsonlBASE-001
base-002-request-id-null.jsonlBASE-002
base-003-request-id-reuse.jsonlBASE-003
base-004-request-answered-twice-cross-flavor.jsonlBASE-004 (the second of an error+result double-answer)
base-004-result-unknown-id.jsonlBASE-004
base-005-notification-with-id.jsonlBASE-005
base-006-error-missing-message.jsonlBASE-006
base-007-error-code-float.jsonlBASE-007
base-008-jsonrpc-version.jsonlBASE-008
base-009-error-unknown-id.jsonlBASE-009
base-010-response-without-result.jsonlBASE-010
base-019-meta-key-bad-prefix.jsonlBASE-019, BASE-020 (shared base.meta-key-format check)
comp-001-capability-undeclared.jsonlCOMP-001
life-001-first-message-not-initialize.jsonlLIFE-001
life-002-initialize-missing-protocolversion.jsonlLIFE-002
life-003-missing-initialized.jsonlLIFE-003
life-004-client-request-before-init-response.jsonlLIFE-004
life-005-server-request-before-initialized.jsonlLIFE-005
life-006-result-version-invalid.jsonlLIFE-006
life-007-initialize-protocolversion-not-string.jsonlLIFE-007
life-009-undeclared-capability-use.jsonlLIFE-009
life-010-initialize-result-missing-capabilities.jsonlLIFE-010
log-001-capability-undeclared.jsonlLOG-001
page-002-cursor-never-issued.jsonlPAGE-002
page-002-fabricated-cursor-rejected.jsonlPAGE-002, and with it PAGE-003’s pass path: the client fabricates a cursor and the server answers -32602 rather than honouring it. The two halves of one event belong to different parties, so the trace falsifies the client’s clause and evidences the server’s.
prom-001-capability-undeclared.jsonlPROM-001
prom-003-image-data-invalid.jsonlPROM-003
prom-004-audio-data-invalid.jsonlPROM-004
prom-005-embedded-resource-malformed.jsonlPROM-005
prom-008-required-argument-unvalidated.jsonlPROM-008
res-001-capability-undeclared.jsonlRES-001
res-004-uri-bad-scheme.jsonlRES-004
res-006-blob-not-base64.jsonlRES-006
tool-001-capability-undeclared.jsonlTOOL-001
tool-003-input-schema-null.jsonlTOOL-003
tool-005-name-length.jsonlTOOL-005
tool-006-name-charset.jsonlTOOL-006
tool-008-name-duplicate.jsonlTOOL-008
tool-009-embedded-resource-without-capability.jsonlTOOL-009
tool-010-structured-without-text.jsonlTOOL-010
tool-011-output-schema-no-structured-result.jsonlTOOL-011
tran-004-stdout-invalid-message.jsonlTRAN-004
tran-005-stdin-invalid-message.jsonlTRAN-005
tran-011-session-id-invisible-ascii.jsonlTRAN-011
tran-013-session-id-not-echoed.jsonlTRAN-013
tran-017-protocol-version-header-missing.jsonlTRAN-017
tran-018-protocol-version-mismatched.jsonlTRAN-018
tran-025-accept-header-missing.jsonlTRAN-025 (a POST with no Accept header at all)
tran-025-accept-header-json-missing.jsonlTRAN-025 (a POST offering text/event-stream only — the half of the clause that read as a pass until the check learned the request method)
tran-039-get-accept-header-missing.jsonlTRAN-039 (a standalone-stream GET offering application/json only)
tran-049-message-not-posted.jsonlTRAN-049 (a ping sent by PUT after a clean handshake)
tran-026-http-post-batch.jsonlTRAN-026 (a batch array POSTed after a clean handshake)
tran-029-content-type-unexpected.jsonlTRAN-029, TRAN-040 (shared transport.success-content-type check)

2026-07-28 captured (corpus/draft/captured/)

Two recordings of the same client driving the same scenarios, differing in one variable: which revision the server was serving. That is what makes the pair worth committing — a single capture shows what one implementation does, while a matched pair shows what the revision costs, and every difference between the two reports is attributable to that one change.

Fieldofficial-suite-2026-07-28-scenarios.jsonlofficial-suite-2026-07-28-stateless.jsonl
ClientThe official MCP conformance suite, 0.2.0-alpha.11, driving its 2026-07-28 scenario set (the pin cargo xtask draft-readiness holds)The same client, the same scenarios, the same run
Servermcp-everything-server serving 2025-11-25 — held to a revision it does not implement, so genuine non-conformance is the expected contentmcp-everything-server --protocol-version 2026-07-28, its stateless mode
Recorded bymcp-everything-server’s tap, during cargo xtask draft-readiness, 2026-08-20same run, second leg
Contents91 events / 22 POST exchanges91 events / 22 POST exchanges
Our verdict60 pass, 1 fail, 0 warn, 64 not observed, 147 excluded61 pass, 0 fail, 0 warn, 64 not observed, 147 excluded
The official runner’s verdict37 passing / 4 failing41 passing / 0 failing

Both carry server/discover, tools/list, tools/call, completion/complete, resources/{list,read}, prompts/{list,get} and progress notifications; every request carries the revision’s per-request _meta envelope and its SEP-2243 metadata headers, and neither carries an initialize anywhere. Their tools/list results differ by one entry, which is the point of recording both: test_url_elicitation is listed by the legacy server and not by the stateless one, because URL-mode elicitation is a feature 2026-07-28 removed.

The one finding. CACH-001 — the legacy server’s complete results carry no ttlMs, on resources/list and three resources/reads. That is our 2025-11-25 server being held to a revision it does not implement, which is the correct answer, and the stateless capture is the control that proves the check reports the server rather than the recording.

Two findings that used to be here were ours, not the implementations’. An earlier version of this pair reported TRAN-058 (client POSTs missing Mcp-Method/Mcp-Name) and TRAN-068 (SSE responses missing X-Accel-Buffering: no). Both were artifacts of the tap’s recording allowlist, which predated SEP-2243 and did not record those headers — the client sent them (rmcp’s transport rejects a 2026-07-28 request without Mcp-Method before dispatch, and the suite scored 23/23) and the server sent them (rmcp sets X-Accel-Buffering on every SSE response it builds). The allowlist was extended; both clauses now pass on both captures. The lesson is recorded rather than quietly fixed: a check can only be as honest as the recording it reads, and a capture path that silently drops evidence manufactures findings against conforming implementations.

The same lesson, a third time, one field over (2026-08-20). The tap recorded an exchange’s headers but not its method, so TRAN-025 and TRAN-039 — the Accept obligations of a POST and of a GET — were enforced as one rule applied to every client request whatever its form. That was wrong twice over, and both halves were found by pointing this workspace’s own reference host at its own reference server over HTTP and validating the tap: the conforming session-teardown DELETE, which owes no Accept header at all, drew two MUST-level failures, while a POST omitting application/json — the verbatim violation TRAN-025 names — was reported pass, because the one shared rule could only demand what both clauses had in common. EventBody::Http now carries the method, each clause is judged against the requests it binds, and an exchange whose method the capture did not record is judged by neither. These two captures were re-recorded so the method is present; that is why their Recorded by row reads 2026-08-20.

Why the runner’s 23/23 and our 58-vs-59 were both right — and how the runner caught up. The suite’s 2026-07-28 scenarios exercise features — list a thing, call a thing, read a thing — and a 2025-11-25 server answers all of them, because rmcp serves a per-request-versioned POST whichever revision the handler advertises. At 0.2.0-alpha.9 that was the whole of the runner’s reading, so it scored both servers 23/23. The registry here judges the specification’s prose instead, so it saw the one place the two servers actually differ: CACH-001, no ttlMs on cacheable results.

0.2.0-alpha.11 found the same thing, six weeks later. Its new wire-schema-valid check validates every message against the negotiated revision’s JSON schema, and it fails the 2025-11-25 leg’s four resource scenarios for must have required property 'cacheScope' / 'ttlMs' — that clause, independently. Re-measured, the legacy leg scores 37 passing / 4 failing and the stateless leg 41 / 0, so the runner now separates the two servers it could not tell apart before. Neither instrument was wrong; they were measuring different things, and one of them saw this first. The pair is still the evidence for taking both readings — now with a worked example of the prose-level reading arriving earlier than the schema-level one.

The 65 not-observed rows are the honest denominator: of the 124 clauses this revision’s registry can judge, these sessions carried subject matter for 59. They open no subscription, present no cursor, draw no error, and send no malformed _meta, so those clauses are neither passed nor failed here — they are untested, and the report says which.

Regenerate both with cargo xtask draft-readiness and copy target/draft-readiness/<revision>/tap/001-stateless.jsonl; the ephemeral port in the host header changes per run, so re-copying is a deliberate act with a golden diff, not something to do casually.

The stdio capture

Fieldreference-host-2026-07-28-stdio.jsonl
ClientThis workspace’s mcp-reference-host, on rmcp’s stateless client lifecycle
Servermcp-everything-server --transport stdio --protocol-version 2026-07-28
Recorded byThe host’s own --trace-dir capture, during cargo xtask draft-capture, 2026-08-18
Contentsserver/discover, a full subscriptions/listen lifecycle, a 16-tool sweep with four MRTR rounds (three elicitations and one sampling), and a discovery-driven sweep of everything that is not a tool: resources/{list,templates/list,read}, prompts/{list,get} for all four prompts, completion/complete, and one read of a URI the catalog does not contain. Every call carries a W3C traceparent in its _meta (BASE-040), and the session closes with a notifications/cancelled naming a request the server had already answered, then one more call the server may answer — the only shape a recording can take for a MUST NOT (TRAN-123/TRAN-124)
Our verdict81 pass, 0 fail, 0 warn, 44 not observed, 147 excluded

It is the only capture that exercises subscriptions/listen. The official suite drives no subscription, so the four judged SUBS clauses — and BASE-039, which binds a subscription stream’s notifications — are backed by this file alone: an acknowledgment, a filter the server narrowed, four tagged notifications and an empty graceful-closure result. The same holds for the MRTR clauses. In the two HTTP captures every one of those rows now reads not observed; here they read pass.

Two defects were found by enriching it, and neither could have been found any other way. Until the session read a resource the catalog does not have, nothing in this workspace had ever exercised the server’s not-found path at this revision — and it answered -32002, which 2026-07-28 withdraws outright (basic/index#error-codes lists it under “Implementations of this protocol version MUST NOT emit these codes” and names -32602 as the replacement). The second was ours: the host’s recorder assigned seq when the send future completed rather than when the message was handed to the transport, so a reply received while that future was still pending landed in the trace ahead of its own request. That reads as the server answering an id nobody asked, and it failed BASE-046 on roughly one capture in three. Both are fixed, both have regression tests, and both are the argument for driving more traffic rather than more fixtures.

That asymmetry is the reason to keep all three, and it used to be invisible. Until the vacuous-pass fix, those rows read identically across the three files — pass everywhere, whether or not the session had opened a stream — and this paragraph recorded that as a known reporting limitation. It is a limitation no longer: every check counts the subjects it examined, so a clause with nothing to look at reports not-observed. Read in the other direction, the captures are complementary rather than redundant: the HTTP pair evidences the transport header and status clauses, which no stdio recording can carry, and this one evidences the subscription, MRTR, prompts, resources, logging, completion and error-code clauses, which theirs never reach.

What the 47 not-observed rows still are, and why. Twenty-three are Streamable HTTP clauses in the TRAN-057TRAN-102 band that a stdio recording structurally cannot carry — the band holds 25 judged clauses, and the two a stdio session does reach (TRAN-060, TRAN-066) are judged here. Nine are server rejection rules — BASE-031, BASE-032, BASE-035, BASE-036, VERS-001, VERS-002, VERS-008, LOG-010, PAGE-011 — reachable only by a client that deliberately sends something malformed, which this session does not: it is the conforming capture, and a probe session is a separate recording with a separate expected report. Six need surface this server does not have (CACH-015/CACH-016 and PAGE-010 need a catalog large enough to paginate; TOOL-033/TOOL-034 need an x-mcp-header designation; PROM-017 needs a prompt carrying audio). TOOL-022 is the interesting one: rmcp’s client caches tools/list under the server’s own ttlMs, so a second listing never reaches the wire — a conforming client cannot exercise the deterministic-order clause within the TTL, which is a property of the caching feature working rather than a gap to close (the official suite’s runner does not cache, and judges it on both of its captures). The remaining eight are reachable and not yet driven: TRAN-123/TRAN-124 cancellation, TRAN-128 and DISC-002’s dual-era probe, MRTR-024, BASE-040, BASE-047, and VERS-004.

The HTTP capture of the same session

Fieldreference-host-2026-07-28-http.jsonl
ClientThe same mcp-reference-host, driving the identical session over Streamable HTTP
Servermcp-everything-server --transport http --protocol-version 2026-07-28
Recorded byThe server’s tap, during cargo xtask draft-capture, 2026-08-18
Contents159 events — the stdio session’s 85 messages plus 74 http events carrying status and headers
Our verdict92 pass, 0 fail, 0 warn, 33 not observed, 147 excluded

Recorded by the server, not the host, and that is the whole point. The host’s recorder sits at rmcp’s Transport seam, which carries protocol messages and nothing else — redaction by construction, and no HTTP framing to record even if it wanted to. So a host-side HTTP recording would report the Streamable HTTP clauses as not observed exactly like the stdio one. The server’s tap sits above the transport and sees the status line and the conformance-relevant headers, which is why this leg is the only recording in the corpus that can bear on them at all. Same session, both ends, one file each: the difference between the two reports is attributable to the transport and to nothing else.

At 92 of the 125 judgeable clauses it is the best-covered capture here. Its 33 not-observed rows are the server-rejection rules a conforming client never triggers, the pagination and x-mcp-header clauses this server’s surface does not reach, and TOOL-022 (rmcp’s client caches tools/list under the server’s own ttlMs, so a second listing never reaches the wire).

The probe session

Fieldprobe-2026-07-28-http.jsonl
Clientmcp-reference-host --probe: nine hand-built HTTP requests, each wrong on purpose
Servermcp-everything-server --transport http --protocol-version 2026-07-28
Recorded byThe server’s tap, during cargo xtask draft-capture, 2026-08-18
ContentsA _meta envelope missing a required field; an unimplemented protocol version and the retry after it; a header/body version mismatch; an unknown method; a log level outside RFC 5424’s eight; a fabricated cursor; a tool needing a capability the request never declared; and the removed initialize handshake
Our verdictJudged against conformance/probe-baseline.json, not for cleanliness

A conforming client cannot exercise a rejection rule. Fifteen clauses of this revision say what a server owes a request it must not serve, and every one of them reported not observed on every recording here, because nothing had ever sent such a request. This file is that request, nine times over.

The probes are built outside rmcp, and that is the point rather than a shortcut: rmcp’s client is what makes the other captures trustworthy — it builds the _meta envelope, mirrors the SEP-2243 headers, and will not emit an ill-formed request — so it is structurally incapable of being the probe. The bytes here are the fixture.

Its verdict is a ledger, not a pass. The probe breaks client-side clauses by construction: BASE-030 because its first request omits a required _meta field, TRAN-071/TRAN-072 because two probes are about exactly those headers, PAGE-010 because a fabricated cursor is a fabricated cursor. Demanding a clean report would mean demanding a probe that probes nothing. So every finding is listed in conformance/probe-baseline.json with a reason, and the gate holds the set in both directions: a finding not in the ledger is a new defect or a regression, and a listed finding that stopped occurring is either a fix that should retire its entry or a check that quietly stopped firing.

It has already worked in both directions. The first probe run drew two server-side findings — the server served a request naming log level "chatty", and honoured a cursor it never issued — which went into the ledger as open defects. When they were fixed, the gate refused the change until their entries were retired, which is the half of a ratchet that is easy to leave out. All ten rejection clauses the probe exercises now pass.

Its provenance is weaker than the pair above, and deliberately labelled so. Both ends of this session are ours: the official suite drives servers over --url only, so there is no third-party client for stdio to be recorded against. What it still supplies that an authored fixture cannot is that neither end was written to satisfy these checks, and that the machinery producing every byte — the stateless lifecycle, the _meta envelope, the MRTR retry loop — is rmcp’s, not this repository’s.

It is also the only capture that judges a verdict rather than pinning a report. cargo xtask draft-capture fails on any finding, because a finding in a session where both ends are ours is a defect here rather than news about somebody else’s code. Regenerate with BLESS=1 cargo xtask draft-capture (which rewrites this file) followed by cargo xtask bless.

What it caught. The first recording of this session reported six failures the HTTP captures could not: TRAN-060/066/119/120 and MRTR-001, because the interactive tools still sent elicitation/create and sampling/createMessage as independent server-to-client requests — the mechanism SEP-2322 replaced — and LOG-008, because the logging tool emitted notifications/message for a request that had not asked for them. The official suite’s 2026-07-28 scenarios exercise no interactive tool and no logging tool, so an HTTP-only corpus would have shipped both defects. That is the argument for capturing both transports rather than treating one as representative.

What the captures evidence, together

The table below is generated from the committed golden reports by cargo xtask draft-coverage and verified by cargo xtask ci — per ADR-0001 these numbers are a projection of the data, not prose anyone maintains. The same gate parses every count stated elsewhere in the living Markdown — the “N of the M judgeable clauses” claims and the “N pass, M fail” verdicts alike — and fails when one disagrees, which is how the counts in this section stopped being wrong. A number quoted rather than asserted goes in backticks, and is read as a specimen instead.

CaptureJudgedpassfailwarnNot observed
official-suite-2026-07-28-scenarios61601064
official-suite-2026-07-28-stateless61610064
probe-2026-07-28-http685610257
reference-host-2026-07-28-http92920033
reference-host-2026-07-28-stdio81810044
Union11411

Across all 5 captures, 114 of the 125 judgeable clauses are evidenced by at least one recording. Each capture’s own judged count is what that recording carried subject matter for; everything else it reports not observed rather than counting it as a pass.

The 11 clauses no capture reaches: CACH-015, CACH-016, MRTR-024, PROM-017, TOOL-033, TOOL-034, TRAN-070, TRAN-079, TRAN-080, TRAN-096, VERS-004.

The probe session closed the largest group — the rejection rules — and the cancellation round closed the next. What is left divides in three, and only the first is ordinary backlog.

Eight need server surface this reference does not have: pagination for CACH-015/CACH-016, an x-mcp-header designation for TOOL-033/TOOL-034 and TRAN-079/TRAN-080/TRAN-096, and a prompt carrying audio for PROM-017.

Two cannot be reached by a conforming host at all, and saying so is more useful than closing them. VERS-004 judges the shape of an extension identifier, so a capture would have to declare an extension — and this host implements none; declaring one to light up a check is precisely the vacuous evidence the corpus exists to avoid. MRTR-024 is a server re-asking after a client’s shortfall, and a retry that omits requested input is the client violation MRTR-015 names: driving it would make the reference host non-conforming, which is the one thing this capture may not be. Both are evidenced by authored violation traces instead, where the fault belongs.

One is a recording gap rather than a session gap: TRAN-070 needs a response stream to close mid-session, and the server’s tap records no lifecycle events at all, so no HTTP capture can carry one however the session is driven.

BASE-047 used to be on this list. It left without a trace being written, because it was never a corpus gap: the check behind it counted only malformed responses as subjects, so every capture full of well-formed results reported the clause not observed — the tool saying it had seen nothing to judge about a session that had carried dozens of them.

2026-07-28 authored (corpus/draft/)

Fixtures for the in-progress revision, judged by the 2026-07-28 registry behind the draft-2026-07-28 feature. Hand-authored rather than captured: no SDK yet serves the stateless surface end to end, so these are minimal traces written to exercise one clause each, which is also what keeps a violation attributable to the check it kills.

Where a trace falsifies more than one requirement, the row says so and why. There are only two reasons, and neither is sloppy authorship: several clauses state one rule across several sections and share a check by design, or one clause’s antecedent is itself a violation by the other party — a server can only fail to reject a bad header if the client sent one. Every other row falsifies exactly the requirement it is named for.

TraceExercises
stateless-session.jsonlConformant 2026-07-28 stateless session over stdio: every request carries its own _meta envelope (protocolVersion + clientCapabilities), every result carries resultType, request ids are reused only after their response. It closes with two calls in flight, a notifications/cancelled for one of them, and the server answering only the other — TRAN-124’s pass path, which a recording can only carry by showing a permitted message where the forbidden one would be. It also carries a complete conforming MRTR round — input_required with an elicitation/create request and a requestState, then a retry under a new id echoing the state and supplying inputResponses — so the MRTR checks pass on real content. The Streamable HTTP clauses report not observed here: there is no HTTP framing for them to read, which is why streamable-http-session.jsonl exists.
base-030-request-meta-missing-required-field.jsonlRequest _meta omits io.modelcontextprotocol/clientCapabilities (BASE-030)
base-031-malformed-meta-answered-with-result.jsonlServer answers a _meta-incomplete request with a result instead of -32602 (BASE-031)
base-032-invalid-params-not-http-400.jsonl-32602 returned with HTTP 200 rather than 400 (BASE-032)
base-034-input-request-for-undeclared-capability.jsonlServer returns input_required asking for elicitation/create the request never declared (BASE-034)
base-035-missing-capability-error-without-capabilities.jsonl-32021 carries no data.requiredCapabilities (BASE-035)
base-036-missing-capability-not-http-400.jsonl-32021 returned with HTTP 500 rather than 400 (BASE-036)
base-039-subscription-notification-untagged.jsonlNotification on a subscriptions/listen stream with no io.modelcontextprotocol/subscriptionId (BASE-039)
base-040-malformed-traceparent.jsonltraceparent that is not W3C Trace Context shaped (BASE-040)
base-045-request-id-reused-in-flight.jsonlRequest id reused while the first is still outstanding (BASE-045) — legal at 2025-11-25 only after a response, and this trace reuses before one
base-048-result-without-result-type.jsonlResult omits the resultType SEP-2322 requires (BASE-048)
base-055-legacy-error-code.jsonlError code -32010 from the closed legacy sub-range (BASE-055)
base-057-undefined-reserved-error-code.jsonlError code -32055: inside the MCP-reserved sub-range but undefined (BASE-057)
base-058-withdrawn-error-code.jsonlError code -32002, withdrawn by this revision (BASE-058)
base-060-app-code-in-reserved-range.jsonlApplication-defined -32500 placed inside the JSON-RPC reserved range (BASE-060)
streamable-http-session.jsonlConformant 2026-07-28 Streamable HTTP session: server/discover under client capabilities advertising a correctly prefixed extension identifier (VERS-004), a paginated tools/list whose two pages agree on cacheScope (CACH-015/016) and which declares x-mcp-header annotations on a string and an integer property, a tools/call mirroring both into Mcp-Param-Region/Mcp-Param-Limit — the integer inside the IEEE 754 safe range (TOOL-034) — over an SSE response with X-Accel-Buffering: no, and a resources/read whose non-ASCII Mcp-Name rides the Base64 sentinel. The continuation page is fetched after a transport-close, which is ordinary here (every POST gets its own response stream) and is the only way a recording can show a server declining to send more for the request that close cancelled — TRAN-070’s pass path.
tran-058-request-metadata-headers-missing.jsonlPOST carries neither Mcp-Method nor Mcp-Name (TRAN-058)
tran-060-client-posts-a-response.jsonlClient POSTs a JSON-RPC response (TRAN-060). Also falsifies BASE-046: at this revision a server cannot issue the request such a response would answer, so an unsolicited id is the only shape the violation can take.
tran-066-independent-server-request.jsonlServer sends elicitation/create as its own request on the response stream instead of an MRTR input request (TRAN-066)
tran-068-sse-without-accel-buffering.jsonlSSE response omits X-Accel-Buffering: no (TRAN-068, SHOULD → warn)
tran-070-message-after-stream-close.jsonlServer answers a request whose response stream had already closed, which this revision treats as cancellation (TRAN-070)
tran-071-protocol-version-header-missing.jsonlPOST request without MCP-Protocol-Version (TRAN-071)
tran-072-protocol-version-header-mismatched.jsonlHeader says 2025-11-25, body _meta says 2026-07-28; the server rejects it correctly, isolating the client’s fault (TRAN-072)
tran-073-header-mismatch-not-rejected.jsonlThe same disagreement, answered with a result (TRAN-073). Necessarily also falsifies TRAN-072 — the client fault is this clause’s antecedent.
tran-074-unsupported-version-without-supported-list.jsonl-32022 without the data.supported list the clause requires (TRAN-074)
tran-074-unsupported-version-accepted.jsonlA request naming a version the server’s own server/discover result omits, answered with a result (TRAN-074)
tran-075-method-not-found-not-404.jsonl-32601 returned with HTTP 200 rather than 404 (TRAN-075)
tran-077-mcp-name-unencoded.jsonlNon-ASCII Mcp-Name carried unencoded (TRAN-077, TRAN-086, TRAN-087 — one rule stated in three sections, one shared transport.header-value-encoding check)
tran-079-x-mcp-header-not-mirrored.jsonltools/call supplies a designated argument without its Mcp-Param-* header (TRAN-079)
tran-080-x-mcp-header-name-invalid.jsonlOne inputSchema with three annotation faults: a non-token name, a case-insensitive duplicate, and an annotation on a number property (TRAN-080)
tran-089-sentinel-markers-miscased.jsonl=?BASE64?…?= — the sentinel markers must be exactly lowercase (TRAN-089). The server rejects it correctly, so nothing else fires.
tran-092-sentinel-pattern-unencoded.jsonlA plain value that matches the sentinel pattern, carried verbatim rather than encoded (TRAN-092)
tran-096-invalid-param-header-accepted.jsonlA recognized Mcp-Param-Region whose value has leading whitespace, answered with a result (TRAN-096). Also falsifies TRAN-077/086/087 — the unencodable value the server had to reject is itself the client’s encoding fault.
tran-077-param-header-value-rejected.jsonlThe same unencodable Mcp-Param-Region, rejected with -32020 and HTTP 400 as the rule requires (TRAN-077, TRAN-086, TRAN-087 — the client’s encoding fault, which is all that is left to fail). TRAN-096’s pass path: the value a server must reject is itself a client fault, so the conforming half of that clause cannot live in good/ and is carried here instead.
tran-097-header-body-mismatch-accepted.jsonlMcp-Param-Region: us-east1 against arguments.region = "us-west1", answered with a result (TRAN-097, TRAN-100 — one rule stated in two sections)
tran-098-header-mismatch-without-400.jsonlHeaderMismatch returned with HTTP 500 rather than 400 (TRAN-098, TRAN-102 — one rule stated in two sections)
tran-074-unsupported-version-without-400.jsonl-32022 returned with HTTP 200 rather than 400 (TRAN-074). The status half of that clause had no trace of its own until transport.unsupported-version-status was split out; it had been riding the kills of the sibling rules it was bundled with.
tool-019-tools-undeclared.jsonltools/call answered though discovery declared no tools capability (TOOL-019)
tool-020-declared-tools-list-unimplemented.jsonltools declared, but tools/list refused with -32601 (TOOL-020)
tool-022-tools-list-order-changes.jsonlTwo tools/list results with the same tools in a different order (TOOL-022)
tool-034-mirrored-integer-out-of-range.jsonlAn x-mcp-header-annotated argument of 2^53, outside the IEEE 754 safe range (TOOL-034)
tool-038-embedded-resource-undeclared.jsonlA tools/call result embedding a resource with no resources capability declared (TOOL-038)
res-012-resources-undeclared.jsonlresources/read answered though discovery declared no resources capability (RES-012)
res-013-declared-resources-list-unimplemented.jsonlresources declared, but resources/list refused with -32601 (RES-013)
res-022-read-empty-contents.jsonlresources/read answered with an empty contents array (RES-022)
prom-012-prompts-undeclared.jsonlprompts/get answered though discovery declared no prompts capability (PROM-012). The result it wrongly serves is itself well formed — base64 audio with a valid MIME type — which is PROM-017’s pass path
prom-013-declared-prompts-list-unimplemented.jsonlprompts declared, but prompts/list refused with -32601 (PROM-013)
comp-007-completions-undeclared.jsonlcompletion/complete answered though the server/discover result declared no completions capability (COMP-007)
log-007-logging-undeclared.jsonlnotifications/message emitted though discovery declared no logging capability (LOG-007)
log-008-log-without-requested-level.jsonlA log notification in a session where no request set io.modelcontextprotocol/logLevel (LOG-008)
log-009-log-on-subscription-stream.jsonlA log notification tagged with a subscription id, so travelling on a subscription’s stream (LOG-009)
log-010-unrecognized-log-level-accepted.jsonlA request declaring log level verbose, served rather than rejected with -32602 (LOG-010)
page-011-unissued-cursor-accepted.jsonlA tools/list presenting a cursor the session never issued, answered with a result (PAGE-011). Also falsifies PAGE-002 at its own revision’s registry, and PAGE-010 here — the fabricated cursor is the client’s defect and this clause’s antecedent.
cach-001-cacheable-result-without-hints.jsonlA complete tools/list result with no ttlMs caching hint (CACH-001)
cach-008-negative-ttl.jsonlttlMs: -1, which servers must never provide (CACH-008)
cach-015-page-scope-changes.jsonlA paginated tools/list whose second page switches from private to public (CACH-015, and CACH-016 — the same rule and its worked example)
subs-001-unrequested-notification-type.jsonlA prompts/list_changed on a subscription whose filter asked only for tools-list changes (SUBS-001)
subs-002-notification-before-acknowledgment.jsonlA notification ahead of notifications/subscriptions/acknowledged on the same subscription id (SUBS-002)
subs-006-graceful-close-result-not-empty.jsonlA graceful-closure response carrying delivered alongside resultType (SUBS-006, and SUBS-005 — one rule stated in the cancellation list and again under Graceful Closure)
mrtr-004-input-required-on-unsupported-method.jsonlinput_required answering a resources/list, which is not one of the three requests that may draw one (MRTR-004)
mrtr-006-input-request-unknown-method.jsonlinputRequests asking for tools/list, which is not an ElicitRequest, CreateMessageRequest or ListRootsRequest (MRTR-006)
mrtr-011-input-required-empty.jsonlinput_required carrying neither inputRequests nor requestState, opening a round that cannot be completed (MRTR-011)
mrtr-015-retry-without-input-responses.jsonlRetry echoes the state but supplies no inputResponses for the input it was asked for (MRTR-015). The server then asks again rather than erroring — MRTR-024’s pass path, which needs a shortfall to answer and so cannot occur in a conforming session
mrtr-016-request-state-not-echoed.jsonlRetry rewrites requestState instead of echoing it (MRTR-016). Also falsifies MRTR-003 and MRTR-017 by design — the same rule stated from the other side, “MUST NOT modify”, sharing one check.
mrtr-018-unsolicited-request-state.jsonlRetry invents a requestState for a round that issued none (MRTR-018)
mrtr-019-retry-reuses-id.jsonlRetry reuses the original request’s JSON-RPC id (MRTR-019) — legal under BASE-045, since the first was already answered, and forbidden here
mrtr-020-request-state-on-another-method.jsonlA prompts/get carrying the requestState a tools/call round issued (MRTR-020)
mrtr-024-shortfall-answered-with-error.jsonlRetry omits requested input and the server answers -32602 rather than asking again (MRTR-024). Necessarily also falsifies MRTR-015 — the client’s shortfall is this clause’s antecedent.
tran-123-cancellation-without-request-id.jsonlnotifications/cancelled carrying only a reason, naming no request to cancel (TRAN-123)
tran-124-message-after-cancel-notification.jsonlServer answers a request the client cancelled by notification (TRAN-124). stdio’s cancellation signal is the notification, not a stream close, so transport.no-messages-after-cancellation — which anchors on a transport-close — cannot see this and would have passed it vacuously.
disc-001-server-discover-method-not-found.jsonlserver/discover answered with -32601, which this revision makes mandatory (DISC-001, and VERS-003 — the same rule stated under versioning). The client then falls back to initialize, carrying a full _meta envelope so the trace isolates the fallback: a dual-era client that probed first is DISC-002/TRAN-128’s pass path, and it can only be witnessed against a server that refused the probe
disc-002-dual-era-client-skips-probe.jsonlA client that speaks both eras — a modern _meta request, then an initialize fallback — whose first request is tools/list rather than the server/discover probe (DISC-002). The initialize carries a full _meta envelope so the trace isolates the missing probe rather than also failing BASE-030, and the handshake is left unanswered so no legacy-shaped result has to be judged for resultType.
vers-002-retry-with-unsupported-version.jsonlAfter a -32022 offering 2026-07-28, the client retries with 1899-01-01 — a version the list it was just handed does not contain (VERS-002)
vers-004-extension-identifier-without-prefix.jsonlClient capabilities advertise the extension ui: a valid _meta key, but extension identifiers require the prefix a _meta key may omit (VERS-004)
vers-008-initialize-error-without-versions.jsonlA modern-only server refuses a legacy initialize with a bare Method not found, naming none of the versions it does speak (VERS-008). Necessarily also falsifies BASE-030: a legacy handshake carries no _meta envelope — that is what makes it legacy — so the antecedent of this clause is a client that cannot satisfy BASE-030. BASE-031 is not also reported, and this trace is why the check was narrowed: basic/versioning’s compatibility matrix makes the rejection code implementation-defined for exactly this exchange.

Conformance results

A conformance toolkit should be willing to measure real implementations — its own and others’ — against the official suite, and publish what it finds. These measurements are stewardship artifacts: every number is the official suite’s number, captured by running it, with this project’s registry used only to read each failure back to a specific spec clause. Every claim has a command that reproduces it.

This project’s reference implementations

The reference server and host are a hard CI gate, not a one-off measurement: on every run, the conformance job drives the pinned official suite against mcp-everything-server (server scenarios) and mcp-reference-host (client scenarios), and replays the captured sessions through the validator for the agreement check. A regression fails the build. The live status is whatever the latest main run reports.

rmcp tier-gap report

The published measurement of the official Rust SDK (rmcp) against the suite is the worked example of the method:

📄 rmcp tier-gap report — 2025-11-25

It records the server-scenario pass count, the requirement-level reading of each failure, a concrete close-the-gap checklist, and the exact commands to reproduce it — including a note on why the suite’s tier-check aggregate is not the figure to trust (it is GitHub-token-gated and its conformance counter carries a known upstream counting bug). The headline finding at the report’s measurement date: two failing server scenarios, neither of which is a 2025-11-25 normative-clause violation — one sits just below the spec’s RFC 2119 floor (template-argument substitution is schema prose), the other is a SEP-1330 serialization bug already filed upstream. That is the kind of distinction a requirement-level tool exists to draw; see the report for the current count and the close-the-gap steps.