Tools
MCP's 2026-07-28 spec removes sessions, so servers can run behind plain load balancers
The new Model Context Protocol revision drops the initialize handshake and Mcp-Session-Id. What changes for load balancers, gateways and serverless, plus a migration checklist.
HackHoster Team · · 10 min read

At a glance
- MCP revision 2026-07-28, published July 28, removes the initialize handshake and the Mcp-Session-Id header, making every request self-contained.
- Streamable HTTP requests must now carry Mcp-Method and, for tool calls and reads, Mcp-Name headers, which servers must check against the JSON body.
- Multi Round-Trip Requests replace server-initiated sampling, elicitation and roots calls; the server returns input_required and the client retries.
- Roots, Sampling, Logging and the legacy HTTP+SSE transport are deprecated under a new policy with a minimum twelve-month window.
- The TypeScript, Python, Go and C# SDKs support the revision at launch, Rust in beta; Tier 1 SDKs see close to half a billion downloads a month.
The Model Context Protocol published specification revision 2026-07-28 on Tuesday, and the headline change is that the protocol no longer has sessions. The initialize/initialized handshake is gone, and so is the Mcp-Session-Id header on the Streamable HTTP transport. Every request now carries its own protocol version, client capabilities and client identity. Lead maintainers David Soria Parra and Den Delimarsky wrote that any request can now land on any server instance behind a plain round-robin load balancer, with no shared storage.
The announcement calls this MCP's most important release since remote MCP first launched. All four Tier 1 SDKs, TypeScript, Python, Go and C#, support the new revision at release, and the Rust SDK supports it in beta. According to the maintainers, Tier 1 SDKs see close to half a billion downloads a month, and the TypeScript and Python SDKs have each passed a billion downloads in total. A change this deep will ripple through a lot of code.
This article covers where MCP came from, why sessions became a problem, what a request looks like now, the new pattern for servers that need input from the client, the authorization changes, and a migration checklist.
From a desktop plug-in protocol to remote infrastructure
Anthropic introduced MCP on November 25, 2024, as an open standard for connecting AI applications to data sources and tools. The pitch was to replace a custom connector for every pairing of model app and data source with one protocol. The first release shipped a specification, SDKs, local server support in Claude Desktop, and reference servers for Google Drive, Slack, GitHub, Git, Postgres and Puppeteer.
The early design reflected that local starting point. A client and server kept a connection open, shook hands once with initialize, and from then on either side could send requests to the other. That bidirectional, stateful model fits a desktop app talking to a local process. It fits less well once servers are remote, multi-tenant services sitting behind gateways and autoscalers.
| Date | Milestone |
|---|---|
| November 2024 | MCP announced; revision 2024-11-05 uses an HTTP+SSE transport for remote servers |
| March 2025 | Revision 2025-03-26 introduces Streamable HTTP and deprecates HTTP+SSE |
| June 2025 | Revision 2025-06-18 adds the MCP-Protocol-Version header |
| November 2025 | Revision 2025-11-25, the previous version, carries experimental tasks in the core protocol |
| December 2025 | Anthropic donates MCP to the Agentic AI Foundation under the Linux Foundation; 97 million monthly SDK downloads |
| July 2026 | Revision 2026-07-28 removes sessions; close to half a billion monthly Tier 1 SDK downloads |
Sources: MCP specification pages and announcements listed below.
Governance changed along the way. In December 2025, Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded by Anthropic, Block and OpenAI. The maintainers said at the time that the existing maintainer structure and decision process would stay as they were. They also reported 97 million monthly SDK downloads and 10,000 active servers. Seven months later, the monthly figure for Tier 1 SDKs alone is roughly five times higher.
Why sessions became a problem
Under the 2025 revisions of Streamable HTTP, a server could assign a session through the Mcp-Session-Id header, which the client sent back on every request and ended with an HTTP DELETE. Clients could also open a standalone GET stream to receive server-initiated messages, and servers could send their own JSON-RPC requests down those streams.
As Appwrite's write-up explains, the trouble was that the state lived on one server instance. Every follow-up request from a client had to reach the instance that handled its handshake. Operators ended up with sticky sessions at the load balancer, or shared storage so any instance could recover the session. Both add infrastructure without adding features, and both made scaling remote servers more complicated than it needed to be.

The web settled this question a long time ago. Roy Fielding's 2000 doctoral dissertation, which defined the REST architectural style, made statelessness one of its constraints: each request carries the context needed to process it, and the server does not hold client session state between requests. Wikipedia's summary of REST lists scalability, visibility of communication to intermediaries and reliability among the benefits of its constraints. The 2026-07-28 revision moves MCP toward that model.

Stateless, in this spec: the protocol keeps no per-client memory between requests. Each request carries its protocol version, capabilities and identity in
_meta. Your application can still have state, but it has to live somewhere explicit, such as a handle passed as a tool argument or your own database.
What a request looks like now
On Streamable HTTP, every message is its own POST to a single MCP endpoint. A tool call from the spec's transport page looks like this:
POST /mcp HTTP/1.1
Content-Type: application/json
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: get_weather
{"jsonrpc": "2.0", "id": 1, "method": "tools/call",
"params": {"name": "get_weather", "arguments": {"location": "Seattle, WA"},
"_meta": {"io.modelcontextprotocol/protocolVersion": "2026-07-28",
"io.modelcontextprotocol/clientInfo": {"name": "ExampleClient", "version": "1.0.0"},
"io.modelcontextprotocol/clientCapabilities": {}}}}
The server answers with either a single JSON object or a Server-Sent Events stream scoped to that one request, carrying progress notifications and then the final result.

Headers that intermediaries can trust
Three details matter for operators. First, the Mcp-Method header is required on every request, and Mcp-Name on tools/call, resources/read and prompts/get, so a gateway, rate limiter or firewall can route and meter without parsing JSON. Names that aren't plain ASCII travel in a Base64 sentinel format, =?base64?…?=.
Second, servers must reject any request whose headers disagree with the body, returning 400 with a HeaderMismatch error, code -32020. The spec explains why: otherwise a load balancer could route on one value while the server executes another. It also advises intermediaries that enforce policy on these headers to reject requests claiming an older protocol version, since those carry no such guarantee.
Third, a tool can mark parameters with x-mcp-header so clients mirror them into Mcp-Param-{Name} headers. The rules are strict: only string, integer and boolean parameters qualify, the property must be reachable through plain properties keys, and clients must drop any tool whose annotations break the rules. That lets you route on, say, a region argument.

Servers must also implement a new server/discover method that advertises supported versions, capabilities and identity. Clients may call it first, but don't have to. A version the server doesn't support gets an UnsupportedProtocolVersionError listing the ones it does.
How servers ask the client for input now
The trickiest removal is server-initiated requests. Sampling, elicitation and roots used to be requests the server sent to the client over an open stream. With Multi Round-Trip Requests (MRTR), the server instead returns a result with resultType: "input_required":
{"jsonrpc": "2.0", "id": 1,
"result": {"resultType": "input_required",
"inputRequests": {"github_login": {"method": "elicitation/create",
"params": {"mode": "form", "message": "Please provide your GitHub username",
"requestedSchema": {"type": "object",
"properties": {"name": {"type": "string"}}, "required": ["name"]}}}},
"requestState": "AEAD-protected blob"}}
The client gathers the answers and retries the original call with a new JSON-RPC ID, the matching inputResponses, and the exact requestState it received. Whichever instance gets the retry has everything it needs. Only tools/call, resources/read and prompts/get may return this kind of result.
The requestState field is where developers will need to be careful. The spec says servers must treat it as attacker-controlled input. If it affects authorization, resource access or business logic, it must be integrity-protected with something like an HMAC or AEAD, and state that fails verification must be rejected. To limit replay, servers should bind it to the authenticated user, give it a short expiry, and tie it to the method and a digest of the original parameters. The spec also warns that none of this guarantees single use; one-time actions still need a server-side check.
Streams, tasks and caching
| Area | Revisions 2025-03-26 to 2025-11-25 | Revision 2026-07-28 |
|---|---|---|
| Setup | initialize handshake, optional Mcp-Session-Id | No handshake; metadata on every request; optional server/discover |
| Server-to-client requests | Sent over SSE streams | Returned as input_required results (MRTR) |
| Change notifications | Standalone GET stream | One subscriptions/listen POST stream |
| Resuming dropped streams | Last-Event-ID redelivery | Removed; client re-sends with a new ID |
| Long-running work | Experimental tasks in core | io.modelcontextprotocol/tasks extension, polled with tasks/get |
| List caching | listChanged notifications | Required ttlMs and cacheScope hints as well |
| Logging | logging/setLevel | Per-request log level in _meta; Logging deprecated |
Source: the 2026-07-28 changelog and transport specification.
A few of these deserve a closer look:
- One long-lived stream remains. Change notifications for tool, prompt and resource lists now arrive on a single
subscriptions/listenPOST response stream. The spec encourages periodic SSE comment lines as keep-alives and tells servers to sendX-Accel-Buffering: noso proxies such as nginx don't hold events back. - State becomes explicit. The maintainers suggest that a server needing cross-call state mint a handle from a tool and let the model pass it back as an ordinary argument. List endpoints no longer vary per connection.
- Long work moves to tasks. The tasks extension replaces the blocking
tasks/resultwith polling viatasks/get, addstasks/updatefor client input, and lets servers return task handles without per-request opt-in. That suits platforms with request timeouts. - Caches get hints. Results from
tools/list,prompts/list,resources/list,resources/templates/listandresources/readmust includettlMsandcacheScope(publicorprivate). Servers should return tools in a deterministic order, which the changelog says improves client caching and LLM prompt-cache hit rates. - Tracing is standardized. The spec documents
traceparent,tracestateandbaggagekeys in_metafor OpenTelemetry context propagation.

Authorization tightens
The authorization changes are smaller but matter for anyone running a remote server with OAuth.
- Issuer validation. Authorization servers should include the
issparameter in authorization responses, per RFC 9207, and clients must validate it against the recorded issuer before redeeming the code. Appwrite notes this, together with credential binding, defends against server mix-up attacks. - Credentials bound to issuers. Clients must key stored credentials by issuer, must never reuse them with a different authorization server, and must re-register when the server changes.
- Registration. Clients must send an
application_typeduring Dynamic Client Registration to avoid OpenID Connect redirect conflicts. Dynamic Client Registration itself is now deprecated in favor of Client ID Metadata Documents, though it stays available for authorization servers that don't support them.

What breaks, and what is still open
The main cost is the transition. Older clients still send initialize. The compatibility rules put the burden on clients: try a modern request first, and fall back to initialize only if a 400 comes back without a recognized modern JSON-RPC error. A server that only speaks the new revision will turn away clients that haven't updated. If you serve desktop apps or tools you don't control, expect to run both behaviors for a while.
Several smaller changes can also break code:
ping,logging/setLevelandnotifications/roots/list_changedare removed outright.- The resource-not-found error code changes from
-32002to-32602. - Every result now carries a required
resultType; clients must treat results without it, from older servers, as complete. - The URL-mode elicitation completion notification and its
elicitationId, added only in 2025-11-25, are gone.
Roots, Sampling, Logging and the legacy HTTP+SSE transport are deprecated rather than removed. Under the new feature lifecycle policy, deprecated features keep working for at least twelve months. The changelog suggests replacements: pass directories through tool parameters instead of Roots, call LLM provider APIs directly instead of Sampling, and log to stderr or OpenTelemetry instead of Logging.
Some questions don't have answers yet. Moving state into handles and requestState gives developers more design responsibility, as Appwrite points out, and gets it wrong in new ways, such as unsigned state blobs. The subscriptions/listen stream is still long-lived, so platforms with hard response-duration caps will see clients reconnecting. And the ecosystem's many third-party servers will move at their own pace. The announcement includes a note from Manufact saying its SDK v2 is about 83 percent smaller and 25 percent faster thanks to a new client-server split, and FastMCP says version 4.0 ships support for the new features, but most servers depend on their maintainers upgrading.
Migration checklist for server operators
- Upgrade to an SDK release that supports 2026-07-28 and read its migration notes.
- Move anything you kept in session memory into explicit handles,
requestState, or your own datastore. - Replace server-initiated sampling, elicitation and roots calls with Multi Round-Trip Requests, and sign any
requestStatethat affects access or logic. - Remove session affinity from your load balancer once old clients are gone.
- Validate
MCP-Protocol-Version,Mcp-MethodandMcp-Nameagainst the body in the server. Gateways that route on these headers should reject requests that claim an older protocol version. - Add
ttlMsandcacheScopeto list results, and sort your tool list. - If you drop older revisions, answer GET and DELETE on the MCP endpoint with
405and ignore incomingMcp-Session-IdandLast-Event-IDheaders. - Swap the Logging feature for stderr or OpenTelemetry. Log level now travels per request in
_meta. - Validate
isson OAuth responses and plan a move from Dynamic Client Registration to Client ID Metadata Documents.
Practical tip for hackathon teams: a new MCP server built against 2026-07-28 can be a single stateless function. Skip sessions entirely, keep state in a handle the model passes back, and deploy behind any HTTP load balancer. Keep a fallback path for
initializeif your demo needs to work with clients that haven't updated.
The transport spec keeps its security basics: servers must validate the Origin header to block DNS rebinding, and should bind to 127.0.0.1 rather than all interfaces when running locally and authenticate every connection.
What to watch next
As of July 30, the work shifts from the spec to the ecosystem. Watch for the Rust SDK to leave beta, for major desktop and IDE clients to ship releases that send per-request metadata, and for hosted platforms and gateways to add routing on Mcp-Method and Mcp-Name. The deprecation clock for Roots, Sampling, Logging and HTTP+SSE has started; under the lifecycle policy, none of them can be removed before July 2027. How quickly authorization servers adopt Client ID Metadata Documents will decide how long Dynamic Client Registration lingers.
Sources
- The 2026-07-28 MCP specification (Model Context Protocol blog)
- Key changes in specification revision 2026-07-28 (modelcontextprotocol.io)
- Streamable HTTP transport, revision 2026-07-28 (modelcontextprotocol.io)
- Multi Round-Trip Requests, revision 2026-07-28 (modelcontextprotocol.io)
- What's new in the MCP 2026-07-28 specification (Appwrite)
- Introducing the Model Context Protocol (Anthropic, November 2024)
- MCP joins the Agentic AI Foundation (Model Context Protocol blog, December 2025)
- REST (Wikipedia)
More from the blog

Tools ·
Google I/O 2026 gives developers Gemini 3.5 Flash, Antigravity 2.0 and one-call managed agents
Google's developer keynote was about agents. A faster but pricier Flash model, a desktop agent app with a CLI and SDK, sandboxed agents from one API call, and a $2 million XPRIZE hackathon.
10 min read

Policy ·
Third Circuit upholds ruling that ROSS's AI training on Westlaw headnotes was not fair use
A federal appeals court affirmed that ROSS Intelligence infringed Thomson Reuters' copyrights by training a legal search tool on Westlaw headnotes. The opinion itself is still sealed.
10 min read

Security ·
An OpenAI research agent got past access blocks on an Australian Medicare statistics portal
Australia's prime minister says an OpenAI model researching medicine spending got past access controls on a government portal in June. OpenAI told the government in September.
10 min read