All posts

Tools

MCP's 2026-07-28 spec removes sessions, so servers can run behind plain load balancers

The new Model Context Protocol revision drops the initialize handshake and Mcp-Session-Id. What changes for load balancers, gateways and serverless, plus a migration checklist.

HackHoster Team · · 10 min read

A large server hall with rows of open rack cabinets under bright ceiling lights
Photo: Florian Hirzinger / Wikimedia Commons, CC BY-SA 3.0

At a glance

  • MCP revision 2026-07-28, published July 28, removes the initialize handshake and the Mcp-Session-Id header, making every request self-contained.
  • Streamable HTTP requests must now carry Mcp-Method and, for tool calls and reads, Mcp-Name headers, which servers must check against the JSON body.
  • Multi Round-Trip Requests replace server-initiated sampling, elicitation and roots calls; the server returns input_required and the client retries.
  • Roots, Sampling, Logging and the legacy HTTP+SSE transport are deprecated under a new policy with a minimum twelve-month window.
  • The TypeScript, Python, Go and C# SDKs support the revision at launch, Rust in beta; Tier 1 SDKs see close to half a billion downloads a month.

The Model Context Protocol published specification revision 2026-07-28 on Tuesday, and the headline change is that the protocol no longer has sessions. The initialize/initialized handshake is gone, and so is the Mcp-Session-Id header on the Streamable HTTP transport. Every request now carries its own protocol version, client capabilities and client identity. Lead maintainers David Soria Parra and Den Delimarsky wrote that any request can now land on any server instance behind a plain round-robin load balancer, with no shared storage.

The announcement calls this MCP's most important release since remote MCP first launched. All four Tier 1 SDKs, TypeScript, Python, Go and C#, support the new revision at release, and the Rust SDK supports it in beta. According to the maintainers, Tier 1 SDKs see close to half a billion downloads a month, and the TypeScript and Python SDKs have each passed a billion downloads in total. A change this deep will ripple through a lot of code.

This article covers where MCP came from, why sessions became a problem, what a request looks like now, the new pattern for servers that need input from the client, the authorization changes, and a migration checklist.

From a desktop plug-in protocol to remote infrastructure

Anthropic introduced MCP on November 25, 2024, as an open standard for connecting AI applications to data sources and tools. The pitch was to replace a custom connector for every pairing of model app and data source with one protocol. The first release shipped a specification, SDKs, local server support in Claude Desktop, and reference servers for Google Drive, Slack, GitHub, Git, Postgres and Puppeteer.

The early design reflected that local starting point. A client and server kept a connection open, shook hands once with initialize, and from then on either side could send requests to the other. That bidirectional, stateful model fits a desktop app talking to a local process. It fits less well once servers are remote, multi-tenant services sitting behind gateways and autoscalers.

DateMilestone
November 2024MCP announced; revision 2024-11-05 uses an HTTP+SSE transport for remote servers
March 2025Revision 2025-03-26 introduces Streamable HTTP and deprecates HTTP+SSE
June 2025Revision 2025-06-18 adds the MCP-Protocol-Version header
November 2025Revision 2025-11-25, the previous version, carries experimental tasks in the core protocol
December 2025Anthropic donates MCP to the Agentic AI Foundation under the Linux Foundation; 97 million monthly SDK downloads
July 2026Revision 2026-07-28 removes sessions; close to half a billion monthly Tier 1 SDK downloads

Sources: MCP specification pages and announcements listed below.

Governance changed along the way. In December 2025, Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded by Anthropic, Block and OpenAI. The maintainers said at the time that the existing maintainer structure and decision process would stay as they were. They also reported 97 million monthly SDK downloads and 10,000 active servers. Seven months later, the monthly figure for Tier 1 SDKs alone is roughly five times higher.

Why sessions became a problem

Under the 2025 revisions of Streamable HTTP, a server could assign a session through the Mcp-Session-Id header, which the client sent back on every request and ended with an HTTP DELETE. Clients could also open a standalone GET stream to receive server-initiated messages, and servers could send their own JSON-RPC requests down those streams.

As Appwrite's write-up explains, the trouble was that the state lived on one server instance. Every follow-up request from a client had to reach the instance that handled its handshake. Operators ended up with sticky sessions at the load balancer, or shared storage so any instance could recover the session. Both add infrastructure without adding features, and both made scaling remote servers more complicated than it needed to be.

Diagram of user traffic passing through a load balancer to three pools of search servers, with backups to object storage
A load balancer spreading requests across server pools, from Wikimedia's 2014 search cluster. Under the new revision an MCP server can sit behind the same kind of setup without sticky sessions. Diagram: ^demon / Wikimedia Commons, CC BY-SA 4.0

The web settled this question a long time ago. Roy Fielding's 2000 doctoral dissertation, which defined the REST architectural style, made statelessness one of its constraints: each request carries the context needed to process it, and the server does not hold client session state between requests. Wikipedia's summary of REST lists scalability, visibility of communication to intermediaries and reliability among the benefits of its constraints. The 2026-07-28 revision moves MCP toward that model.

A man in a short-sleeved shirt speaking at a podium marked OSCON on a dark stage
Roy Fielding speaking at OSCON in 2008. His 2000 dissertation defined REST, with stateless requests as a core constraint. Photo: Phil Whitehouse / Wikimedia Commons, CC BY 2.0

Stateless, in this spec: the protocol keeps no per-client memory between requests. Each request carries its protocol version, capabilities and identity in _meta. Your application can still have state, but it has to live somewhere explicit, such as a handle passed as a tool argument or your own database.

What a request looks like now

On Streamable HTTP, every message is its own POST to a single MCP endpoint. A tool call from the spec's transport page looks like this:

POST /mcp HTTP/1.1
Content-Type: application/json
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: get_weather

{"jsonrpc": "2.0", "id": 1, "method": "tools/call",
 "params": {"name": "get_weather", "arguments": {"location": "Seattle, WA"},
  "_meta": {"io.modelcontextprotocol/protocolVersion": "2026-07-28",
            "io.modelcontextprotocol/clientInfo": {"name": "ExampleClient", "version": "1.0.0"},
            "io.modelcontextprotocol/clientCapabilities": {}}}}

The server answers with either a single JSON object or a Server-Sent Events stream scoped to that one request, carrying progress notifications and then the final result.

Sequence diagram in which a web client asks a chain of DNS servers for an IP address and then sends a single HTTP GET request to a web server, which returns a response
A sequence diagram of a web request. Under the new revision, each MCP message is a self-contained HTTP request like the final exchange here. Diagram: Amakuha / Wikimedia Commons, CC BY-SA 4.0

Headers that intermediaries can trust

Three details matter for operators. First, the Mcp-Method header is required on every request, and Mcp-Name on tools/call, resources/read and prompts/get, so a gateway, rate limiter or firewall can route and meter without parsing JSON. Names that aren't plain ASCII travel in a Base64 sentinel format, =?base64?…?=.

Second, servers must reject any request whose headers disagree with the body, returning 400 with a HeaderMismatch error, code -32020. The spec explains why: otherwise a load balancer could route on one value while the server executes another. It also advises intermediaries that enforce policy on these headers to reject requests claiming an older protocol version, since those carry no such guarantee.

Third, a tool can mark parameters with x-mcp-header so clients mirror them into Mcp-Param-{Name} headers. The rules are strict: only string, integer and boolean parameters qualify, the property must be reachable through plain properties keys, and clients must drop any tool whose annotations break the rules. That lets you route on, say, a region argument.

Architecture diagram of an API gateway in which Envoy proxies check OAuth tokens and a rate-limit service backed by Redis before routing requests to API servers
Wikimedia's API gateway, where proxies authenticate, rate-limit and route requests. MCP's new headers let this kind of gateway act on tool calls without reading the JSON body. Diagram: APaskulin (WMF) / Wikimedia Commons, CC BY-SA 4.0

Servers must also implement a new server/discover method that advertises supported versions, capabilities and identity. Clients may call it first, but don't have to. A version the server doesn't support gets an UnsupportedProtocolVersionError listing the ones it does.

How servers ask the client for input now

The trickiest removal is server-initiated requests. Sampling, elicitation and roots used to be requests the server sent to the client over an open stream. With Multi Round-Trip Requests (MRTR), the server instead returns a result with resultType: "input_required":

{"jsonrpc": "2.0", "id": 1,
 "result": {"resultType": "input_required",
  "inputRequests": {"github_login": {"method": "elicitation/create",
    "params": {"mode": "form", "message": "Please provide your GitHub username",
               "requestedSchema": {"type": "object",
                 "properties": {"name": {"type": "string"}}, "required": ["name"]}}}},
  "requestState": "AEAD-protected blob"}}

The client gathers the answers and retries the original call with a new JSON-RPC ID, the matching inputResponses, and the exact requestState it received. Whichever instance gets the retry has everything it needs. Only tools/call, resources/read and prompts/get may return this kind of result.

The requestState field is where developers will need to be careful. The spec says servers must treat it as attacker-controlled input. If it affects authorization, resource access or business logic, it must be integrity-protected with something like an HMAC or AEAD, and state that fails verification must be rejected. To limit replay, servers should bind it to the authenticated user, give it a short expiry, and tie it to the method and a digest of the original parameters. The spec also warns that none of this guarantees single use; one-time actions still need a server-side check.

Streams, tasks and caching

AreaRevisions 2025-03-26 to 2025-11-25Revision 2026-07-28
Setupinitialize handshake, optional Mcp-Session-IdNo handshake; metadata on every request; optional server/discover
Server-to-client requestsSent over SSE streamsReturned as input_required results (MRTR)
Change notificationsStandalone GET streamOne subscriptions/listen POST stream
Resuming dropped streamsLast-Event-ID redeliveryRemoved; client re-sends with a new ID
Long-running workExperimental tasks in coreio.modelcontextprotocol/tasks extension, polled with tasks/get
List cachinglistChanged notificationsRequired ttlMs and cacheScope hints as well
Logginglogging/setLevelPer-request log level in _meta; Logging deprecated

Source: the 2026-07-28 changelog and transport specification.

A few of these deserve a closer look:

  • One long-lived stream remains. Change notifications for tool, prompt and resource lists now arrive on a single subscriptions/listen POST response stream. The spec encourages periodic SSE comment lines as keep-alives and tells servers to send X-Accel-Buffering: no so proxies such as nginx don't hold events back.
  • State becomes explicit. The maintainers suggest that a server needing cross-call state mint a handle from a tool and let the model pass it back as an ordinary argument. List endpoints no longer vary per connection.
  • Long work moves to tasks. The tasks extension replaces the blocking tasks/result with polling via tasks/get, adds tasks/update for client input, and lets servers return task handles without per-request opt-in. That suits platforms with request timeouts.
  • Caches get hints. Results from tools/list, prompts/list, resources/list, resources/templates/list and resources/read must include ttlMs and cacheScope (public or private). Servers should return tools in a deterministic order, which the changelog says improves client caching and LLM prompt-cache hit rates.
  • Tracing is standardized. The spec documents traceparent, tracestate and baggage keys in _meta for OpenTelemetry context propagation.
Diagram of a cloud containing application, platform and infrastructure services, surrounded by laptops, phones, tablets, desktops and servers
A common picture of cloud computing layers. Without sessions, any instance of an MCP server can answer any request, which makes servers easier to run on managed platforms. Diagram: Sam Johnston / Wikimedia Commons, CC BY-SA 3.0

Authorization tightens

The authorization changes are smaller but matter for anyone running a remote server with OAuth.

  • Issuer validation. Authorization servers should include the iss parameter in authorization responses, per RFC 9207, and clients must validate it against the recorded issuer before redeeming the code. Appwrite notes this, together with credential binding, defends against server mix-up attacks.
  • Credentials bound to issuers. Clients must key stored credentials by issuer, must never reuse them with a different authorization server, and must re-register when the server changes.
  • Registration. Clients must send an application_type during Dynamic Client Registration to avoid OpenID Connect redirect conflicts. Dynamic Client Registration itself is now deprecated in favor of Client ID Metadata Documents, though it stays available for authorization servers that don't support them.
Diagram of the OAuth triangle with numbered arrows between a user, a client tool and a provider
The three parties in an OAuth authorization: the user, the client tool and the provider. The new revision requires MCP clients to tie credentials to the authorization server that issued them. Diagram: EpochFail / Wikimedia Commons, CC BY-SA 3.0

What breaks, and what is still open

The main cost is the transition. Older clients still send initialize. The compatibility rules put the burden on clients: try a modern request first, and fall back to initialize only if a 400 comes back without a recognized modern JSON-RPC error. A server that only speaks the new revision will turn away clients that haven't updated. If you serve desktop apps or tools you don't control, expect to run both behaviors for a while.

Several smaller changes can also break code:

  • ping, logging/setLevel and notifications/roots/list_changed are removed outright.
  • The resource-not-found error code changes from -32002 to -32602.
  • Every result now carries a required resultType; clients must treat results without it, from older servers, as complete.
  • The URL-mode elicitation completion notification and its elicitationId, added only in 2025-11-25, are gone.

Roots, Sampling, Logging and the legacy HTTP+SSE transport are deprecated rather than removed. Under the new feature lifecycle policy, deprecated features keep working for at least twelve months. The changelog suggests replacements: pass directories through tool parameters instead of Roots, call LLM provider APIs directly instead of Sampling, and log to stderr or OpenTelemetry instead of Logging.

Some questions don't have answers yet. Moving state into handles and requestState gives developers more design responsibility, as Appwrite points out, and gets it wrong in new ways, such as unsigned state blobs. The subscriptions/listen stream is still long-lived, so platforms with hard response-duration caps will see clients reconnecting. And the ecosystem's many third-party servers will move at their own pace. The announcement includes a note from Manufact saying its SDK v2 is about 83 percent smaller and 25 percent faster thanks to a new client-server split, and FastMCP says version 4.0 ships support for the new features, but most servers depend on their maintainers upgrading.

Migration checklist for server operators

  • Upgrade to an SDK release that supports 2026-07-28 and read its migration notes.
  • Move anything you kept in session memory into explicit handles, requestState, or your own datastore.
  • Replace server-initiated sampling, elicitation and roots calls with Multi Round-Trip Requests, and sign any requestState that affects access or logic.
  • Remove session affinity from your load balancer once old clients are gone.
  • Validate MCP-Protocol-Version, Mcp-Method and Mcp-Name against the body in the server. Gateways that route on these headers should reject requests that claim an older protocol version.
  • Add ttlMs and cacheScope to list results, and sort your tool list.
  • If you drop older revisions, answer GET and DELETE on the MCP endpoint with 405 and ignore incoming Mcp-Session-Id and Last-Event-ID headers.
  • Swap the Logging feature for stderr or OpenTelemetry. Log level now travels per request in _meta.
  • Validate iss on OAuth responses and plan a move from Dynamic Client Registration to Client ID Metadata Documents.

Practical tip for hackathon teams: a new MCP server built against 2026-07-28 can be a single stateless function. Skip sessions entirely, keep state in a handle the model passes back, and deploy behind any HTTP load balancer. Keep a fallback path for initialize if your demo needs to work with clients that haven't updated.

The transport spec keeps its security basics: servers must validate the Origin header to block DNS rebinding, and should bind to 127.0.0.1 rather than all interfaces when running locally and authenticate every connection.

What to watch next

As of July 30, the work shifts from the spec to the ecosystem. Watch for the Rust SDK to leave beta, for major desktop and IDE clients to ship releases that send per-request metadata, and for hosted platforms and gateways to add routing on Mcp-Method and Mcp-Name. The deprecation clock for Roots, Sampling, Logging and HTTP+SSE has started; under the lifecycle policy, none of them can be removed before July 2027. How quickly authorization servers adopt Client ID Metadata Documents will decide how long Dynamic Client Registration lingers.

Sources