MCP Went Stateless and Broke My Lambda MCP Server Anyway
Blog article/Blog archive

MCP Went Stateless and Broke My Lambda MCP Server Anyway

The 2026-07-28 MCP spec deleted the initialize handshake and Mcp-Session-Id, which should have been free for a Lambda server that was already stateless. It wasn't: a CORS allow-list, a missing server/discover, a 404 that should have been a 405, and a notification stream an HTTP API can't hold open all broke behind API Gateway.

Aug 19, 202613 min read0 comments37 views
MCPAWSLambdaServerlessTerraformAI

My MCP server went dark on a Tuesday. Claude Desktop showed the hammer icon with zero tools under it. No error dialog, no red banner, just an empty list where get_user_stats, list_recent_events, and query_documents used to be.

CloudWatch had the answer, and it was embarrassing. The client had started speaking a protocol my server couldn't answer, and the way it failed looked like nothing at all: no error in the app, no rejected request you could reason about, just an empty tool list that turned out to be a negotiation failure in disguise.

That was the 2026-07-28 MCP spec landing on a server I built for the 2025 spec. I wrote about that server in Deploying an MCP Server on AWS Lambda. This post is the repair job.

The annoying part is that I thought I was immune. The headline change in 2026-07-28 is that MCP dropped protocol-level sessions, and my server never had sessions. I'd set sessionIdGenerator: undefined on day one because Lambda can't guarantee the same instance handles consecutive requests. I read the changelog, saw "sessions removed," and figured I'd already paid that tax.

The session removal was free. Everything the spec did around it was not.

What 2026-07-28 actually deleted

Worth being precise here, because the summaries floating around all lead with "MCP is stateless now" and stop there.

The initialize request and the notifications/initialized notification are gone. So is the Mcp-Session-Id header and the whole session lifecycle attached to it. Instead, every single request carries its own context in _meta:

{
  "jsonrpc": "2.0",
  "id": 1,
  "method": "tools/call",
  "params": {
    "name": "get_user_stats",
    "arguments": { "userId": "u_8812", "period": "30d" },
    "_meta": {
      "io.modelcontextprotocol/protocolVersion": "2026-07-28",
      "io.modelcontextprotocol/clientInfo": {
        "name": "ExampleClient",
        "version": "1.0.0"
      },
      "io.modelcontextprotocol/clientCapabilities": {}
    }
  }
}

Protocol version, client identity, and capabilities ride along on every call. Nothing is negotiated once and remembered.

Then there's the part nobody put in the headline. Selected body fields are now mirrored into HTTP headers, and they're required:

POST /mcp HTTP/1.1
Content-Type: application/json
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: get_user_stats

Mcp-Method mirrors method. Mcp-Name mirrors params.name or params.uri, and it's required on tools/call, resources/read, and prompts/get. The point is that a load balancer or WAF can route and rate-limit on the operation without parsing a JSON body. Good idea. It also means every box between the client and your function now has an opinion about three headers it's never seen.

There's more in the changelog that matters if you're on Lambda:

  1. server/discover is a new RPC that servers MUST implement, advertising supported protocol versions, capabilities, and identity.
  2. The HTTP GET endpoint is gone, along with resources/subscribe. Change notifications now come from a long-lived subscriptions/listen POST-response stream.
  3. SSE resumability is gone. No Last-Event-ID, no event IDs, no redelivery. A broken stream means the client re-issues the request with a new ID.
  4. ping, logging/setLevel, and notifications/roots/list_changed are removed.
  5. Server-initiated requests like sampling and elicitation are replaced by Multi Round-Trip Requests. The server returns resultType: "input_required" and the client retries the original request with the answers attached.
  6. Roots, Sampling, and Logging are deprecated, with a twelve-month minimum removal window under the new feature lifecycle policy.

Six things on that list touched my three-tool server. Let me go through the ones that actually cost me time.

Break 1: the request path didn't know the new headers

This is the failure I reached for first, and it's the dumbest one in the post.

The first thing I reached for was the CORS config, because that's the layer where a browser-based client dies with no Lambda logs. The CloudFormation template in my original post had this:

CorsConfiguration:
  AllowOrigins:
    - "*"
  AllowMethods:
    - POST
    - OPTIONS
  AllowHeaders:
    - Authorization
    - Content-Type

Authorization and Content-Type. That was the complete set of headers MCP needed in 2025. Now the client sends MCP-Protocol-Version, Mcp-Method, and Mcp-Name on top of those, and the path between client and function has to carry three more headers than it did before.

I've since moved this stack to Terraform. Here's the corrected API:

resource "aws_apigatewayv2_api" "hrr_mcp" {
  name          = "hrr-mcp-server"
  protocol_type = "HTTP"

  cors_configuration {
    allow_origins = ["https://claude.ai"]
    allow_methods = ["POST", "OPTIONS"]
    allow_headers = [
      "authorization",
      "content-type",
      "mcp-protocol-version",
      "mcp-method",
      "mcp-name",
    ]
    max_age = 600
  }
}

Two notes on that block.

I dropped allow_origins = ["*"] while I was in there. A wildcard origin on an endpoint that reads my production database was always lazy, and the migration was a decent excuse to fix it.

The CORS list matters for browser-based clients specifically. If the client runs in a page (Claude.ai's web app, or any embedded MCP client), the browser enforces CORS, and a missing allow-list entry shows up as a failed preflight: the browser sends OPTIONS with the requested headers, API Gateway compares them against the allow-list, and the preflight fails before Lambda ever runs. That's the whole job of allow_headers — it only shapes the preflight response. Native desktop clients skip the preflight entirely, and API Gateway's HTTP APIs forward every request header to Lambda regardless of this block. So for the desktop failure the culprit is elsewhere in the path: the note below about WAF, CloudFront, and ALB is the one that matters for native clients. Where headers do arrive intact, Break 4's header-vs-body validation is what decides the request lives or dies.

The second note is a trap if you use the new x-mcp-header feature. Servers can annotate a tool parameter in its inputSchema and the client mirrors that value into an Mcp-Param-{Name} header. Handy for routing on a tenant or region without cracking the body open. But API Gateway's allow_headers doesn't take wildcards, so there's no mcp-param-*. You enumerate every single one:

allow_headers = [
  "authorization",
  "content-type",
  "mcp-protocol-version",
  "mcp-method",
  "mcp-name",
  "mcp-param-tenant",
  "mcp-param-region",
]

Add a new annotated parameter, add a line here, or that tool silently stops working for browser clients while it keeps working fine in your local tests. I'd write a comment above the block pointing at whichever file holds your tool schemas. I did.

If you have a WAF, CloudFront, or an ALB in the path, walk the whole chain. Anything that strips unknown Mcp-* headers fails every request, and the failure looks like a protocol bug rather than an infrastructure one. For a native desktop client this is a silent outage: no browser, no preflight, the headers just don't arrive intact, and the app renders the result as an empty tool list.

Break 2: server/discover is mandatory, and mine returned -32601

Clients can call server/discover before anything else to find out which protocol versions you speak. Servers MUST implement it. Mine didn't exist, so the request fell through to the SDK's unknown-method path and came back as JSON-RPC -32601, Method not found.

This is the break that actually took my server down. A client running version negotiation in auto mode uses server/discover as its "is this server modern?" probe. Fail the probe and the client concludes you're a 2025-era server and falls back to sending initialize. Which my server also didn't handle anymore, once I'd upgraded the SDK. Two failed probes, zero tools, no useful error — that's the empty hammer menu I opened the post with.

The v2 TypeScript SDK handles server/discover for you as long as you're building the server through its handler factory. That was the real fix: stop hand-rolling the transport.

Break 3: a 404 that should have been a 405

My API Gateway had exactly one route, POST /mcp. Correct for 2025, correct for 2026, and still wrong.

The spec says a server that only speaks 2026-07-28 and receives a GET or DELETE on the MCP endpoint SHOULD respond 405 Method Not Allowed. Mine responded with API Gateway's stock 404:

{ "message": "Not Found" }

That body is the problem. The backward-compatibility flow in the spec has clients POST first, then inspect the response before falling back. If the body is a recognized modern JSON-RPC error, the client knows the server is modern and retries properly. If the body is empty or unrecognized, the client assumes it's talking to a legacy HTTP+SSE server and goes looking for a GET stream with an endpoint event.

{"message": "Not Found"} is not a JSON-RPC error. So an older client hitting my endpoint would misclassify it and start a fallback dance that could never succeed.

The fix is a route that reaches Lambda so Lambda can answer properly:

resource "aws_apigatewayv2_route" "hrr_mcp_post" {
  api_id    = aws_apigatewayv2_api.hrr_mcp.id
  route_key = "POST /mcp"
  target    = "integrations/${aws_apigatewayv2_integration.hrr_mcp.id}"
}

# GET and DELETE exist only so the function can return a spec-correct 405
# instead of API Gateway's stock {"message": "Not Found"}.
resource "aws_apigatewayv2_route" "hrr_mcp_legacy_methods" {
  for_each = toset(["GET /mcp", "DELETE /mcp"])

  api_id    = aws_apigatewayv2_api.hrr_mcp.id
  route_key = each.value
  target    = "integrations/${aws_apigatewayv2_integration.hrr_mcp.id}"
}

And a short circuit at the top of the handler, before auth, because there's no point validating a bearer token on a request you're rejecting on method alone. Note the error code here: I'm reusing -32601 on purpose. A legacy client sends GET, gets back a JSON-RPC-shaped -32601, recognizes it as a JSON-RPC error, and knows the server is modern — it just used a method this server doesn't serve. (A bare 405 with no body, by contrast, would read as "no server here" and kick off the fallback.) The nuance is that in Break 2 the same code meant "I'm not a modern server, fall back to initialize"; here it means "I am a modern server, use a different method." Same numeric, opposite directions — the client tells them apart by which RPC it was calling.

const HRR_ALLOWED_METHOD = "POST";

function hrr_methodNotAllowed(event) {
  const method = event.requestContext?.http?.method;
  if (method === HRR_ALLOWED_METHOD) return null;

  return {
    statusCode: 405,
    headers: { "Content-Type": "application/json", Allow: "POST" },
    body: JSON.stringify({
      jsonrpc: "2.0",
      id: null,
      error: { code: -32601, message: `${method} is not supported on this endpoint` },
    }),
  };
}

Small thing. Costs nothing. Saves a client from an unwinnable fallback loop.

Break 4: header and body have to agree, and you have to check

The mirrored headers aren't just hints. Servers that process the body MUST verify the header values match the body values, and reject mismatches with HTTP 400 and JSON-RPC error -32020, HeaderMismatch.

The reasoning is worth understanding because it changes how you think about the endpoint. If a load balancer routes on Mcp-Name: get_user_stats while your function executes params.name = "delete_everything" from the body, two components in the same request path are working from different truths. That's a confused deputy waiting to happen. The spec closes it by making the server the thing that refuses to let them disagree.

Same rule covers the required headers being missing entirely. Missing MCP-Protocol-Version, Mcp-Method, or Mcp-Name is a -32020, not a generic 400.

The v2 SDK's handler validates all three on every modern request, so if you're using it you get this behavior without writing it. Which is exactly why I stopped writing it. There's also -32022 for UnsupportedProtocolVersion, which returns the list of versions you do support so the client can retry instead of guessing.

One detail that'll bite you if you hand-roll the comparison: Mcp-Name values that aren't plain ASCII get Base64-encoded with a sentinel wrapper, =?base64?{value}?=. It's lowercase, there's no charset field, and it's not an RFC 2047 encoded-word despite looking like one — the spec just uses those two markers to say "what follows is Base64." Decode before comparing, or every tool with a non-ASCII name in its arguments fails validation.

Break 5: subscriptions/listen can't live behind an HTTP API

This one isn't a bug I hit. It's a wall I walked up to and decided not to climb.

Change notifications used to come over a standalone GET SSE stream. That's gone. Now a client opens a subscriptions/listen POST and the response is an SSE stream that stays open, delivering the notification types the client opted into: toolsListChanged, resourcesListChanged, and friends. The spec even suggests emitting a periodic SSE comment line as a keep-alive so intermediaries don't hang up on it during quiet periods.

"Stays open" and "API Gateway HTTP API" don't go together. HTTP APIs cap the integration timeout at 30 seconds, and that's a hard ceiling. API Gateway added response streaming for Lambda proxy integrations in November 2025, but only on REST APIs, not HTTP APIs. My original post picked HTTP API specifically because it's cheaper and lower latency. That decision now costs me a protocol feature.

The options, in increasing order of effort:

  1. Don't advertise it. Skip the subscription capabilities entirely. A well-behaved client won't call subscriptions/listen if you never claimed to support it.
  2. Lambda Function URL. Native response streaming, no API Gateway in front of it, and the invoke mode is a one-line change. You give up API Gateway's throttling, custom domains, and usage plans, so you're back to doing auth yourself in the handler. Which I was already doing.
  3. Move to a REST API with responseTransferMode set to STREAM. You get streaming and keep API Gateway, at REST API pricing and REST API idle-timeout rules.

I took option one. My three tools are request-response reads against my own database. The tool list changes when I deploy, which is not an event any client needs pushed to it in real time. Advertising a capability I can't hold open on this stack would just be a promise I'd break under load.

Worth saying plainly: if your MCP server genuinely needs push notifications, API Gateway HTTP API is the wrong front door now. That's a real architectural constraint the new spec introduces for serverless, and no amount of SDK upgrading routes around it.

The SDK move, and deleting code I was proud of

The TypeScript SDK split. @modelcontextprotocol/sdk reached 1.30.0 and the new generation shipped as separate packages at 2.0.0: @modelcontextprotocol/core, /client, /server, and /server-legacy, plus adapters for Express, Fastify, Hono, and Node.

My old handler created a fresh StreamableHTTPServerTransport per invocation, connected the server to it, processed the request, then closed both. I'd written a whole paragraph in the last post about learning the hard way that you can't cache the transport across invocations.

That advice is now obsolete in the best way. The v2 server package hands you a factory that builds a fresh server per request by default, because that's what the stateless spec implies:

import { createMcpHandler, McpServer } from "@modelcontextprotocol/server";
import { z } from "zod";

const hrr_handler = createMcpHandler(() => {
  const server = new McpServer(
    { name: "hrr-saas-tools", version: "2.0.0" },
    { capabilities: { tools: {} } }
  );

  server.registerTool(
    "get_user_stats",
    {
      description: "Get activity statistics for a specific user including login count, API calls, and last active date",
      inputSchema: z.object({
        userId: z.string().describe("The user ID to look up"),
        period: z.enum(["7d", "30d", "90d"]).describe("Time period for the stats"),
      }),
    },
    async ({ userId, period }) => {
      const stats = await fetchUserStats(userId, period);
      return { content: [{ type: "text", text: JSON.stringify(stats, null, 2) }] };
    }
  );

  return server;
});

// createMcpHandler's .fetch() expects a web-standard Request and returns a
// Response. API Gateway hands Lambda a proxy event instead, so these two
// convert both ways. The SDK ships toNodeHandler() for Express/Fastify/Hono/
// raw http, but nothing Lambda-specific, so this is the small adapter that's
// left to write by hand.
function hrr_toWebRequest(event) {
  const host = event.headers?.host ?? event.requestContext.domainName;
  const query = event.rawQueryString ? `?${event.rawQueryString}` : "";

  return new Request(`https://${host}${event.rawPath}${query}`, {
    method: event.requestContext.http.method,
    headers: event.headers,
    body: event.isBase64Encoded ? Buffer.from(event.body, "base64") : event.body,
  });
}

async function hrr_toApiGatewayResponse(response) {
  return {
    statusCode: response.status,
    headers: Object.fromEntries(response.headers),
    body: await response.text(),
  };
}

export const handler = async (event) => {
  const notAllowed = hrr_methodNotAllowed(event);
  if (notAllowed) return notAllowed;

  const auth = validateApiKey(event);
  if (!auth.valid) {
    return {
      statusCode: 401,
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({ error: auth.error }),
    };
  }

  const response = await hrr_handler.fetch(hrr_toWebRequest(event));
  return hrr_toApiGatewayResponse(response);
};

The schemas and business logic came over unchanged; only the registration call changed shape. The Zod schemas, the descriptions I spent hours tuning, the business logic behind them, all of it survived. What I deleted was transport plumbing: the manual construct, the server.connect(transport), the finally block closing both. That's the code the spec change made unnecessary, and it's the code I'd got wrong twice.

My bearer-token check stayed exactly where it was, in front of everything. Nothing in 2026-07-28 changes how you authenticate a request. It does deprecate OAuth Dynamic Client Registration in favor of Client ID Metadata Documents, but that's a concern for servers doing real OAuth, not for one API key in an environment variable.

One thing worth being precise about, because it's easy to mistake for a session replacement: requestState is not general cross-call state, and my three tools don't touch it. It's scoped to Multi Round-Trip Requests specifically — the opaque string a server can attach to an InputRequiredResult when it needs more input (an elicitation, a sampling call) before it can finish a request. The client echoes it back byte for byte, but only on the retry of that exact request; the spec says outright it MUST NOT be used for any other request the client happens to be sending in parallel. If the value can influence authorization or business logic, the server has to protect its integrity — the spec's own wording is "e.g. HMAC or AEAD," not one mandated algorithm — and reject anything that fails verification. The SDK ships a codec for sealing it. Think of it less as a session and more as a claim check for one unfinished request, and for Lambda that's still a better fit than a session ID that implies an instance remembers you.

What the migration actually improved

I've been grumbling for a whole post, so here's the other side.

Under the old spec, a fresh client connection cost three round trips before a single tool ran: initialize, notifications/initialized, tools/list. On Lambda, each of those could land on a cold instance. Now the first thing a client sends can be the tool call itself, with everything the server needs in _meta.

Sessions were also the reason a horizontally scaled MCP server needed sticky routing or a shared session store. That's gone. Any request can land on any instance, which is what Lambda was always going to do anyway. The protocol finally agrees with the runtime.

I'm still running one provisioned concurrency instance at roughly $5/month, for the same reason as before: tool calls feel synchronous to whoever's typing, and a cold start on top of model thinking time is noticeable. But the case for it is weaker now that there's no handshake to pay for on top.

The checklist

If you've got an MCP server on Lambda built against the 2025 spec, here's the order I'd do it in. Most of the pain is in the first few steps, and none of them involve your tool code.

  1. Confirm server/discover answers. Curl it. If it returns -32601, clients will misidentify you as a legacy server and fall back to an initialize your upgraded SDK won't serve — the empty-tool-list outage. This is the one that took my server down.
  2. Fix CORS. Add mcp-protocol-version, mcp-method, and mcp-name to allow_headers, plus every mcp-param-* you use. Browser-based clients fail hard here with no Lambda logs.
  3. Walk the whole request path. CloudFront, WAF, ALB, custom authorizers. Anything that strips unknown Mcp-* headers fails every request. For a desktop client this is a silent outage — the headers vanish before Lambda, and the app just shows an empty tool list.
  4. Add GET and DELETE routes pointing at the same integration, and return a JSON-RPC-shaped 405 from the handler.
  5. Move to the v2 SDK packages and rebuild the handler around the factory. Your tool definitions port over as they are.
  6. Decide about subscriptions/listen. Either don't advertise the capability, or move off HTTP API to a Function URL or a streaming REST API.
  7. Audit for deprecated features. Roots, Sampling, and Logging still work but are on a twelve-month clock. Log to stderr or OpenTelemetry instead.
  8. Test with an actual client, not just curl. The header mirroring, the encoding rules, and the version negotiation all live on the client side, and curl won't do any of it for you.

The whole repair took me an afternoon, and most of that was staring at API Gateway access logs before I understood that the outage was a version-negotiation failure, not a crash or an error I could search for. The lesson I'll actually keep is that "we removed sessions" sounded like a change that couldn't touch a server without sessions, and the thing that actually took me down was a handshake my client assumed still existed.

Hope this saves you the afternoon. If you're migrating an MCP server on AWS and you hit something I didn't, or you found a cleaner way to keep a notification stream alive on serverless, come tell me on X at https://x.com/harundotdev.

Get the next note

I email when a new post goes up. One send a week, and only if there's something new.

Want this applied on your account? Start with Infrastructure as Code.

Related reading

More posts, sliding underneath the article

Kept below the post instead of in a sidebar, with a slow continuous motion for a cleaner editorial feel.

On this post

Comments

A reply stays under the note it answers.

0 comments

No comments yet.

If you have a note on MCP Went Stateless and Broke My Lambda MCP Server Anyway, sign in and leave it.