
Lambda SnapStart with Terraform - Cutting Cold Starts Without the Hacks
How to enable Lambda SnapStart for Java and Python runtimes using Terraform, measure cold start improvements, and compare costs against provisioned concurrency.

How to enable Lambda SnapStart for Java and Python runtimes using Terraform, measure cold start improvements, and compare costs against provisioned concurrency.
Cold starts are the most complained-about thing in serverless. I've seen teams abandon Lambda entirely because they couldn't get init times under control. They moved to ECS Fargate or EC2, traded cold starts for always-on costs, and regretted it six months later when the bill came in.
Lambda SnapStart changes the math. It works by taking a snapshot of the initialized execution environment (the Firecracker microVM) and caching it. When a new invocation comes in, Lambda loads the snapshot instead of running init from scratch. The result: cold starts drop from multiple seconds to sub-200ms for Java, and you get measurable improvements for Python too.
In this post I'll show you how to enable SnapStart with Terraform, measure the improvement with CloudWatch Logs, compare costs against provisioned concurrency and keep-warm tricks, and lay out when you should and shouldn't use it.
When a Lambda function is invoked for the first time (or after being idle), AWS has to spin up a new execution environment. That process has three phases:
Phase 2 (init) is what kills you. A Spring Boot Lambda with Hibernate, JDBC connection pooling, and Jackson serialization can take 6-10 seconds just to initialize. Even a well-optimized Python function with a heavy library like Pandas or a large Pydantic model can spend 1-3 seconds in init.
SnapStart works by running init once in a sandboxed Firecracker microVM, taking a full snapshot of memory and disk state, and encrypting that snapshot. Future invocations restore from the snapshot instead of re-initializing. The initialization still happens once per snapshot version, but after that every cold start is just a restore.
The key constraint: the snapshot happens after init but before your first invocation handler runs. So anything in your handler code or class initialization that depends on runtime state (environment variables, Secrets Manager lookups, database connection state) needs to be refreshed on each invocation. More on this later.
As of writing, SnapStart is available for:
Node.js, Go, .NET, and Ruby don't support SnapStart yet. For Node.js and Go the cold starts are usually fast enough that you don't need it, but if you're running a heavy Node.js framework or a compiled Go binary with lots of reflect-based initialization, it would be a welcome addition.
Python support is relatively new and still has some rough edges. The snapshot captures Python's global interpreter state, which means objects created during module import time get preserved. Dynamic imports and runtime code generation (common with Pydantic v2 and some ORMs) can cause problems.
Here's the Terraform configuration I use in my projects. It's not complicated, but there are a few gotchas to watch for.
You enable SnapStart on the Lambda function resource with the snap_start block and set the apply_on value to PublishedVersions. SnapStart only works on published versions, not the $LATEST alias.
resource "aws_lambda_function" "api" {
filename = "function.zip"
function_name = "my-api-function"
role = aws_iam_role.lambda.arn
handler = "com.example.App::handleRequest"
runtime = "java21"
memory_size = 1024
timeout = 30
snap_start {
apply_on = "PublishedVersions"
}
}
resource "aws_lambda_function_event_invoke_config" "api" {
function_name = aws_lambda_function.api.function_name
qualifier = aws_lambda_function.api.version
destination_config {
on_failure {
destination = aws_sqs_queue.dlq.arn
}
}
}
output "lambda_version" {
value = aws_lambda_function.api.version
}
That's it for the basic setup. But there's an important detail: SnapStart only applies to published versions. The $LATEST alias never uses SnapStart, so if you're invoking $LATEST in development you won't see any difference.
You need to explicitly publish a version and point your alias at it. I use aws_lambda_function with publish = true to auto-publish on every deploy, then create an alias for my staging/production traffic.
resource "aws_lambda_function" "api" {
filename = "function.zip"
function_name = "my-api-function"
role = aws_iam_role.lambda.arn
handler = "com.example.App::handleRequest"
runtime = "java21"
memory_size = 1024
timeout = 30
publish = true
snap_start {
apply_on = "PublishedVersions"
}
}
resource "aws_lambda_alias" "prod" {
name = "prod"
function_name = aws_lambda_function.api.function_name
function_version = aws_lambda_function.api.version
}
resource "aws_lambda_function_event_invoke_config" "prod" {
function_name = aws_lambda_function.api.function_name
qualifier = aws_lambda_alias.prod.name
}
One thing that tripped me up: SnapStart does not work with provisioned concurrency on the same version. If you have both enabled, SnapStart is silently disabled and you get regular cold starts. The Terraform plan won't warn you about this. I caught it only after noticing cold start times were still 4 seconds despite SnapStart being "enabled."
The Python configuration is identical. The same snap_start block works for any supported runtime.
resource "aws_lambda_function" "inference" {
filename = "function.zip"
function_name = "inference-processor"
role = aws_iam_role.lambda.arn
handler = "handler.lambda_handler"
runtime = "python3.13"
memory_size = 512
timeout = 60
publish = true
snap_start {
apply_on = "PublishedVersions"
}
}
resource "aws_lambda_alias" "prod" {
name = "prod"
function_name = aws_lambda_function.inference.function_name
function_version = aws_lambda_function.inference.version
}
If you're using the AWS CLI or SDK instead of Terraform, the equivalent is passing --apply-on PublishedVersions to update-function-configuration.
Let's talk about real numbers. I set up a test with a Java 21 Quarkus REST API that connects to DynamoDB and serializes JSON responses. Here are the cold start times I measured before and after SnapStart, averaged over 50 invocations each (all with 1024MB memory):
Java gets the most dramatic wins because its JVM startup and class loading is the slowest phase. A Spring Boot app that took 8.5 seconds to cold start now takes under half a second. Python benefits too -- the import pandas overhead that used to take 1.5 seconds is now eliminated.
You can measure this yourself with CloudWatch Logs. Look for the Init Duration field in the REPORT log line:
REPORT RequestId: abc-123 Duration: 45.12 ms Billed Duration: 46 ms
Memory Size: 1024 MB Max Memory Used: 178 MB
Init Duration: 180.23 ms
With SnapStart, the Init Duration is the time to restore the snapshot. Without it, Init Duration is the full init phase including class loading and static initialization.
You can query this at scale with a CloudWatch Logs Insights query:
filter @type = "REPORT"
| stats avg(@initDuration) as avgInit,
pct(@initDuration, 95) as p95Init,
pct(@initDuration, 99) as p99Init
by bin(1h)
| sort @timestamp desc
| limit 50
Run this query on a function before enabling SnapStart, then after. The difference is immediate and obvious.
SnapStart isn't free. There are specific tradeoffs you need to understand before enabling it across your whole estate.
The snapshot captures the /tmp directory state at snapshot time. Any files written to /tmp during your handler invocation are lost on the next invocation because the snapshot is restored cleanly. This is actually the correct behavior -- you don't want stale temp files from one invocation leaking into another. But if your code relies on writing downloaded models or cache files to /tmp and expecting them to persist across invocations, you need to redesign that flow.
Firecracker microVMs don't support the fork() system call. If your code spawns child processes (common with multiprocessing in Python or ProcessBuilder in Java), those calls will fail at restore time. The snapshot is taken after init, so if you fork during init, the child process state is captured and restored, but forking during invocation is not supported.
Database connections, HTTP clients, and SDK clients that were established during init get captured in the snapshot. When the snapshot is restored on a different host, those connections are stale. Lambda SnapStart handles some of this automatically -- AWS SDK v2 for Java has built-in SnapStart support that refreshes credentials and connections on restore. But if you're using raw JDBC connections or custom socket pools, you need to implement a restore hook that refreshes them on each snapshot restore.
For Java, you register a com.amazonaws.services.lambda.runtime.snapstart.SnapStartRestoreListener:
public class MyRestoreHook implements SnapStartRestoreListener {
@Override
public void restore() {
// Refresh database connections
DatabaseConnectionPool.refresh();
// Refresh HTTP client connections
HttpClientPool.refresh();
// Re-resolve Secrets Manager secrets
SecretsCache.refresh();
}
}
This hook runs on every snapshot restore, before the handler invocation starts. If you skip this, your first few requests after a scale-up will hit stale connections and throw exceptions. AWS SDK v2 handles this automatically. Custom libraries usually don't.
For Python, the equivalent is a bit less formal. You can use the __init__ module pattern or a decorator that checks a flag on each invocation. The important thing is to not assume that anything created during import time (like an HTTP session or database engine) still works without being refreshed.
I mentioned this above but it bears repeating: SnapStart and provisioned concurrency are mutually exclusive on the same version. If you want both, you need to split traffic across two versions -- one with SnapStart for the base capacity and one with provisioned concurrency for the buffer. This is awkward and I don't recommend it. Pick one strategy per function.
Each time you publish a new Lambda version with SnapStart enabled, AWS has to run init once and take the snapshot. This adds 30-90 seconds to your deployment pipeline depending on how long your init phase takes. For most teams this is a non-issue -- you deploy a few times a day at most -- but if you're doing continuous deployment with dozens of iterations per hour, the delay adds up.
Here's how I think about these three strategies. I've used all of them in production and none is universally better.
| Strategy | Cost/Month (1024MB) | Cold Start Latency | Scale Handling | Setup Complexity |
|---|---|---|---|---|
| Standard (no optimization) | $0 extra | 1-10 seconds | Unlimited (with cold starts) | None |
| Keep-Warm | ~$0.50 | 1-10 seconds (except 1 warm) | Does not scale | Trivial |
| Provisioned Concurrency (10) | ~$15 + usage | Near zero | Fixed N, overflow cold starts | Low |
| SnapStart | $0 extra | 100-400ms | Unlimited (with fast restores) | Low |
I use SnapStart for the vast majority of my Java and Python Lambda functions now. The only functions where I still use provisioned concurrency are the ones that serve synchronous user-facing traffic with sub-50ms latency requirements, which is a small fraction of my total.
SnapStart is a great tool, but it's not universal. Here are the situations where I've chosen not to use it:
If your function is invoked thousands of times per second and already has high concurrency, SnapStart won't help much. The execution environments are already warm. The cold start problem is largely solved by the fact that your concurrency keeps all the init phases happening continuously. SnapStart won't hurt, but it won't be noticeable either.
If your function creates a new database connection on every invocation (or every few invocations) and doesn't persist anything during init, SnapStart has nothing to optimize. Your init phase is already near zero. The restore time will actually be slightly slower than a trivial init, since restoring a snapshot has some overhead.
As mentioned above, forked processes don't survive the snapshot/restore cycle. If your code depends on multiprocessing.Pool, subprocess with forking, or any other fork-based concurrency during the handler invocation, SnapStart will break it. You can work around this by moving fork operations out of the handler, but at that point the complexity cost might not be worth it.
If you rely on /tmp for caching downloaded model weights, compiled templates, or generated files across invocations, SnapStart changes the semantics. The snapshot captures /tmp at snapshot time, but any subsequent writes to /tmp during invocation are lost on restore. You can use /tmp as a scratch directory within a single invocation, but don't expect anything to persist.
Some libraries are incompatible with SnapStart because they generate code at runtime. Pydantic v2 with model_rebuild(), some JIT compilers, and certain bytecode manipulation frameworks can leave the snapshot in an inconsistent state. You'll catch this during the snapshot-taking phase (it will error out and the version won't be published), but it's worth testing early rather than discovering it in production.
SnapStart is the easiest cold start optimization you can enable for Java and Python Lambda functions. One block in your Terraform config, a version publish, and you're done. No keep-warm cron jobs, no provisioned concurrency spend, no custom runtime hacks.
Here's my current playbook:
@initDuration exceeds 500ms for SnapStart-enabled functions -- that usually means one of your restore hooks is too slow or the snapshot is failing and falling back to full init.SnapStart isn't a magic bullet. It doesn't fix cold starts for Node.js or Go. It doesn't make your init phase disappear -- it just moves it to deploy time. And it comes with constraints around connections, forked processes, and ephemeral storage that you need to understand.
But for the specific problem it solves -- Java and Python cold starts that take multiple seconds -- it's the best solution I've seen. No hacks, no workarounds, no extra infrastructure. Just a faster Lambda function with a single Terraform attribute.
I've been using SnapStart in production for about eight months across a dozen functions. My p99 latency dropped from 3.2 seconds to 340ms on the worst offender (a Spring Boot DynamoDB-backed API). The monthly Lambda bill didn't change. The deployment times went up by about 45 seconds per deploy, which I barely notice.
Try it on one function first. Run the CloudWatch query before and after. If the numbers look good, roll it out across the board. Your users will thank you.
I email when a new post goes up. One send a week, and only if there's something new.
Want this applied on your account? Start with Infrastructure as Code.
Related reading
Kept below the post instead of in a sidebar, with a slow continuous motion for a cleaner editorial feel.
On this post
Comments
A reply stays under the note it answers.
No comments yet.
If you have a note on Lambda SnapStart with Terraform - Cutting Cold Starts Without the Hacks, sign in and leave it.