AWS WAF + CloudFront for a Solo Dev - Rate Limiting, Bot Control, and IP Blocking
Blog article/Blog archive

AWS WAF + CloudFront for a Solo Dev - Rate Limiting, Bot Control, and IP Blocking

How to set up AWS WAF in front of CloudFront for a solo-built application, with rate limiting, bot control, and IP block rules, all managed with Terraform.

Jun 16, 202610 min read0 comments35 views
AWSWAFCloudFrontSecurityTerraform

I run a handful of personal projects behind CloudFront. For a long time I thought WAF was overkill for a solo dev. A waste of money and complexity. Then I watched a single scraper hit my API 50,000 times in an afternoon and spike my Lambda costs by $40. That changed my mind.

AWS WAF in front of CloudFront costs about $6 per month for the web ACL plus whatever you add in rule groups. For that you get rate limiting, IP blocking, bot detection, and managed threat protection. Compared to the cost of one bad day with an unchecked scraper, it pays for itself in the first hour.

This post walks through what I actually run. Real Terraform configs, real thresholds, real gotchas. No academic architecture diagrams. Just what works for a single developer who wants protection without a security team.

Why WAF for a Solo Dev

The standard argument against WAF at small scale is that you can handle bad traffic at the application layer. Write a rate limiter middleware, parse user-agent headers, maintain an IP blocklist in Redis. I have done all of that. It works until it does not.

Application-level rate limiting runs inside your compute. Every request that gets rate limited still consumed CPU cycles, database connections, and memory. It still counts toward your Lambda invocation costs. WAF stops requests before they reach your application. The request never touches your origin. You pay nothing for compute or data transfer on the requests you blocked.

The other reason is maintenance. I do not want to deploy code changes every time I need to block a range of IPs or update a rate limit threshold. WAF rules update in minutes. No deployment pipeline. No code review. Just a Terraform apply or a console change.

The Terraform Module

Here is the full module I use. It creates a WAFv2 web ACL attached to a CloudFront distribution, with rate limiting, managed rule groups, bot control, and an IP set for manual blocks.

I keep this in modules/waf-cloudfront/main.tf and reference it from my root configuration.

# modules/waf-cloudfront/main.tf

locals {
  name_prefix = var.name_prefix != "" ? var.name_prefix : "waf"
}

resource "aws_wafv2_ip_set" "blocked_ips" {
  count = length(var.blocked_ip_addresses) > 0 ? 1 : 0

  name        = "${local.name_prefix}-blocked-ips"
  description = "IP addresses to block"
  scope       = "CLOUDFRONT"
  ip_address_version = "IPV4"
  addresses   = var.blocked_ip_addresses
}

resource "aws_wafv2_web_acl" "main" {
  name        = "${local.name_prefix}-web-acl"
  description = "Web ACL for CloudFront protection"
  scope       = "CLOUDFRONT"

  default_action {
    allow {}
  }

  # ---------- Rate limiting ----------

  rule {
    name     = "rate-limit"
    priority = 0

    action {
      block {}
    }

    statement {
      rate_based_statement {
        limit              = var.rate_limit
        aggregate_key_type = "IP"
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "${local.name_prefix}RateLimit"
      sampled_requests_enabled   = true
    }
  }

  # ---------- AWS Managed Rule Groups (Baseline) ----------

  rule {
    name     = "aws-managed-baseline"
    priority = 10

    override_action {
      none {}
    }

    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesCommonRuleSet"
        vendor_name = "AWS"

        rule_action_override {
          name = "NoUserAgent_HEADER"
          action_to_use {
            block {}
          }
        }
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "${local.name_prefix}Baseline"
      sampled_requests_enabled   = true
    }
  }

  # ---------- SQL Injection ----------

  rule {
    name     = "aws-managed-sqli"
    priority = 20

    override_action {
      none {}
    }

    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesSQLiRuleSet"
        vendor_name = "AWS"
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "${local.name_prefix}Sqli"
      sampled_requests_enabled   = true
    }
  }

  # ---------- XSS ----------

  rule {
    name     = "aws-managed-xss"
    priority = 30

    override_action {
      none {}
    }

    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesKnownBadInputsRuleSet"
        vendor_name = "AWS"
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "${local.name_prefix}Xss"
      sampled_requests_enabled   = true
    }
  }

  # ---------- Bot Control ----------

  rule {
    name     = "bot-control"
    priority = 40

    override_action {
      none {}
    }

    statement {
      managed_rule_group_statement {
        name        = "AWSManagedRulesBotControlRuleSet"
        vendor_name = "AWS"
        version     = "WAFBotControlLatest"

        managed_rule_group_configs {
          aws_managed_rules_bot_control_rule_set {
            inspection_level = "COMMON"
          }
        }
      }
    }

    visibility_config {
      cloudwatch_metrics_enabled = true
      metric_name                = "${local.name_prefix}BotControl"
      sampled_requests_enabled   = true
    }
  }

  # ---------- IP Block List ----------

  dynamic "rule" {
    for_each = length(var.blocked_ip_addresses) > 0 ? [1] : []
    content {
      name     = "block-ip-set"
      priority = 50

      action {
        block {}
      }

      statement {
        ip_set_reference_statement {
          arn = aws_wafv2_ip_set.blocked_ips[0].arn
        }
      }

      visibility_config {
        cloudwatch_metrics_enabled = true
        metric_name                = "${local.name_prefix}BlockedIps"
        sampled_requests_enabled   = true
      }
    }
  }

  visibility_config {
    cloudwatch_metrics_enabled = true
    metric_name                = "${local.name_prefix}WebAcl"
    sampled_requests_enabled   = true
  }
}

resource "aws_wafv2_web_acl_association" "cloudfront" {
  resource_arn = var.cloudfront_arn
  web_acl_arn  = aws_wafv2_web_acl.main.arn
}

# ---------- CloudWatch Alarms ----------

resource "aws_cloudwatch_metric_alarm" "blocked_requests_high" {
  alarm_name          = "${local.name_prefix}-blocked-requests-high"
  comparison_operator = "GreaterThanThreshold"
  evaluation_periods  = 2
  metric_name         = "BlockedRequests"
  namespace           = "AWS/WAFV2"
  period              = 300
  statistic           = "Sum"
  threshold           = var.alarm_blocked_threshold
  alarm_description   = "Alert when blocked requests exceed threshold in 10 minutes"
  treat_missing_data  = "notBreaching"

  dimensions = {
    Rule   = "ALL"
    Region = "global"
    WebACL = aws_wafv2_web_acl.main.name
  }

  alarm_actions = [var.sns_topic_arn]
}
# modules/waf-cloudfront/variables.tf

variable "name_prefix" {
  description = "Prefix for resource names"
  type        = string
  default     = "waf"
}

variable "cloudfront_arn" {
  description = "ARN of the CloudFront distribution to associate"
  type        = string
}

variable "rate_limit" {
  description = "Maximum requests per 5-minute window per IP"
  type        = number
  default     = 2000
}

variable "blocked_ip_addresses" {
  description = "List of IP addresses or CIDR blocks to block"
  type        = list(string)
  default     = []
}

variable "sns_topic_arn" {
  description = "ARN of SNS topic for alarm notifications"
  type        = string
  default     = ""
}

variable "alarm_blocked_threshold" {
  description = "Threshold for blocked requests alarm (sum over 5 minutes)"
  type        = number
  default     = 1000
}

Rate Limiting: How to Set Thresholds

Rate limiting is the single most useful rule for a solo dev. It protects your origin from brute force attacks, scraping, and accidental runaway clients.

The rate limit in the module above is set per IP over a 5-minute rolling window. A threshold of 2,000 requests per 5 minutes works out to about 6.6 requests per second per IP. That is generous enough for normal API usage but aggressive enough to stop most abuse.

I set my threshold by looking at CloudFront access logs for the previous week. I took the 95th percentile of requests per IP per 5 minutes and doubled it. That gave me a number that would never trip on legitimate traffic but catches spikes early.

For a content site (not an API), you can set it higher. For a login endpoint, set it lower. You can have multiple rate limit rules with different thresholds targeting different URI paths.

# Example: stricter rate limit on login endpoint
rule {
  name     = "rate-limit-login"
  priority = 1

  action {
    block {}
  }

  statement {
    rate_based_statement {
      limit              = 100
      aggregate_key_type = "IP"

      scope_down_statement {
        byte_match_statement {
          field_to_match {
            uri_path {}
          }
          positional_constraint = "STARTS_WITH"
          search_string        = "/login"
        }
      }
    }
  }

  visibility_config {
    cloudwatch_metrics_enabled = true
    metric_name                = "RateLimitLogin"
    sampled_requests_enabled   = true
  }
}

The scope_down_statement limits the rate counter to only requests matching a pattern. In this case, only requests to paths starting with /login count toward the 100 request per 5 minute limit per IP. This is how you protect authentication endpoints without penalizing the rest of your site.

AWS Managed Rule Groups

The managed rule groups are pre-built rule sets maintained by the AWS security team. They update automatically as new threats emerge. You do not need to follow CVE announcements or update regex patterns.

I use three of them:

AWSManagedRulesCommonRuleSet covers the OWASP top 10 basics. Known bad headers, path traversal, RFI, SSRF. I overrode the NoUserAgent_HEADER rule to block instead of count. Most legitimate clients send a user-agent header. If a request lacks one, I do not want it reaching my application.

AWSManagedRulesSQLiRuleSet catches SQL injection attempts in query parameters, body, and URI paths. If you have a database-backed application, this is non-negotiable.

AWSManagedRulesKnownBadInputsRuleSet covers XSS and other input-based attacks. It checks for common XSS payload patterns in request parameters and bodies.

All three run in count mode by default in my config, which means they log matches but do not block. I let them run for a week, review the logs, whitelist any false positives, then switch to block mode. For the config above I left them as override_action { none {} } which means they use the managed group's default action (usually count). If you want to force block, change it to:

override_action {
  count {}
}

Wait. That is count, not block. To override to block, you do not set an override_action at all and set the rule action. Actually, the correct pattern for forcing managed rules to block is more nuanced. The managed rule group has its own rule actions. Set override_action { none {} } and then use rule_action_override on individual rules within the group if you want to change specific ones. For most solo projects, the default actions work fine. The managed groups will block the most severe detections by default and count the rest.

Bot Control

Bot Control is a managed rule group that classifies bots into categories: verified (Googlebot, Bingbot), unverified (scrapers, crawlers), and malicious (bad bots, credential stuffers).

It costs about $10 per month on top of the base WAF ACL cost. Worth it if you run a public API or a content site that attracts scrapers.

What it catches:

  • Scrapers that fake browser user-agent strings but behave differently at the HTTP level (no JavaScript execution, predictable request timing, no Accept-Language headers).
  • Headless browsers used for automated form submissions.
  • HTTP libraries (requests, curl, httpx) that are sometimes used for legitimate purposes but also for abuse.

The inspection level in my config is COMMON which is the cheaper and less aggressive option. TARGETED is more thorough but costs more and has a higher chance of false positives. Start with COMMON.

Bot Control in count mode gives you a label on each request. You can see in CloudWatch which requests were classified as bots without blocking anything. I ran it in count mode for two weeks before enabling blocking.

IP Set Rules for Blocking Known Bad Actors

Some IPs and ASNs abuse your application repeatedly. A single IP hammering your API from a datacenter. A range of IPs from a hosting provider you do not do business with. These belong in an IP set.

The module above creates a aws_wafv2_ip_set resource and a rule that blocks all traffic from those IPs. I update the list manually when I see abuse in the logs.

You can also subscribe to threat intelligence feeds that publish lists of known bad IPs. AbuseIPDB, Spamhaus, and AlienVault OTX all have free or low-cost feeds. With a scheduled Lambda function you can pull these feeds, parse them, and update your IP set automatically. That is a separate project, but the WAF side of it is straightforward once you have the IP set resource configured.

CloudWatch Metrics and Alarms

WAF publishes metrics to CloudWatch automatically. The important ones:

  • AllowedRequests - requests that passed all rules
  • BlockedRequests - requests that were blocked
  • CountedRequests - requests that matched a rule but were not blocked (count mode)
  • PassedRequests - requests that did not match any rule

I set an alarm on BlockedRequests to notify me when blocked traffic spikes. If I see a sudden increase, I investigate. Sometimes it is a legitimate change in traffic patterns. Sometimes it is an attack that my rate limit is catching.

The alarm in the module triggers when the sum of blocked requests exceeds a threshold over two consecutive 5-minute periods. I set mine to 1,000 blocked requests per 5 minutes. You will want to tune this based on your normal traffic volume.

Testing: curl and Artillery

Before you put a WAF in production, test that it works the way you expect. Here is how I test mine.

Rate limiting test with curl:

# Run this in a loop to trigger rate limiting
for i in $(seq 1 100); do
  curl -s -o /dev/null -w "%{http_code}\n" https://your-domain.com/
done

After hitting the rate limit threshold, you should see 503 responses with a waf block reason. The 503 is WAF's way of saying "I blocked this, but the origin never saw it."

Load test with Artillery:

npm install -g artillery

cat > load-test.yaml <<EOF
config:
  target: "https://your-domain.com"
  phases:
    - duration: 60
      arrivalRate: 10
      name: "Warm up"
    - duration: 60
      arrivalRate: 50
      name: "Sustained load"
  defaults:
    headers:
      User-Agent: "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7)"

scenarios:
  - flow:
      - get:
          url: "/"
EOF

artillery run load-test.yaml

Watch the CloudWatch metrics during the test. You should see BlockedRequests increase when the rate limit rule activates. If you see 200s throughout, your threshold is too high. If you see 503s during the warm-up phase, your threshold is too low.

SQL injection test:

curl "https://your-domain.com/?id=1%27%20OR%20%271%27%3D%271"

This should return a 403 or 503 depending on which managed rule catches it. The SQLi rule group will block this request.

Bot test with a raw HTTP client:

curl -A "" https://your-domain.com/

With the NoUserAgent_HEADER override I set above, this request gets blocked. You should see a 403 response.

Gotchas

WAF Regional vs. CloudFront Global

This is the most important thing to understand. WAF can be deployed in two scopes: REGIONAL and CLOUDFRONT.

When you attach WAF to CloudFront, the scope must be CLOUDFRONT. This deploys the WAF to AWS's edge locations globally. The IP set for a CloudFront-scoped WAF must use IPv4 or IPv6 addresses, not ranges that overlap with AWS internal ranges.

If you attach WAF to an Application Load Balancer, API Gateway, or AppSync, the scope is REGIONAL. The rules are evaluated in that region only.

You cannot reuse a regional WAF ACL with CloudFront, and you cannot reuse a CloudFront-scoped WAF ACL with anything else. They are separate resource types in the API even though they look identical in the console.

In Terraform, the scope attribute on aws_wafv2_web_acl controls this. CloudFront requires scope = "CLOUDFRONT". ALB requires scope = "REGIONAL".

60-Second Propagation Delay

When you create or update a WAF ACL, the changes propagate to edge locations within about 60 seconds. During that propagation window, old rules may still be in effect at some edge locations while new rules are active at others.

This is not usually a problem. The propagation is eventually consistent. If you are blocking an IP that is actively attacking you, the 60-second delay means they might get a few more requests through before the block applies everywhere. That is fine.

What you should not do is make a change, test it immediately from a single location, and assume the result represents global behavior. Give it a minute.

CloudFront WAF vs. ALB WAF Pricing Differences

CloudFront-scoped WAF does not charge for data transfer on blocked requests because the request never reaches an AWS region. Regional WAF on an ALB still processes the request through the load balancer before WAF evaluates it, so you pay for the ALB processing even for blocked requests.

For a solo dev serving traffic globally, CloudFront + WAF is both cheaper and more effective than ALB + WAF. You get edge caching, DDoS protection (AWS Shield Standard is included with CloudFront), and WAF blocking at the edge. Three services for the price of one plus WAF.

Rate Limit Is Per Edge Location

The rate limit counter is per individual CloudFront edge location, not global. If your traffic is distributed across 50 edge locations, a single IP could send up to 50 times your rate limit before getting blocked.

This is a known limitation. AWS documents it explicitly. If you need strict global rate limits, you need to implement them at the application layer or use CloudFront origin-facing rate limiting with Lambda@Edge. For most solo projects, the per-edge-location rate limit is good enough to stop abuse without requiring global precision.

What This Costs

Here is the real cost breakdown for a single web ACL with the rules above, serving fewer than 10 million requests per month:

  • Web ACL: $5.00 per month (prorated, ~$0.007 per hour)
  • Rule group usage: free for AWS managed rules (you pay per rule over your first 10, but managed group rules in the first 10 count differently - the pricing page is confusing, so check the current rates)
  • Bot Control: $10.00 per month if enabled
  • IP Set: free (included in ACL price)
  • CloudWatch metrics and alarms: a few cents

Total: around $6 per month without Bot Control, $16 with it. If you are running a Lambda-based API that costs $20-50 per month in compute, adding WAF might increase your AWS bill by 10-30%. For me, the peace of mind and the prevention of one bad incident per year makes it worth it.

Usage

Here is how I call the module from my root configuration:

# locals.tf
locals {
  cloudfront_arn = "arn:aws:cloudfront::123456789012:distribution/ABCDEF1234"
  domain         = "myapp.com"
}

# waf.tf
module "waf" {
  source = "./modules/waf-cloudfront"

  name_prefix          = "myapp"
  cloudfront_arn       = local.cloudfront_arn
  rate_limit           = 2000
  blocked_ip_addresses = [
    "192.0.2.0/24",
    "198.51.100.10/32",
  ]
  sns_topic_arn        = aws_sns_topic.alerts.arn
}

That is the entire setup. One terraform apply and your CloudFront distribution has rate limiting, SQL injection protection, XSS protection, bot detection, and an IP block list.

Wrapping Up

WAF is not just for enterprise teams with dedicated security engineers. For a solo developer running a production application on CloudFront, the setup cost is one afternoon of Terraform work and the ongoing cost is less than a coffee subscription.

The module in this post is what I run in production. It has stopped scrapers, blocked credential stuffing attempts, and absorbed traffic spikes that would have cost me real money in Lambda execution time. You can copy it directly, adjust the rate limit to your traffic patterns, and deploy it in one Terraform apply.

One thing I left out: the WAF logs. You should enable logging to S3 or CloudWatch Logs to see what is being blocked and tune your rules. That is a topic on its own, but the short version is you pass a logging_configuration block to the web ACL resource and point it at an S3 bucket or a CloudWatch log group. Do not skip this. Without logs, you are flying blind.

Get the next note

I email when a new post goes up. One send a week, and only if there's something new.

Want this applied on your account? Start with Infrastructure as Code.

Related reading

More posts, sliding underneath the article

Kept below the post instead of in a sidebar, with a slow continuous motion for a cleaner editorial feel.

On this post

Comments

A reply stays under the note it answers.

0 comments

No comments yet.

If you have a note on AWS WAF + CloudFront for a Solo Dev - Rate Limiting, Bot Control, and IP Blocking, sign in and leave it.