Moving Past Vibe Coding: The Pre-Merge Checklist I Now Run on Every AI-Generated PR
Blog article/Blog archive

Moving Past Vibe Coding: The Pre-Merge Checklist I Now Run on Every AI-Generated PR

AI agents write code fast. They also write bugs, security holes, and nonsense. Here is the 12-item checklist I run before merging anything an agent touched.

Jul 7, 202613 min read0 comments43 views
AWSAIDevOpsTerraformCI/CD

I have been leaning hard into AI coding agents for the last few months. Claude Code, Copilot, the whole stack. They generate Terraform, Lambda handlers, CI workflows, and blog content. And honestly? They are shockingly productive.

But here is the thing nobody talks about when they tell you to vibe code. The same agent that writes a perfect IAM policy in 30 seconds will also quietly introduce a security hole that a human would never write. It will import a package that does not exist. It will suggest a Terraform resource that was deprecated three years ago. And it will do all of this with total confidence.

I learned this the hard way. An agent generated a PR that looked perfect. Tests passed. Linting clean. I merged it. The next morning, our staging environment was serving 403 errors on every Bedrock call because the agent had changed the IAM role trust policy in a way that looked right but broke everything.

That was the day I stopped treating agent-written code like human-written code. And the day I started building a proper pre-merge checklist specifically for AI-generated PRs.

This post is that checklist. Twelve items, concrete commands, and real examples from my portfolio repo. I run every single one before merging any agent-written change. You should too.

Developer reviewing code on screen with a checklist Photo by Jakub Zerdzicki on Pexels

Why Agent Code Needs a Different Review Process

Human developers make mistakes too. But the mistakes are different. A human who writes a broken IAM policy usually knows the policy is broken and will flag it in the PR description. An agent has no such self-awareness.

Agents have several failure modes that humans do not:

  1. Hallucinated dependencies. The agent imports a package that sounds plausible but does not exist. I have seen aws-sdk-bedrock-runtime-agent suggested in a PR. That package is not real. An agent made it up.
  2. Overly permissive IAM. Agents optimize for "does this work?" not "is this secure?" They tend to add Resource: "*" and Action: "*" because it unblocks the code. And they do it silently, in the middle of a 400-line diff.
  3. Confidence in wrong approaches. An agent will implement a deprecated API pattern with the same confidence it uses for a best-practice pattern. It does not know the difference.
  4. Silent regressions. The agent changes a line that looks unrelated. The change passes tests. But it breaks something subtle in production. I had an agent change the timeout on a Lambda function from 300 seconds to 30 seconds because "30 is a round number." The function handles long Bedrock calls. It started failing silently.

These failures have a pattern. They are invisible to standard CI checks. Lint, typecheck, and unit tests pass because the code is syntactically valid. The problem is semantic.

So here is the checklist I use. Twelve items, arranged by priority. I run them in order, and I do not merge until all twelve pass.

Robot and human comparing code side by side Photo by cottonbro studio on Pexels

Checklist Item 1: Terraform Plan Review (the Hard Way)

The first thing I check on any agent-written PR that touches infrastructure is the Terraform plan output. Not the .tf file diff. The actual terraform plan output.

The diff lies. The plan tells the truth.

I run this command:

cd infra/
terraform plan -no-color | grep -E '(\+|\\-|~|#)' | head -80

What I look for:

Any Resource: "*" in IAM policies that was not there before. Agents love adding wildcards. I reject any PR that introduces a new wildcard IAM statement unless there is a documented exception.

Any resource deletion. Agents sometimes delete resources they think are "unused" without asking. I have caught an agent trying to delete a DynamoDB table that was referenced in three other repos.

Any count or for_each changes. Agents frequently change resource counts without understanding the blast radius.

Real example: An agent added a new IAM policy to allow Lambda to write to CloudWatch Logs. The policy looked reasonable. But the plan showed it also removed the existing Deny effect on deleting log groups. Merged together, the change would have allowed the Lambda function to delete log groups. The plan caught it. The PR did not.

Terraform infrastructure code displayed on a laptop Photo by Daniil Komov on Pexels

Checklist Item 2: IAM Diff with iam-lint

Terraform plan catches what changed. But it does not tell you whether the change is safe. For that, I use a simple script I wrote called hrr_iam_lint. It compares every IAM statement before and after the agent's change and flags patterns I never want to see.

The script checks for these patterns and marks them as FAIL:

  • Action: "*" combined with Resource: "*"
  • Action: "iam:*" on a resource that is not the account root
  • Effect: "Allow" on sts:AssumeRole without a Condition block
  • A policy attached to a user instead of a role

I run it like this:

./scripts/hrr_iam_lint.sh infra/policies/

If any check returns FAIL, the PR goes back to drafting. No exceptions.

The most common agent mistake I catch with this is attaching policies directly to IAM users instead of roles. Agents seem to prefer the simpler path even though roles are the AWS best practice.

Security shield protecting infrastructure code Photo by RealToughCandy.com on Pexels

Checklist Item 3: Infrastructure Security Scan with Checkov

I run Checkov on every agent-written PR. Not just on the changed files, but on the entire Terraform directory. Agents sometimes add insecure defaults to unrelated resources.

checkov -d infra/ --quiet --compact \
  --skip-check CKV_AWS_272,CKV_AWS_273 \
  | grep -E '(FAIL|ERROR)' | head -30

I skip CKV_AWS_272 and CKV_AWS_273 because those are false positives for my setup. Everything else gets reviewed.

What I look for:

Any high-severity finding that the agent introduced. Medium and low findings are acceptable if there is a documented reason. High findings block the merge.

The agent that tried to deploy an S3 bucket without encryption is a classic. Checkov flagged it immediately. The agent had written the bucket resource and simply forgot the server_side_encryption_configuration block. A human would probably have caught it too. But the agent did it because it was following a minimal example from a tutorial.

Checkov catches these automatically. I do not even need to think about it.

Security scan results showing pass and fail checks Photo by Tima Miroshnichenko on Pexels

Checklist Item 4: Dependency Audit

Agents add dependencies liberally. They import packages, add npm modules, and reference GitHub Actions that you have never heard of. I audit every new dependency that an agent introduces.

For Node.js:

npm audit --json | jq '.vulnerabilities | keys'

For Python:

pip-audit --desc | grep -E '(HIGH|CRITICAL)'

For GitHub Actions:

grep -r 'uses:' .github/workflows/ | grep -v '@v[0-9]' | grep -v '/'

But the audit goes beyond security. I also check:

Did the agent add a package that already exists in the project? Agents frequently add a second HTTP client library because they did not check package.json first. I have seen axios, got, node-fetch, and undici all imported in the same PR by an agent that kept rewriting the same function.

Did the agent pin a version or use a range? I reject any PR that introduces an unpinned dependency. ^1.0.0 is not acceptable for production infrastructure code.

Did the agent import a package that does not exist yet? This is harder to catch. I run the build and check for npm warnings. But the real defense is knowing that agents hallucinate package names and always running the full dependency install before review.

Network visualization of package dependencies Photo by U.Lucas Dube-Cantin on Pexels

Checklist Item 5: AI Detection on Doc Strings

This one is specific to AI-generated PRs. Agents write code with verbose, overly formal doc strings. They sound like a textbook. And the problem is not aesthetic. Overly formal doc strings make the codebase harder to scan because every block of text looks the same.

I run a simple check on any new or modified doc strings in the PR:

grep -r 'Parameters:|Returns:|Raises:' src/ | grep -v '__pycache__'

If I see the Google-style doc string format (Parameters, Returns, Raises on their own lines) I know an agent wrote it. I flag it for a human rewrite.

This is not just style. The verbose doc strings trigger AI detection tools when the codebase is used as context for other agents. I want the codebase to stay clean and human-sounding.

The rule is simple: doc strings should explain why not what. The code explains what. If the agent wrote a doc string that reads like a Java tutorial, I replace it with a one-liner.

Chatbot interface alongside a code editor Photo by Daniil Komov on Pexels

Checklist Item 6: Eval Harness for LLM-Touching Code

If the PR changes any code that makes LLM calls (Bedrock, OpenAI, Anthropic), I run the eval harness. This is not optional.

My eval harness is simple. It takes a set of known inputs, runs them through the new code, and checks that the outputs match expected values. The inputs are real customer requests that have been sanitized.

npx tsx src/evals/run.ts --changed-only \
  --diff-baseline main \
  --output /tmp/eval-report.json

What I look for:

Any regression in output quality. An agent might change a prompt to make it faster. That is good. But it might also change the prompt in a way that loses accuracy for edge cases. The eval harness catches this.

Any increase in token usage. Agents sometimes add extra context to prompts without understanding the cost impact. A 20 percent token increase means a 20 percent cost increase. That needs a documented reason.

The most common eval failure I see is an agent "optimizing" a prompt by removing instructions it considers redundant. The eval then shows a 15 percent drop in accuracy for the most common input type. The agent saved 50 tokens per call and lost accuracy. Not a good trade.

AI neural network evaluation dashboard showing metrics Photo by Google DeepMind on Pexels

Checklist Item 7: Manual Smoke Test Against Staging

CI tests run. CI tests pass. But CI tests only cover what you thought to test. The agent might have broken something that no test covers.

I run one manual smoke test on staging before any agent-written PR gets merged:

  1. Deploy the agent's branch to staging
  2. Hit the main user-facing endpoint
  3. Check that the response is valid
  4. Check CloudWatch logs for unexpected errors
  5. Check the Lambda cold start and duration metrics

The whole thing takes five minutes. But it catches things that automated tests miss.

I once had an agent change the content-type header in an API Gateway response from application/json to text/plain. All automated tests passed because they checked the status code and the body structure, not the content type. The manual smoke test caught it immediately because the frontend displayed raw JSON instead of parsed data.

The smoke test is fast. It is not comprehensive. But it is the last line of defense before the code hits production.

Server deployment dashboard with health metrics Photo by Atlantic Ambience on Pexels

Checklist Item 8: Diff Size Check

This is the simplest check on the list. If an agent-written PR changes more than 20 files, I do not merge it until it is broken into smaller PRs.

Agents have no sense of scope. They will refactor three unrelated systems in the same PR because the prompt said "make this better." And they do this without asking whether each change is in scope.

I run:

git diff main HEAD --stat | tail -1

If the line count is over 500 or the file count is over 20, the PR goes back for splitting. No exceptions.

The reason is not just reviewability. It is safety. A 500-line agent-written diff has a much higher defect density than a 50-line one. I have seen this empirically. The first 100 lines are usually solid. The next 200 are a mixed bag. Everything after line 300 is where the hallucinations start.

Keep agent PRs small. Review them carefully. Deploy them fast.

Git merge conflict resolution interface Photo by Daniil Komov on Pexels

Checklist Item 9: Structured Logging and Telemetry

Agents write code, but they rarely add logging. If the PR adds a new Lambda function or API endpoint, I check whether it includes structured logging.

I look for:

console.log(JSON.stringify({ event: "lambda_start", requestId, timestamp }))

Not:

console.log("Function started")

The structured log entry is 30 characters longer. It saves hours of debugging because you can query it in CloudWatch Logs Insights.

I also check that the agent added a metric for the new code path. A simple mono_metric call for latency, error count, and invocation count is enough. Without it, you are flying blind on the new code.

The agent that added a new API endpoint without any logging at all is my favorite example. The endpoint worked. But when it started failing under load, there was no way to tell why. CloudWatch showed errors but no context. I added the logging myself, found the bug, fixed it, and then added this check to the list.

Server logs and monitoring metrics on a dashboard Photo by Tima Miroshnichenko on Pexels

Checklist Item 10: Hardcoded Secrets Scan

Agents love hardcoding API keys. Not because they think it is a good practice. Because the example in their training data hardcodes the key.

I run Gitleaks on every agent-written PR:

gitleaks detect --source . --verbose --redact | grep -E '(\+|\\-)'

If Gitleaks finds anything, the PR is blocked. But I also look for a subtler pattern: the agent adding a .env.example file with real-looking placeholder values.

I had an agent add a .env.example with AWS_ACCESS_KEY_ID=AKIAIOSFODNN7EXAMPLE. That is a well-known test key, and Gitleaks did not flag it because it is intentionally public. But the agent also added BEDROCK_API_KEY=sk-test-abcdef123456. Not a real key. But close enough to a real key format that a developer might accidentally use it in staging.

The rule: any file that contains what looks like a credential gets reviewed by a human. No exceptions.

Cybersecurity concept with lock and key Photo by Nathan Thomas on Pexels

Checklist Item 11: Human Code Review (This Is Not Optional)

None of the automated checks replace a human reading the diff. The twelve items above are filters. They catch the common failures. But a human reviewer catches the uncommon ones.

I read the entire diff of every agent-written PR. Not a skim. A line-by-line read. And I look for things that the automated checks cannot catch:

Does this change make the codebase harder to understand? An agent might introduce a clever abstraction that saves three lines but makes the flow incomprehensible.

Does this change introduce coupling that did not exist before? An agent might wire two services together in a way that looks efficient but creates a dependency cycle.

Is the error handling consistent with the rest of the codebase? Agents handle errors differently in every function. I check that new code uses the same error patterns as the existing code.

The human review takes the longest. It is also the most valuable check on this list. Do not skip it just because the automated checks passed.

Two developers reviewing code on large monitors Photo by ThisIsEngineering on Pexels

Checklist Item 12: Deploy to Staging First, Watch for 30 Minutes

The last check happens after the merge. I deploy to staging. And then I watch.

For 30 minutes I check:

  • Error rate in CloudWatch (should be zero or near zero)
  • P99 latency (should not increase by more than 10 percent)
  • Lambda invocation count (should match expected traffic)
  • Any new log entries at ERROR level
  • Any unexpected cost changes in the staging billing dashboard

If any of these metrics look wrong, I roll back. Not investigate first. Roll back first, investigate second. The staging environment has real data and the rollback takes two minutes.

This saved me once already. An agent-optimized Lambda function reduced cold start time by 40 percent. Great. But it also introduced a memory leak that showed up after 15 minutes of sustained traffic. The 30-minute watch caught it before production users were affected. I rolled back, fixed the leak, and deployed again.

The 30-minute watch is the cheapest insurance you will ever buy.

CloudWatch monitoring dashboard showing cloud metrics Photo by RDNE Stock project on Pexels

The Checklist in Practice

Here is how I run the entire thing:

# 1. Terraform plan
cd infra && terraform plan -no-color | grep -E '(\+|\\-|~|#)' | head -80

# 2. IAM lint
./scripts/hrr_iam_lint.sh infra/policies/

# 3. Security scan
checkov -d infra/ --quiet --compact | grep -E '(FAIL|ERROR)' | head -30

# 4. Dependency audit
npm audit --json | jq '.vulnerabilities | keys'
grep -r 'uses:' .github/workflows/ | grep -v '@v[0-9]'

# 5. Doc string check
grep -r 'Parameters:|Returns:|Raises:' src/

# 6. Eval harness
npx tsx src/evals/run.ts --changed-only --diff-baseline main

# 7. Manual smoke test
# Deploy to staging, hit endpoint, check logs

# 8. Diff size
git diff main HEAD --stat | tail -1

# 9. Structured logging
grep -r 'console.log(' src/ --include='*.{ts,js,py}' | grep -v 'JSON.stringify'

# 10. Secrets scan
gitleaks detect --source . --verbose --redact | grep -E '(\+|\\-)'

# 11. Human code review
git diff main HEAD | less

# 12. Deploy to staging, watch 30 min

The whole thing takes me about 45 minutes for a medium PR. That sounds like a lot. But it has saved me from at least a dozen production incidents in the last three months.

You do not need to do all twelve every time. Start with items 1, 2, 3, and 10. Those catch the most common and most dangerous failures. Add the rest as you find gaps in your process.

The key insight is simple. AI coding agents are incredible tools. But they are not human engineers. They need a different review process. One that accounts for their specific failure modes: wildcard IAM, hallucinated dependencies, overly permissive defaults, and silent regressions.

Build the checklist. Run it every time. Your production environment will thank you.

Hope you enjoyed this. If you have a checklist item I missed, I would love to hear about it. Find me on X at https://x.com/harundotdev and let me know.

Developer typing a checklist on a laptop Photo by cottonbro studio on Pexels

Get the next note

I email when a new post goes up. One send a week, and only if there's something new.

Want this applied on your account? Start with Infrastructure as Code.

Related reading

More posts, sliding underneath the article

Kept below the post instead of in a sidebar, with a slow continuous motion for a cleaner editorial feel.

On this post

Comments

A reply stays under the note it answers.

0 comments

No comments yet.

If you have a note on Moving Past Vibe Coding: The Pre-Merge Checklist I Now Run on Every AI-Generated PR, sign in and leave it.