Building Bedrock Agents in Production - Knowledge Bases, Action Groups, and Guardrails
Blog article/Blog archive

Building Bedrock Agents in Production - Knowledge Bases, Action Groups, and Guardrails

How to build a production-ready Bedrock Agent with S3 knowledge bases, Lambda action groups, and guardrails, all configured with Terraform.

Jun 18, 202611 min read0 comments51 views
AWSBedrockAILambdaTerraform

I have been building AI-powered features on AWS for the last two years. Most of them follow the same pattern. A Lambda function takes a prompt, calls a Bedrock model, and returns a response. It works. But there is a ceiling on what you can do with a single model call and a static prompt.

A Bedrock Agent is different from a Lambda function that calls Bedrock. When you call bedrock:InvokeModel from a Lambda function, you get exactly one response. No tool use. No document retrieval. No multi-step reasoning. The agent wraps orchestration around the model. It can call knowledge bases for context, invoke Lambda functions as tools, follow guardrails, and chain multiple steps together based on what the model decides to do next.

I have built three production Bedrock Agents so far. One for internal customer support, one for automated incident response, and one for document summarization across a large S3 corpus. Each one taught me something about where the complexity really lives. It is not in the agent itself. It is in the surrounding pieces. Knowledge base ingestion, action group schemas, IAM policies, and guardrails.

This post covers everything I had to figure out to put a Bedrock Agent into production with Terraform. S3 knowledge bases with proper chunking, Lambda action groups with OpenAPI schemas the agent can actually parse, guardrails that block bad inputs and outputs, IAM policies that are least-privilege without being too restrictive, and the cost and latency tradeoffs you need to plan for.

What a Bedrock Agent Actually Is

The term "agent" gets thrown around loosely. In Bedrock, an agent is a managed runtime that exposes an orchestrated AI workflow. You define what the agent knows and what it can do, and the agent figures out how to combine those capabilities to answer a user's request.

Here is what happens when a user sends a message to a Bedrock Agent:

  1. Input hits the agent via the Bedrock Agent Runtime API (or the console).
  2. Pre-processing runs the input through guardrails. Content filters, topic policies, PII redaction. Anything blocked stops here.
  3. Orchestration sends the (filtered) input to the foundation model with the agent's system prompt and available tools.
  4. Tool calls happen when the model decides it needs more context or wants to perform an action. The agent routes to knowledge bases for retrieval or Lambda functions for action execution.
  5. Post-processing runs the model output through guardrails again before returning to the user.

Each step is configurable. You control which knowledge bases the agent can query, which Lambda functions it can invoke, what guardrails apply at each stage, and how many retries the agent attempts before giving up.

This is the architecture you are building toward:

User --> Bedrock Agent Runtime API
              |
        [Guardrails]
              |
        [Foundation Model]
           /          \
  [Knowledge Base]  [Action Groups]
     (S3 + RDS)      (Lambda)

The agent decides when to query the knowledge base, when to call a Lambda function, and how to combine results. You do not hardcode a workflow. You define the capabilities and let the model route between them.

Setting Up an S3 Knowledge Base

The knowledge base is where your agent looks for context. Bedrock Agents support two types of knowledge bases. S3-based (documents stored as files in S3) and RDS-based (vectors stored in a PostgreSQL or Aurora database). For most use cases I start with S3 because it is the simplest ingestion pipeline.

Here is what a knowledge base needs:

  • A source S3 bucket containing your documents.
  • A vector index (Bedrock manages this for you in its OpenSearch Serverless collection, or you can use your own).
  • A data source configuration that defines how documents get parsed and chunked.
  • An ingestion job that reads documents from S3, chunks them, generates embeddings, and stores them in the vector index.

The chunking strategy matters more than most tutorials admit. Bedrock supports two chunking strategies: fixed-size and hierarchical. Fixed-size is simpler. You split documents into chunks of N tokens with overlap. Hierarchical breaks documents into sections first, then chunks within sections. The second option preserves more contextual boundaries.

For knowledge base documents, use hierarchical chunking if your source files have structure (headings, sections, paragraphs). Use fixed-size if you have unstructured text. I have seen better retrieval quality with hierarchical for anything that resembles documentation.

# knowledge_base.tf

resource "aws_bedrockagent_knowledge_base" "hrr_docs" {
  name     = "hrr-docs-knowledge-base"
  role_arn = aws_iam_role.hrr_bedrock_agent.arn

  knowledge_base_configuration {
    type = "VECTOR"

    vector_knowledge_base_configuration {
      embedding_model_arn = "arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0"
    }
  }

  storage_configuration {
    type = "OPENSEARCH_SERVERLESS"

    opensearch_serverless_configuration {
      collection_arn    = aws_opensearchserverless_collection.hrr_kb.arn
      vector_index_name = "hrr-bedrock-kb-index"
      field_mapping {
        metadata_field = "AMAZON_BEDROCK_METADATA"
        text_field     = "AMAZON_BEDROCK_TEXT_CHUNK"
      }
    }
  }
}

resource "aws_bedrockagent_data_source" "hrr_docs_source" {
  name                 = "hrr-prod-docs"
  knowledge_base_id    = aws_bedrockagent_knowledge_base.hrr_docs.id
  data_source_type     = "S3"

  data_deletion_policy = "RETAIN"

  s3_configuration {
    bucket_arn = aws_s3_bucket.hrr_knowledge_docs.arn
    inclusion_prefixes = ["documents/"]
  }

  vector_ingestion_configuration {
    chunking_configuration {
      chunking_strategy = "HIERARCHICAL"
      hierarchical_chunking_configuration {
        level_configurations {
          max_tokens = 1500
        }
        level_configurations {
          max_tokens = 512
        }
        overlap_tokens = 100
      }
    }

    parsing_configuration {
      parsing_strategy = "BEDROCK_FOUNDATION_MODEL"
    }
  }
}

A few things I learned the hard way:

  • Set data_deletion_policy = "RETAIN" unless you want to lose your vector index every time you destroy the Terraform resource. The default deletes vectors when you remove a data source, and rebuilding a large ingestion job takes time.
  • Use inclusion_prefixes to scope which S3 paths the knowledge base watches. Without this, Bedrock indexes every file in the bucket and you pay for storage on everything.
  • Use the Bedrock foundation model parser (BEDROCK_FOUNDATION_MODEL) instead of the default parser for PDFs and complex documents. The default parser strips formatting, tables, and structure. The Bedrock model parser preserves enough context for the agent to make sense of the content.
  • Start with 1500/512 tokens as your hierarchical chunk sizes. 1500 for the parent section level, 512 for the leaf chunks. 100 token overlap. This gives you enough context per chunk without ballooning your vector index storage. Adjust based on your actual document types.

Ingestion Scheduling

Knowledge bases do not sync automatically. You need to trigger an ingestion job when your source docs change. The simplest approach is an S3 event notification that triggers a Lambda function, which calls StartIngestionJob on the knowledge base.

resource "aws_s3_bucket_notification" "hrr_kb_ingest" {
  bucket = aws_s3_bucket.hrr_knowledge_docs.id

  lambda_function {
    lambda_function_arn = aws_lambda_function.hrr_trigger_ingest.arn
    events              = ["s3:ObjectCreated:*", "s3:ObjectRemoved:*"]
    filter_prefix       = "documents/"
  }
}

resource "aws_lambda_function" "hrr_trigger_ingest" {
  filename      = "lambda/trigger_ingest.zip"
  function_name = "hrr-trigger-kb-ingest"
  role          = aws_iam_role.hrr_lambda.arn
  handler       = "index.handler"
  runtime       = "nodejs22.x"

  environment {
    variables = {
      KNOWLEDGE_BASE_ID = aws_bedrockagent_knowledge_base.hrr_docs.id
      DATA_SOURCE_ID    = aws_bedrockagent_data_source.hrr_docs_source.id
    }
  }
}

The Lambda function itself is about ten lines of code.

// index.js
import { BedrockAgent } from "@aws-sdk/client-bedrock-agent";

const client = new BedrockAgent();
const KB_ID = process.env.KNOWLEDGE_BASE_ID;
const DS_ID = process.env.DATA_SOURCE_ID;

export const handler = async (event) => {
  await client.startIngestionJob({
    knowledgeBaseId: KB_ID,
    dataSourceId: DS_ID,
  });
  return { statusCode: 202 };
};

Ingestion jobs take time. A few hundred documents finish in a couple of minutes. Thousands of documents can take an hour or more. The agent cannot query newly ingested content until the job completes, so time your syncs accordingly. I run a nightly sync for the main corpus and a real-time trigger for any document that gets updated during the day.

Defining Action Groups with Lambda Functions

Action groups are how your agent performs actions. Each action group wraps one or more Lambda functions and tells the agent what those functions can do via an OpenAPI schema.

The agent does not call your Lambda directly. It inspects the OpenAPI schema, decides which operation matches the user's request, constructs the input parameters from the conversation context, and then invokes the function. The Lambda response gets fed back into the model as a tool result, and the model decides what to do next.

Here is the Terraform for an action group:

# action_group.tf

resource "aws_bedrockagent_agent_action_group" "hrr_support" {
  agent_id        = aws_bedrockagent_agent.hrr_support.id
  agent_version   = "DRAFT"
  action_group_name = "hrr-support-actions"
  description     = "Actions for customer support agent"

  action_group_executor {
    lambda = aws_lambda_function.hrr_support_actions.arn
  }

  api_schema {
    s3 {
      s3_bucket = aws_s3_bucket.hrr_action_schemas.id
      s3_key    = "schemas/support-actions.json"
    }
  }
}

The OpenAPI schema defines each operation the agent can call. Here is an example for a support agent that can look up orders and process refunds:

{
  "openapi": "3.0.0",
  "info": {
    "title": "Support Actions",
    "version": "1.0.0"
  },
  "paths": {
    "/lookup-order": {
      "post": {
        "summary": "Look up an order by ID",
        "operationId": "lookupOrder",
        "parameters": [
          {
            "name": "orderId",
            "in": "query",
            "required": true,
            "schema": {
              "type": "string"
            },
            "description": "The order identifier"
          }
        ],
        "responses": {
          "200": {
            "description": "Order details",
            "content": {
              "application/json": {
                "schema": {
                  "type": "object",
                  "properties": {
                    "orderId": { "type": "string" },
                    "status": { "type": "string" },
                    "items": { "type": "array", "items": { "type": "string" } },
                    "total": { "type": "number" }
                  }
                }
              }
            }
          }
        }
      }
    },
    "/process-refund": {
      "post": {
        "summary": "Process a refund for an order",
        "operationId": "processRefund",
        "parameters": [
          {
            "name": "orderId",
            "in": "query",
            "required": true,
            "schema": { "type": "string" }
          },
          {
            "name": "reason",
            "in": "query",
            "required": true,
            "schema": { "type": "string" }
          }
        ],
        "responses": {
          "200": {
            "description": "Refund result",
            "content": {
              "application/json": {
                "schema": {
                  "type": "object",
                  "properties": {
                    "refundId": { "type": "string" },
                    "status": { "type": "string" }
                  }
                }
              }
            }
          }
        }
      }
    }
  }
}

Three things I have learned about writing action group schemas for Bedrock:

  • Keep operation names simple and descriptive. The agent uses the operationId to decide which function to call. lookupOrder and processRefund are clear. Avoid abbreviations or technical jargon the model might misinterpret.
  • Provide good descriptions for each parameter. The model uses these descriptions to extract the right values from the conversation. If a parameter description says "The customer's email address" and the user says "my email is foo@bar.com," the model knows what to extract. Vague descriptions lead to wrong parameter extraction.
  • Return structured JSON from your Lambda. The response is passed back to the model as-is (after guardrail filtering). Return objects with clear field names and types. Flat strings make it harder for the model to extract specific values in subsequent reasoning steps.

Here is what the Lambda function looks like on the other end:

// support-actions/index.js
import { DynamoDBClient } from "@aws-sdk/client-dynamodb";
import { DynamoDBDocumentClient, GetCommand, UpdateCommand } from "@aws-sdk/lib-dynamodb";

const db = DynamoDBDocumentClient.from(new DynamoDBClient({}));

export const handler = async (event) => {
  // The agent sends the invoked operationId and parameters
  const { operationId, parameters } = JSON.parse(event.body);

  if (operationId === "lookupOrder") {
    const order = await db.send(new GetCommand({
      TableName: process.env.ORDERS_TABLE,
      Key: { orderId: parameters.orderId },
    }));
    return {
      statusCode: 200,
      body: JSON.stringify(order.Item || { error: "Order not found" }),
    };
  }

  if (operationId === "processRefund") {
    await db.send(new UpdateCommand({
      TableName: process.env.ORDERS_TABLE,
      Key: { orderId: parameters.orderId },
      UpdateExpression: "SET refundStatus = :status, refundReason = :reason",
      ExpressionAttributeValues: {
        ":status": "PENDING",
        ":reason": parameters.reason,
      },
    }));
    return {
      statusCode: 200,
      body: JSON.stringify({ refundId: `REF-${parameters.orderId}`, status: "PENDING" }),
    };
  }

  return {
    statusCode: 400,
    body: JSON.stringify({ error: `Unknown operation: ${operationId}` }),
  };
};

Guardrails That Actually Block Things

Bedrock Guardrails are your safety layer. They sit between the user and the model, and between the model and the user. Every input and output passes through before the agent sees it or returns it.

I configure guardrails with three policies for every production agent:

Content Filters

These block specific categories of harmful content. Configure severity thresholds for each category: hate, insults, sexual, violence, misconduct, and prompt attacks.

Topic Policies

Define what the agent is allowed to discuss. Topic policies block entire categories of conversation at a higher level than content filters. For a customer support agent, you might allow "Returns" and "Shipping" but block "Account deletion" and "Billing disputes" (routing those to human agents instead).

PII Redaction

Strip personally identifiable information from inputs and outputs before they reach the model or the user. Bedrock supports regex-based patterns for common PII types. Credit card numbers, email addresses, phone numbers, SSNs. You can choose to mask or block.

# guardrails.tf

resource "aws_bedrockagent_guardrail" "hrr_support" {
  name        = "hrr-support-guardrails"
  description = "Guardrails for the customer support agent"
  blocked_input_messaging  = "I cannot process that request."
  blocked_outputs_messaging = "I cannot provide that information."

  content_policy_config {
    filters_config {
      type      = "HATE"
      input_strength  = "HIGH"
      output_strength = "HIGH"
    }
    filters_config {
      type      = "INSULTS"
      input_strength  = "MEDIUM"
      output_strength = "MEDIUM"
    }
    filters_config {
      type      = "SEXUAL"
      input_strength  = "HIGH"
      output_strength = "HIGH"
    }
    filters_config {
      type      = "VIOLENCE"
      input_strength  = "HIGH"
      output_strength = "HIGH"
    }
    filters_config {
      type      = "MISCONDUCT"
      input_strength  = "HIGH"
      output_strength = "HIGH"
    }
    filters_config {
      type      = "PROMPT_ATTACK"
      input_strength  = "HIGH"
      output_strength = "NONE"
    }
  }

  topic_policy_config {
    topics_config {
      name       = "Account Deletion"
      definition = "Requests to delete or permanently disable user accounts"
      examples = [
        "I want to delete my account",
        "Close my account permanently",
        "Remove my user profile"
      ]
      type = "DENY"
    }
    topics_config {
      name       = "Billing Disputes"
      definition = "Charges, refunds, or billing issues"
      examples = [
        "I was charged twice",
        "This bill is wrong",
        "I need a refund for an unauthorized charge"
      ]
      type = "DENY"
    }
  }

  sensitive_information_policy_config {
    pii_entities_config {
      type    = "EMAIL"
      action  = "BLOCK"
    }
    pii_entities_config {
      type    = "PHONE"
      action  = "BLOCK"
    }
    pii_entities_config {
      type    = "CREDIT_DEBIT_CARD_NUMBER"
      action  = "BLOCK"
    }
    pii_entities_config {
      type    = "US_SOCIAL_SECURITY_NUMBER"
      action  = "BLOCK"
    }
  }
}

I set PROMPT_ATTACK input strength to HIGH and output strength to NONE. Prompt injection is an input problem. The guardrail should catch attempts to override the agent's system prompt on the way in, but there is no reason to filter model outputs for prompt attacks.

For PII, I use BLOCK instead of MASK or ANONYMIZE. Masking replaces the PII with a placeholder, which means the agent still processes content that references credit card numbers. Blocking stops the request entirely. For a support agent, there is no legitimate reason to send raw credit card numbers through the model.

IAM Permissions for the Agent Execution Role

The agent execution role is the most commonly misconfigured piece of a Bedrock Agent deployment. The role needs a specific trust policy that allows Bedrock to assume it, plus permissions for whatever the agent needs to access.

# iam.tf

resource "aws_iam_role" "hrr_bedrock_agent" {
  name = "hrr-bedrock-agent-role"

  assume_role_policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Effect = "Allow"
        Principal = {
          Service = "bedrock.amazonaws.com"
        }
        Action = "sts:AssumeRole"
        Condition = {
          StringEquals = {
            "aws:SourceAccount" = data.aws_caller_account.current.account_id
          }
        }
      }
    ]
  })
}

resource "aws_iam_role_policy" "hrr_bedrock_agent_kb" {
  role = aws_iam_role.hrr_bedrock_agent.name

  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Effect = "Allow"
        Action = [
          "bedrock:Retrieve",
          "bedrock:RetrieveAndGenerate"
        ]
        Resource = aws_bedrockagent_knowledge_base.hrr_docs.arn
      },
      {
        Effect = "Allow"
        Action = [
          "aoss:APIAccessAll"
        ]
        Resource = aws_opensearchserverless_collection.hrr_kb.arn
      },
      {
        Effect = "Allow"
        Action = [
          "bedrock:InvokeModel",
          "bedrock:InvokeModelWithResponseStream"
        ]
        Resource = "arn:aws:bedrock:us-east-1::foundation-model/anthropic.claude-sonnet-4-6"
      }
    ]
  })
}

resource "aws_iam_role_policy" "hrr_bedrock_agent_lambda" {
  role = aws_iam_role.hrr_bedrock_agent.name

  policy = jsonencode({
    Version = "2012-10-17"
    Statement = [
      {
        Effect = "Allow"
        Action = "lambda:InvokeFunction"
        Resource = [
          aws_lambda_function.hrr_support_actions.arn,
          "${aws_lambda_function.hrr_support_actions.arn}:*"
        ]
      }
    ]
  })
}

Key things to get right:

  • The trust principal is bedrock.amazonaws.com. Not agents.amazonaws.com. This was the most common gotcha I hit early on. AWS Bedrock Agents use the Bedrock service principal.
  • Scope the bedrock:InvokeModel resource to specific model ARNs. A wildcard gives the agent access to every model in Bedrock, including ones you may not want it using for cost or compliance reasons. Pin it to the models you actually use.
  • Add aoss:APIAccessAll for OpenSearch Serverless. The knowledge base needs this permission to query the vector index. Without it, the agent returns "Knowledge base not found" errors that take forever to debug.
  • The Lambda invoke permission needs the :* qualifier. Bedrock Agents invoke Lambda functions using a version-specific ARN internally. If you only grant access to the unqualified ARN, the invocation fails when the agent tries to call a specific version.

CloudWatch Logging for Agent Traces

Agent traces are the single most useful debugging tool you have. Every step the agent takes gets logged. The raw model input, the model output, which knowledge base queries were executed, which action group functions were invoked, what parameters were passed, and what returned.

resource "aws_bedrockagent_agent" "hrr_support" {
  agent_name              = "hrr-support-agent"
  agent_resource_role_arn = aws_iam_role.hrr_bedrock_agent.arn
  foundation_model        = "anthropic.claude-sonnet-4-6"
  instruction             = <<-EOT
You are a customer support agent. Use the knowledge base to answer
questions about products, policies, and procedures. Use action group
functions to look up orders and process returns. If the user asks
about account deletion or billing disputes, explain that you cannot
handle those and offer to transfer them to a human agent.
EOT

  # Enable traces
  customer_encryption_key_arn = aws_kms_key.hrr_bedrock.arn

  # Optional: enable trace logging to CloudWatch
  # (enabled via the Bedrock Agent API or console)
}

resource "aws_cloudwatch_log_group" "hrr_agent_traces" {
  name              = "/aws/bedrock/agents/hrr-support"
  retention_in_days = 30
}

data "aws_iam_policy_document" "hrr_bedrock_logging" {
  statement {
    effect = "Allow"
    principals {
      type        = "Service"
      identifiers = ["bedrock.amazonaws.com"]
    }
    actions = ["logs:PutLogEvents"]
    resources = ["${aws_cloudwatch_log_group.hrr_agent_traces.arn}:*"]
  }
}

You need to enable trace logging through the Bedrock Agent API or the console. Terraform does not yet support all the agent-level configuration options for CloudWatch logging, so I use a manual step or a local-exec provisioner to set it up after the agent is created.

aws bedrock-agent update-agent \
  --agent-id $(terraform output -raw agent_id) \
  --agent-name hrr-support-agent \
  --agent-collaboration CUSTOM \
  --customer-encryption-key-arn $(terraform output -raw kms_arn)

When something goes wrong, traces tell you exactly where. I have debugged issues where the model was generating malformed tool call parameters, where the guardrail was too aggressive and blocking legitimate inputs, and where the knowledge base returned irrelevant chunks because the chunking strategy was wrong. Every one of those was visible in the trace logs within seconds.

Complete Terraform Example

Here is a full Terraform config that ties everything together. This omits the S3 bucket and OpenSearch Serverless collection definitions to keep it focused, but the full setup needs those too.

# main.tf (condensed)

provider "aws" {
  region = "us-east-1"
}

data "aws_caller_account" "current" {}

# --- Agent ---

resource "aws_bedrockagent_agent" "hrr_support" {
  agent_name              = "hrr-support-agent"
  agent_resource_role_arn = aws_iam_role.hrr_bedrock_agent.arn
  foundation_model        = "anthropic.claude-sonnet-4-6"
  instruction             = "You are a customer support agent..."
  idle_ttl                = 300
  prepare_agent   = true
}

# --- Knowledge Base ---

resource "aws_bedrockagent_knowledge_base" "hrr_docs" {
  name     = "hrr-docs-kb"
  role_arn = aws_iam_role.hrr_bedrock_agent.arn

  knowledge_base_configuration {
    type = "VECTOR"
    vector_knowledge_base_configuration {
      embedding_model_arn = "arn:aws:bedrock:us-east-1::foundation-model/amazon.titan-embed-text-v2:0"
    }
  }

  storage_configuration {
    type = "OPENSEARCH_SERVERLESS"
    opensearch_serverless_configuration {
      collection_arn    = aws_opensearchserverless_collection.hrr_kb.arn
      vector_index_name = "hrr-bedrock-kb-index"
      field_mapping {
        metadata_field = "AMAZON_BEDROCK_METADATA"
        text_field     = "AMAZON_BEDROCK_TEXT_CHUNK"
      }
    }
  }
}

resource "aws_bedrockagent_data_source" "hrr_docs_source" {
  name              = "hrr-prod-docs"
  knowledge_base_id = aws_bedrockagent_knowledge_base.hrr_docs.id
  data_source_type  = "S3"
  data_deletion_policy = "RETAIN"

  s3_configuration {
    bucket_arn = aws_s3_bucket.hrr_knowledge_docs.arn
    inclusion_prefixes = ["documents/"]
  }

  vector_ingestion_configuration {
    chunking_configuration {
      chunking_strategy = "HIERARCHICAL"
      hierarchical_chunking_configuration {
        level_configurations { max_tokens = 1500 }
        level_configurations { max_tokens = 512 }
        overlap_tokens = 100
      }
    }
    parsing_configuration {
      parsing_strategy = "BEDROCK_FOUNDATION_MODEL"
    }
  }
}

# --- Action Group ---

resource "aws_bedrockagent_agent_action_group" "hrr_support" {
  agent_id          = aws_bedrockagent_agent.hrr_support.id
  agent_version     = "DRAFT"
  action_group_name = "hrr-support-actions"
  description       = "Support-related actions"

  action_group_executor {
    lambda = aws_lambda_function.hrr_support_actions.arn
  }

  api_schema {
    s3 {
      s3_bucket = aws_s3_bucket.hrr_action_schemas.id
      s3_key    = "schemas/support-actions.json"
    }
  }
}

# --- Guardrails ---

resource "aws_bedrockagent_agent_knowledge_base_association" "hrr_docs_assoc" {
  agent_id       = aws_bedrockagent_agent.hrr_support.id
  agent_version  = "DRAFT"
  description    = "Associate knowledge base with agent"
  knowledge_base_id = aws_bedrockagent_knowledge_base.hrr_docs.id
}

resource "aws_bedrockagent_agent_guardrail_association" "hrr_guardrail_assoc" {
  agent_id      = aws_bedrockagent_agent.hrr_support.id
  agent_version = "DRAFT"
  guardrail_id  = aws_bedrockagent_guardrail.hrr_support.id
}

Testing with the Bedrock Console

The Bedrock console has a built-in test panel for agents. You can send messages, inspect the full trace, see which steps the agent took, and test individual action group functions without wiring up a client application.

When testing, pay attention to three things in the trace output:

  • Orchestration loop iterations. Count how many times the model calls tools before responding. If it is calling five or six tools for a simple question, something is wrong. Either the model is confused about the tool capabilities or the action group schema is too vague.
  • Knowledge base relevance scores. Each retrieved chunk comes with a score. If the scores are low (below 0.6 for Titan embeddings), the chunking strategy or embedding model may not be a good fit for your document types.
  • Guardrail triggers. If legitimate inputs keep getting blocked, check the content filter thresholds and topic policy definitions. I had a case where the agent classified "I need help with billing" as a prompt attack because the user's phrasing triggered the prompt injection filter. Had to dial that back.

Cost Considerations

Bedrock Agents introduce costs beyond the base model inference. Here is the breakdown I track per agent:

  • Model inference. Per-token pricing for the foundation model. Claude Sonnet 4.6 is $3.00/M input tokens and $15.00/M output tokens. Agent orchestration increases token consumption because of the system prompt, tool definitions, and conversation history that get included in every turn.
  • Knowledge base storage. OpenSearch Serverless collection costs based on OCU (OpenSearch Capacity Units). A minimal collection with one OCU costs around $300/month. This is the biggest surprise cost for most people. If your knowledge base is large, the OpenSearch costs dominate the monthly bill.
  • Knowledge base retrieval. Bedrock charges per character for knowledge base retrieval. At scale this adds up, but for most workloads it is a fraction of the inference cost.
  • Lambda invocations. Standard Lambda pricing applies. For action groups that get called frequently, these costs are negligible relative to the Bedrock inference costs.
  • Guardrail processing. Per-text-unit pricing for content filtering and PII redaction. Usually less than $0.50 per million characters processed.

My production support agent handling about 5,000 conversations per month costs roughly $450. About $200 of that is the OpenSearch Serverless collection, $200 is model inference, and the remaining $50 is everything else. The OpenSearch cost is essentially a floor. Even with zero traffic, you pay the OCU allocation.

If you want to reduce knowledge base costs, consider using the Aurora PostgreSQL vector engine instead of OpenSearch Serverless. You share an existing RDS instance and avoid the separate OpenSearch bill. The tradeoff is provisioning and managing the Aurora cluster yourself.

Production Considerations

Provisioned Throughput vs On-Demand

On-demand inference is fine for development and low-volume workloads. For production agents handling sustained traffic, provisioned throughput gives you predictable latency and avoids throttling.

Throttling shows up as ThrottlingException in your agent traces. The agent retries automatically, but each retry adds seconds of latency. For customer-facing agents, users notice the delay.

Provisioned throughput costs more but eliminates the variance. Budget 1.5x to 2x the on-demand model cost for provisioned capacity, depending on the commitment term.

Latency

A complete agent turn involves multiple model calls. The model calls itself (for reasoning), then potentially knowledge base retrieval (adds 500ms-2s depending on index size), then one or more Lambda invocations (500ms-2s each), and finally the model generates the response. End-to-end latency for a multi-step agent interaction is usually 10-30 seconds.

Do not try to make a Bedrock Agent respond in under a second. It is not architected for real-time use cases. If you need low-latency responses, use direct model inference and handle tool routing in your own code.

The biggest latency improvements I have found:

  • Reduce idle TTL on the agent (idle_ttl in Terraform). Default is 300 seconds. Lowering it to 60 seconds reduces cold start frequency for bursty traffic.
  • Use the same Lambda runtime for all action groups. If some are Python and some are Node.js, the agent has to initialize different runtimes.
  • Pre-warm the OpenSearch collection by running periodic health-check queries if your traffic pattern has sudden spikes.

Agent Versions and Aliases

Bedrock Agents use a versioning system similar to Lambda. You create a DRAFT version during development, then create numbered versions for production. Aliases point to specific versions, and you can shift traffic between versions for canary deployments.

resource "aws_bedrockagent_agent_alias" "hrr_support_prod" {
  agent_id    = aws_bedrockagent_agent.hrr_support.id
  alias_name  = "prod"
  routing_configuration {
    agent_version = "1"
  }
}

I use DRAFT for all development and testing, then promote to version 1 and alias it as "prod" for production traffic. When I need to make changes, I update the DRAFT, test in the console, create a new version, and point the prod alias at it.

Practical Summary

Bedrock Agents are not the right tool for every AI use case. If you need a single-turn Q&A endpoint, just call bedrock:InvokeModel from a Lambda function and skip the agent overhead. Use an agent when you need multi-step reasoning, conditional tool usage, or retrieval-augmented generation that the model itself decides to trigger.

Here are the decisions I would make starting from scratch today:

  • Start with one knowledge base and one action group. You can add more later. The complexity grows with each additional tool the agent can choose from, and the model makes worse decisions when overwhelmed with options.
  • Use hierarchical chunking for anything with structure. Fixed-size is a fallback, not a default.
  • Pin the foundation model ARN in IAM. Never use a wildcard for Bedrock model access in production.
  • Set guardrails before the first user request. Content filtering and topic policies are easier to configure upfront than to retrofit after something gets through.
  • Budget for the OpenSearch Serverless collection. It is the biggest fixed cost and the easiest to overlook. Consider Aurora PostgreSQL for vector storage if the cost is a concern.
  • Use the Bedrock console test panel for trace debugging. It shows you exactly what the model sees and does at every step. That visibility is worth more than any documentation.

The agent pattern is powerful for the right use case, but the infrastructure around it matters more than the agent configuration itself. Get the knowledge base chunking right, write clear action group schemas, lock down the IAM policies, and set guardrails that block the right things. The agent handles the rest.

Get the next note

I email when a new post goes up. One send a week, and only if there's something new.

Want this applied on your account? Start with Infrastructure as Code.

Related reading

More posts, sliding underneath the article

Kept below the post instead of in a sidebar, with a slow continuous motion for a cleaner editorial feel.

On this post

Comments

A reply stays under the note it answers.

0 comments

No comments yet.

If you have a note on Building Bedrock Agents in Production - Knowledge Bases, Action Groups, and Guardrails, sign in and leave it.