日本語 ← Back to home
Generative AI

How Do You Measure AI Agent Accuracy and Explainability? AWS's Three-Layer Evaluation Framework

AWS demonstrates a three-layer approach to multi-agent evaluation with Amazon Bedrock AgentCore: built-in quality checks, business constraints, and independent explainability scoring.

Article ID: TC-0068 Published:

On October 5, 2026, AWS published a technical guide to evaluating multi-agent systems for both business correctness and explainability. Using Amazon Bedrock AgentCore Evaluations, it demonstrates a three-layer assessment framework in a fictional retail supply-chain scenario: built-in quality metrics, custom business-rule checks, and independent explanation checks.

WHY FLUENT ANSWERS ARE NOT ENOUGH: An agent can sound convincing while choosing the wrong tool or violating inventory, budget, or delivery constraints. A routing recommendation is not operationally valid if it exceeds carrier capacity, even when the explanation reads well. Evaluation must cover the full workflow.

A FICTIONAL RETAIL TESTBED: AWS uses AnyCompany Retail, an invented multinational retailer with ecommerce channels, stores, and fulfillment centers. The scenario involves inventory allocation and logistics decisions. It is a reference implementation for evaluation, not a claim that a real retailer achieved measured improvements.

FIVE AGENT ROLES: A central orchestrator delegates to four specialists for optimization, distribution, routing, and analytics. The example uses Strands Agents SDK and Amazon Bedrock AgentCore Runtime. Separating responsibilities makes it easier to investigate whether an error came from routing, data retrieval, optimization, or explanation.

LAYER ONE — BUILT-IN EVALUATORS: AgentCore's built-in evaluators establish baseline quality through helpfulness and agent-specific checks such as tool-selection accuracy, response relevance, instruction following, and faithfulness. The right metric depends on each agent's primary failure mode.

LAYER TWO — BUSINESS-SPECIFIC VALIDITY: Custom evaluators check constraints that generic language metrics miss. Inventory recommendations must respect budgets, forecast demand, and warehouse capacity. Routing recommendations must satisfy delivery windows, carrier capacity, regional restrictions, and cost limits.

AN INVENTORY RULE EXAMPLE: The reference evaluator checks whether recommended stock is at least forecast demand but no more than twice that demand. It also verifies remaining budget and warehouse capacity. These thresholds belong to the demonstration and are not universal inventory-planning rules.

LAYER THREE — EXPLAINABILITY: AWS defines six separate explanation checks: decision rationale, evidence attribution, constraint reasoning, trade-off explanation, tool-use explanation, and disclosure of assumptions. Measuring explanation independently from correctness helps identify whether an agent needs better decisions or clearer communication.

CORRECT BUT OPAQUE, OR CLEAR BUT WRONG: A 1,500-unit inventory recommendation might satisfy all constraints yet fail to explain the choice. Another recommendation might offer a polished rationale while exceeding warehouse capacity. These are different failure modes and require different remedies.

ON-DEMAND VERSUS ONLINE EVALUATION: On-demand evaluations support development benchmarks, regression testing, and CI/CD gates. Online evaluations sample production traces and continuously score behavior. AWS illustrates sampling rates such as 1–10 percent; appropriate rates depend on operational risk and cost.

LINKING SCORES TO EXECUTION TRACES: Online evaluation reads AgentCore Observability traces and streams results to Amazon CloudWatch dashboards and alarms. Teams need to identify which agent and tool call caused a failure, while limiting sensitive information recorded in traces.

A REPRODUCIBLE TEST FLOW: The sample provides Terraform deployment, a test client with 20 questions across four specialist categories, and evaluation scripts that use generated session IDs. Results are processed asynchronously and written to Amazon S3. Testing requires attention to permissions, regional support, costs, and resource cleanup.

EVALUATION IS NOT ENFORCEMENT: AgentCore Evaluations measures behavior; it does not replace access controls that block unauthorized actions. AWS discusses Bedrock Guardrails as complementary runtime safeguards. Consequential operations still require permission enforcement and, where appropriate, human approval.

A PRACTICAL ROLLOUT: Define business success and failure cases, create representative test data, establish built-in baselines, add deterministic or model-assisted business checks, and score explanations separately. Repeat the same evaluations when models, prompts, tools, or workflows change.

WHAT TO WATCH: The competitive question for enterprise agents is moving beyond fluent responses toward correct tool use, business-rule compliance, and defensible explanations. Continuous evaluation tied to real execution traces can help teams improve reliability after deployment.

Source

AWS Artificial Intelligence Blog ↗