The mental model behind confirmation loops is the same as a two-person rule in nuclear command: no single autonomous actor should be able to initiate a catastrophic action unilaterally. For AI agents, the "catastrophic" bar is much lower than nuclear, but the principle holds. The agent is one actor; the human operator is the second. Both must agree before the trigger is pulled.
Under the hood, a well-designed loop has four components. First, a risk classifier that tags every candidate action before it enters the tool-call pipeline. Classification criteria should include reversibility (can this be undone in under five minutes?), blast radius (how many resources, users, or dollars are affected?), and authorization scope (does the action touch resources outside the agent's normal operating envelope?). Second, a structured proposal builder that serializes not just what the agent wants to do, but why -- the chain of reasoning, the evidence it considered, and what it expects to happen. Third, a delivery and wait mechanism: this could be a Slack message, a push notification, a webhook, or a dedicated approval UI, with a hard timeout and a safe default. Fourth, a state machine that transitions the action from PENDING to APPROVED, REJECTED, or TIMED_OUT, and writes each transition to an append-only log.
Here is a real-world scenario. An agent managing a SaaS company's Postgres cluster detects that a table has 40 million rows and no recent reads. Its tool suite includes a DROP TABLE function. Without a confirmation loop, it might drop the table as part of a cleanup task. With a confirmation loop, it emits a proposal: "I plan to drop table user_sessions_archive_2021. Reasoning: zero reads in 180 days, 12 GB storage. Affected resources: 1 table, 0 foreign key references. Reversible: no (no scheduled backup in the next 4 hours). Risk tier: HIGH." That proposal routes to two on-call engineers. Both must approve within 30 minutes. If either rejects or if the window lapses, the action aborts and logs TIMED_OUT. The agent moves on. This is what a senior engineer would build on day one of giving an agent write access to production.
The tradeoff versus alternative approaches is worth examining. The simplest alternative is a read-only agent -- the agent only recommends, a human executes. That eliminates the risk entirely but also eliminates the value of automation for high-frequency decisions. Another alternative is rule-based pre-filters: blocklists of forbidden actions (never DROP, never DELETE without WHERE). These are cheap and fast but brittle; agents find creative paths around rigid rules. Confirmation loops sit in the middle: they allow the agent to propose any action, but they enforce human review for the risky ones. The cost is latency and operator attention. The benefit is that you catch the cases your blocklist never anticipated.
At scale, the design changes significantly. At 10 users, a Slack DM to the founder works fine. At 10,000 users running agents concurrently, you might generate hundreds of confirmation requests per hour. Now you need an approval queue with SLA tracking, auto-escalation when an approver goes dark, and probably a tiered system where a trained reviewer handles MEDIUM tier approvals while HIGH tier still requires a senior engineer. At 10 million users, most actions must be auto-approved by a secondary AI reviewer rather than a human -- you're running a meta-agent that evaluates proposals against a policy document and approves low-risk ones programmatically, only escalating genuine outliers to humans. This is sometimes called "AI-in-the-loop" rather than "human-in-the-loop." The confirmation loop architecture stays the same; only the approver changes.
Cost and latency implications are real. A synchronous confirmation loop introduces latency equal to human response time, which might be seconds for a Slack notification or hours for email. Design your agent to continue other work while waiting. Use async patterns -- the tool call returns a PENDING status immediately, and the agent polls or receives a webhook when the approval resolves. For cost, the overhead is cheap: a few extra API calls to serialize and deserialize the proposal, plus whatever your notification channel costs. The expensive failure mode is not the loop itself but the missed loop: one uncaught DROP TABLE can cost more to recover from than months of confirmation infrastructure.
Key Takeaways
- Classify actions by reversibility and blast radius before deciding confirmation tier.
- Always include the agent's reasoning and expected side effects in the confirmation payload.
- Enforce a safe-default timeout: if no approval arrives, the action must abort automatically.
- Log every proposal, approval, rejection, and timeout with structured metadata for audits.
Pro tips
- Never make the safe default "approve." If your polling loop errors out or times out, the action must abort. Defaulting to approval on ambiguity defeats the entire mechanism -- you've built a rubber stamp, not a safety gate.
- Encode the blast radius in human terms, not system terms. "Affects 3 S3 bucket policies" means nothing to an on-call engineer at 2 a.m. "Affects billing access for all 4,000 paying customers" triggers the right instinct immediately.
- Track confirmation fatigue. If your MEDIUM tier generates 50 approvals per day, engineers start approving without reading within a week. Instrument approval latency and acceptance rate -- if latency drops below 10 seconds for a HIGH tier action, your loop has become theater.
- For async agents, return a PENDING action ID immediately and let the agent continue other tasks. Blocking the entire agent thread on human response time is an anti-pattern that kills throughput and makes the confirmation loop feel expensive -- it isn't, if you do it async.
Common pitfalls
- Mistake: Showing operators only the action description, not the reasoning chain. Fix: Always serialize and display the agent's full reasoning so the reviewer can evaluate whether the logic is sound, not just whether the output looks plausible.
- Mistake: Using a single global timeout for all risk tiers. Fix: Set tier-specific timeouts -- LOW can be milliseconds (auto), MEDIUM minutes, HIGH potentially hours -- and configure them via environment variables, not hardcoded constants.
- Mistake: Logging only the final outcome (approved/rejected) and discarding the proposal payload. Fix: Append-only log every state transition plus the full proposal JSON so post-incident analysis can reconstruct exactly what the agent intended and why.
- Mistake: Treating the confirmation loop as a UI problem and building it last. Fix: Design the ActionProposal schema and the state machine on day one; the delivery channel (Slack, email, webhook) is a plugin that can be swapped without touching core safety logic.
When to use which confirmation tier
| Option | Use when | Avoid when |
|---|---|---|
| Auto-approve (LOW tier) | Action is fully reversible, affects fewer than 10 resources, and is within the agent's routine operating scope. | Any irreversibility exists, or the action touches resources outside the agent's normal envelope. |
| Single-human approval (MEDIUM tier) | Action is irreversible or affects a moderate number of resources, but a single on-call engineer has enough context to evaluate it quickly. | Financial, security, or compliance implications require documented multi-party accountability. |
| Multi-party approval (HIGH tier) | Action is irreversible, high blast radius, touches security boundaries, or has regulatory compliance requirements. | Approvers are unavailable and the SLA cannot tolerate multi-hour waits -- in that case, block the action entirely until availability is confirmed. |
| AI-in-the-loop secondary reviewer | Volume of MEDIUM tier approvals exceeds human capacity and actions follow predictable patterns a policy-checking model can evaluate. | Actions are novel, high-stakes, or legally significant -- AI reviewers should never be the sole gate for HIGH tier actions. |
Code Example
# openai>=1.0.0, pydantic>=2.0
from pydantic import BaseModel
from enum import Enum
import time
class RiskTier(str, Enum):
LOW = "low" # auto-approve
MEDIUM = "medium" # single human approval
HIGH = "high" # multi-party approval
class ActionProposal(BaseModel):
action_id: str
description: str
reasoning: str
affected_resources: list[str]
reversible: bool
risk_tier: RiskTier
def request_confirmation(proposal: ActionProposal, timeout_seconds: int = 120) -> bool:
if proposal.risk_tier == RiskTier.LOW:
return True # auto-approved
print(f"[CONFIRMATION REQUIRED]\n{proposal.model_dump_json(indent=2)}")
deadline = time.time() + timeout_seconds
response = input(f"Type 'APPROVE' to continue (timeout {timeout_seconds}s): ")
if time.time() > deadline:
print("Timeout reached. Action ABORTED.")
return False
return response.strip().upper() == "APPROVE"How this code works
This code establishes a crucial safety mechanism for AI agents, often called a "confirmation loop." Its main job is to ensure that AI-proposed actions, especially those with significant impact, receive appropriate human review and explicit approval. It prevents agents from automatically executing high-stakes operations, instead requiring a human "sign-off" based on the action's perceived risk. This system uses pydantic's BaseModel to structure proposed actions consistently and enum to categorize their risk levels.
The ActionProposal BaseModel defines key details for any action, including its description, affected_resources, and a risk_tier. The RiskTier Enum categorizes actions as LOW, MEDIUM, or HIGH, dictating the necessary approval process. The core logic resides in the request_confirmation function. It instantly auto-approves LOW risk actions. For higher risks, it prints the full proposal using model_dump_json and prompts for human input, specifically 'APPROVE'. A subtle but vital detail is the timeout_seconds parameter; if a human doesn't respond within this timeframe (defaulting to 120 seconds), the action is automatically ABORTED, preventing an unconfirmed, potentially risky operation from hanging indefinitely. Only an explicit 'APPROVE' within the deadline will allow the action to proceed.
Production-grade example
Adds retry delivery, structured JSON logging, env-var config, timeout enforcement, and safe-default abort on failure.
# openai>=1.0.0, pydantic>=2.0, tenacity>=8.0
import os
import uuid
import time
import logging
from enum import Enum
from datetime import datetime, timezone
from pydantic import BaseModel
from tenacity import retry, stop_after_attempt, wait_exponential
log = logging.getLogger("agent.confirmation")
logging.basicConfig(level=logging.INFO, format='{"time":"%(asctime)s","level":"%(levelname)s","msg":%(message)s}')
class RiskTier(str, Enum):
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
class ConfirmationStatus(str, Enum):
PENDING = "pending"
APPROVED = "approved"
REJECTED = "rejected"
TIMED_OUT = "timed_out"
class ActionProposal(BaseModel):
action_id: str
description: str
reasoning: str
affected_resources: list[str]
reversible: bool
risk_tier: RiskTier
requested_at: str
class ConfirmationResult(BaseModel):
action_id: str
status: ConfirmationStatus
approved_by: str | None = None
resolved_at: str | None = None
# Approval store -- replace with Redis or a DB in real deployments
_pending_approvals: dict[str, ConfirmationResult] = {}
TIMEOUT_SECONDS = int(os.environ.get("CONFIRMATION_TIMEOUT_SECONDS", "300"))
REQUIRED_APPROVERS = {"high": 2, "medium": 1, "low": 0}
def classify_action(description: str, reversible: bool, affected_count: int) -> RiskTier:
if not reversible and affected_count > 100:
return RiskTier.HIGH
if not reversible or affected_count > 10:
return RiskTier.MEDIUM
return RiskTier.LOW
@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def _deliver_notification(proposal: ActionProposal) -> None:
"""Deliver via your real channel: Slack webhook, PagerDuty, etc."""
webhook_url = os.environ.get("APPROVAL_WEBHOOK_URL")
if not webhook_url:
raise EnvironmentError("APPROVAL_WEBHOOK_URL not set")
# import httpx; httpx.post(webhook_url, json=proposal.model_dump(), timeout=10)
log.info(f'"event":"notification_sent","action_id":"{proposal.action_id}","tier":"{proposal.risk_tier}"')
def request_confirmation(proposal: ActionProposal) -> ConfirmationResult:
result = ConfirmationResult(action_id=proposal.action_id, status=ConfirmationStatus.PENDING)
_pending_approvals[proposal.action_id] = result
if proposal.risk_tier == RiskTier.LOW:
result.status = ConfirmationStatus.APPROVED
result.approved_by = "auto"
result.resolved_at = datetime.now(timezone.utc).isoformat()
log.info(f'"event":"auto_approved","action_id":"{proposal.action_id}"')
return result
try:
_deliver_notification(proposal)
except Exception as exc:
log.error(f'"event":"notification_failed","action_id":"{proposal.action_id}","error":"{exc}"')
result.status = ConfirmationStatus.REJECTED
result.resolved_at = datetime.now(timezone.utc).isoformat()
return result
deadline = time.monotonic() + TIMEOUT_SECONDS
poll_interval = 5
while time.monotonic() < deadline:
current = _pending_approvals.get(proposal.action_id)
if current and current.status != ConfirmationStatus.PENDING:
log.info(f'"event":"resolved","action_id":"{proposal.action_id}","status":"{current.status}"')
return current
time.sleep(poll_interval)
result.status = ConfirmationStatus.TIMED_OUT
result.resolved_at = datetime.now(timezone.utc).isoformat()
_pending_approvals[proposal.action_id] = result
log.warning(f'"event":"timed_out","action_id":"{proposal.action_id}","timeout_seconds":{TIMEOUT_SECONDS}')
return result
def execute_action_with_confirmation(
description: str,
reasoning: str,
affected_resources: list[str],
reversible: bool,
action_fn,
):
proposal = ActionProposal(
action_id=str(uuid.uuid4()),
description=description,
reasoning=reasoning,
affected_resources=affected_resources,
reversible=reversible,
risk_tier=classify_action(description, reversible, len(affected_resources)),
requested_at=datetime.now(timezone.utc).isoformat(),
)
result = request_confirmation(proposal)
if result.status == ConfirmationStatus.APPROVED:
log.info(f'"event":"executing","action_id":"{proposal.action_id}"')
return action_fn()
log.warning(f'"event":"action_blocked","action_id":"{proposal.action_id}","reason":"{result.status}"')
return NoneHow this code works
This code establishes a "confirmation loop" to ensure human oversight for high-stakes AI agent operations. Its primary job is to pause potentially impactful actions, request approval from a human, and only proceed if approved within a given timeframe, enhancing agent safety and guardrails.
The system first defines clear states using Enum for RiskTier (e.g., HIGH, LOW) and ConfirmationStatus (e.g., PENDING, APPROVED). Action details are structured with ActionProposal and ConfirmationResult using Pydantic's BaseModel for data validation. The classify_action function automatically determines an action's risk. The request_confirmation function then initiates the loop: low-risk actions are auto_approved. For higher risks, it attempts to send a notification via _deliver_notification. A subtle but important detail is the @retry decorator on _deliver_notification; this ensures the notification is retried up to three times if the initial attempt fails, improving reliability before an action is rejected due to a transient communication error. Finally, it enters a while loop, continuously checking the _pending_approvals store until the action is approved, rejected, or TIMED_OUT after a specified period.
Practice & master
Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.
Exercise
Build a confirmation loop for a hypothetical agent that can send bulk email campaigns. Implement risk classification based on recipient count and whether the email list was manually uploaded vs. automatically generated. Route appropriately to auto-approve, single-approval, or multi-party approval. Enforce a 60-second timeout with a TIMED_OUT default.
import time
import uuid
from enum import Enum
from pydantic import BaseModel
class RiskTier(str, Enum):
LOW = "low"
MEDIUM = "medium"
HIGH = "high"
class EmailCampaignProposal(BaseModel):
action_id: str
subject: str
recipient_count: int
list_source: str # "manual" or "auto-generated"
risk_tier: RiskTier
def classify_email_action(recipient_count: int, list_source: str) -> RiskTier:
# TODO: return HIGH if auto-generated and >1000 recipients
# TODO: return MEDIUM if manual and >500, or auto-generated and <=1000
# TODO: return LOW otherwise
pass
def request_confirmation(proposal: EmailCampaignProposal, timeout_seconds: int = 60) -> str:
# TODO: auto-approve LOW tier
# TODO: for MEDIUM/HIGH, print proposal and wait for input
# TODO: enforce timeout and return "TIMED_OUT" if exceeded
# TODO: return "APPROVED", "REJECTED", or "TIMED_OUT"
pass
if __name__ == "__main__":
proposal = EmailCampaignProposal(
action_id=str(uuid.uuid4()),
subject="Q4 Newsletter",
recipient_count=1500,
list_source="auto-generated",
risk_tier=classify_email_action(1500, "auto-generated"),
)
result = request_confirmation(proposal)
print(f"Result: {result}")Quick check
An agent's confirmation loop defaults to APPROVED when the polling service returns a 503 error. What is wrong with this design?
Why should an ActionProposal include the agent's reasoning chain, not just the proposed action description?
You observe that on-call engineers are approving HIGH tier agent actions in under 8 seconds on average. What does this most likely indicate?