Phase 4: AI Agents & Autonomous Systems

Confirmation loops for high-stakes operations

Advanced ~16 min read
Think of it this way A friendly analogy. Read this if the technical version feels dense. Show Hide

Imagine you’re building a super-duper complicated LEGO castle. It’s huge, has lots of towers, secret passages, and maybe even a working drawbridge! You’re building so fast, piecing things together left and right. But every now and then, you come across a really important part – maybe it’s a giant cornerstone for the main wall, or a special electronic brick that makes the drawbridge move. If you put this piece in the wrong spot, it could mess up a whole section, or worse, make the entire castle wobbly and unstable.

This is where a "confirmation loop" comes in. Think of it like a special rule you make for your LEGO building. Before you click that really important, hard-to-move piece into place, you pause. You show it to an adult (or a master LEGO builder friend) and say, "Hey, I think this big green brick goes here to hold up the main arch because it's super strong. What do you think?" You explain why you want to put it there. Then, you wait for them to give you a clear, "Yes, that's perfect, go for it!" Not just a nod, but a definite approval. If they don't answer after a while, or if they say "Hmm, maybe not," you just put the piece aside and don't force it.

In the world of computers, when people build really smart AI programs that act like super-fast helpers, sometimes these helpers want to do something big and important. For example, an AI might want to send an email to all the parents in your school about an event, or change a setting on a very important computer system, or even make a small purchase using real money. These are like putting those huge, critical LEGO bricks in place. If the AI makes a mistake on one of these, it’s not easy to undo, and it could cause a lot of trouble – like sending the wrong email to everyone!

So, a confirmation loop is like telling your AI helper: "Before you do something huge and difficult to change, you must stop, explain your plan to a human, and wait for them to give you a clear 'Go ahead!'." This means when you get older and start building your own amazing AI tools, you'll know how to add these smart safety stops. You can make sure your AI helpers are not just super-smart, but also super-safe, so they don’t accidentally break your digital castle while trying to build it faster.

The mental model behind confirmation loops is the same as a two-person rule in nuclear command: no single autonomous actor should be able to initiate a catastrophic action unilaterally. For AI agents, the "catastrophic" bar is much lower than nuclear, but the principle holds. The agent is one actor; the human operator is the second. Both must agree before the trigger is pulled.

Under the hood, a well-designed loop has four components. First, a risk classifier that tags every candidate action before it enters the tool-call pipeline. Classification criteria should include reversibility (can this be undone in under five minutes?), blast radius (how many resources, users, or dollars are affected?), and authorization scope (does the action touch resources outside the agent's normal operating envelope?). Second, a structured proposal builder that serializes not just what the agent wants to do, but why -- the chain of reasoning, the evidence it considered, and what it expects to happen. Third, a delivery and wait mechanism: this could be a Slack message, a push notification, a webhook, or a dedicated approval UI, with a hard timeout and a safe default. Fourth, a state machine that transitions the action from PENDING to APPROVED, REJECTED, or TIMED_OUT, and writes each transition to an append-only log.

Here is a real-world scenario. An agent managing a SaaS company's Postgres cluster detects that a table has 40 million rows and no recent reads. Its tool suite includes a DROP TABLE function. Without a confirmation loop, it might drop the table as part of a cleanup task. With a confirmation loop, it emits a proposal: "I plan to drop table user_sessions_archive_2021. Reasoning: zero reads in 180 days, 12 GB storage. Affected resources: 1 table, 0 foreign key references. Reversible: no (no scheduled backup in the next 4 hours). Risk tier: HIGH." That proposal routes to two on-call engineers. Both must approve within 30 minutes. If either rejects or if the window lapses, the action aborts and logs TIMED_OUT. The agent moves on. This is what a senior engineer would build on day one of giving an agent write access to production.

The tradeoff versus alternative approaches is worth examining. The simplest alternative is a read-only agent -- the agent only recommends, a human executes. That eliminates the risk entirely but also eliminates the value of automation for high-frequency decisions. Another alternative is rule-based pre-filters: blocklists of forbidden actions (never DROP, never DELETE without WHERE). These are cheap and fast but brittle; agents find creative paths around rigid rules. Confirmation loops sit in the middle: they allow the agent to propose any action, but they enforce human review for the risky ones. The cost is latency and operator attention. The benefit is that you catch the cases your blocklist never anticipated.

At scale, the design changes significantly. At 10 users, a Slack DM to the founder works fine. At 10,000 users running agents concurrently, you might generate hundreds of confirmation requests per hour. Now you need an approval queue with SLA tracking, auto-escalation when an approver goes dark, and probably a tiered system where a trained reviewer handles MEDIUM tier approvals while HIGH tier still requires a senior engineer. At 10 million users, most actions must be auto-approved by a secondary AI reviewer rather than a human -- you're running a meta-agent that evaluates proposals against a policy document and approves low-risk ones programmatically, only escalating genuine outliers to humans. This is sometimes called "AI-in-the-loop" rather than "human-in-the-loop." The confirmation loop architecture stays the same; only the approver changes.

Cost and latency implications are real. A synchronous confirmation loop introduces latency equal to human response time, which might be seconds for a Slack notification or hours for email. Design your agent to continue other work while waiting. Use async patterns -- the tool call returns a PENDING status immediately, and the agent polls or receives a webhook when the approval resolves. For cost, the overhead is cheap: a few extra API calls to serialize and deserialize the proposal, plus whatever your notification channel costs. The expensive failure mode is not the loop itself but the missed loop: one uncaught DROP TABLE can cost more to recover from than months of confirmation infrastructure.

Key Takeaways

  • Classify actions by reversibility and blast radius before deciding confirmation tier.
  • Always include the agent's reasoning and expected side effects in the confirmation payload.
  • Enforce a safe-default timeout: if no approval arrives, the action must abort automatically.
  • Log every proposal, approval, rejection, and timeout with structured metadata for audits.

Pro tips

  • Never make the safe default "approve." If your polling loop errors out or times out, the action must abort. Defaulting to approval on ambiguity defeats the entire mechanism -- you've built a rubber stamp, not a safety gate.
  • Encode the blast radius in human terms, not system terms. "Affects 3 S3 bucket policies" means nothing to an on-call engineer at 2 a.m. "Affects billing access for all 4,000 paying customers" triggers the right instinct immediately.
  • Track confirmation fatigue. If your MEDIUM tier generates 50 approvals per day, engineers start approving without reading within a week. Instrument approval latency and acceptance rate -- if latency drops below 10 seconds for a HIGH tier action, your loop has become theater.
  • For async agents, return a PENDING action ID immediately and let the agent continue other tasks. Blocking the entire agent thread on human response time is an anti-pattern that kills throughput and makes the confirmation loop feel expensive -- it isn't, if you do it async.

Common pitfalls

  • Mistake: Showing operators only the action description, not the reasoning chain. Fix: Always serialize and display the agent's full reasoning so the reviewer can evaluate whether the logic is sound, not just whether the output looks plausible.
  • Mistake: Using a single global timeout for all risk tiers. Fix: Set tier-specific timeouts -- LOW can be milliseconds (auto), MEDIUM minutes, HIGH potentially hours -- and configure them via environment variables, not hardcoded constants.
  • Mistake: Logging only the final outcome (approved/rejected) and discarding the proposal payload. Fix: Append-only log every state transition plus the full proposal JSON so post-incident analysis can reconstruct exactly what the agent intended and why.
  • Mistake: Treating the confirmation loop as a UI problem and building it last. Fix: Design the ActionProposal schema and the state machine on day one; the delivery channel (Slack, email, webhook) is a plugin that can be swapped without touching core safety logic.

When to use which confirmation tier

Option Use when Avoid when
Auto-approve (LOW tier) Action is fully reversible, affects fewer than 10 resources, and is within the agent's routine operating scope. Any irreversibility exists, or the action touches resources outside the agent's normal envelope.
Single-human approval (MEDIUM tier) Action is irreversible or affects a moderate number of resources, but a single on-call engineer has enough context to evaluate it quickly. Financial, security, or compliance implications require documented multi-party accountability.
Multi-party approval (HIGH tier) Action is irreversible, high blast radius, touches security boundaries, or has regulatory compliance requirements. Approvers are unavailable and the SLA cannot tolerate multi-hour waits -- in that case, block the action entirely until availability is confirmed.
AI-in-the-loop secondary reviewer Volume of MEDIUM tier approvals exceeds human capacity and actions follow predictable patterns a policy-checking model can evaluate. Actions are novel, high-stakes, or legally significant -- AI reviewers should never be the sole gate for HIGH tier actions.

Code Example

python
# openai>=1.0.0, pydantic>=2.0
from pydantic import BaseModel
from enum import Enum
import time

class RiskTier(str, Enum):
    LOW = "low"       # auto-approve
    MEDIUM = "medium" # single human approval
    HIGH = "high"     # multi-party approval

class ActionProposal(BaseModel):
    action_id: str
    description: str
    reasoning: str
    affected_resources: list[str]
    reversible: bool
    risk_tier: RiskTier

def request_confirmation(proposal: ActionProposal, timeout_seconds: int = 120) -> bool:
    if proposal.risk_tier == RiskTier.LOW:
        return True  # auto-approved
    print(f"[CONFIRMATION REQUIRED]\n{proposal.model_dump_json(indent=2)}")
    deadline = time.time() + timeout_seconds
    response = input(f"Type 'APPROVE' to continue (timeout {timeout_seconds}s): ")
    if time.time() > deadline:
        print("Timeout reached. Action ABORTED.")
        return False
    return response.strip().upper() == "APPROVE"

How this code works

This code establishes a crucial safety mechanism for AI agents, often called a "confirmation loop." Its main job is to ensure that AI-proposed actions, especially those with significant impact, receive appropriate human review and explicit approval. It prevents agents from automatically executing high-stakes operations, instead requiring a human "sign-off" based on the action's perceived risk. This system uses pydantic's BaseModel to structure proposed actions consistently and enum to categorize their risk levels.

The ActionProposal BaseModel defines key details for any action, including its description, affected_resources, and a risk_tier. The RiskTier Enum categorizes actions as LOW, MEDIUM, or HIGH, dictating the necessary approval process. The core logic resides in the request_confirmation function. It instantly auto-approves LOW risk actions. For higher risks, it prints the full proposal using model_dump_json and prompts for human input, specifically 'APPROVE'. A subtle but vital detail is the timeout_seconds parameter; if a human doesn't respond within this timeframe (defaulting to 120 seconds), the action is automatically ABORTED, preventing an unconfirmed, potentially risky operation from hanging indefinitely. Only an explicit 'APPROVE' within the deadline will allow the action to proceed.

Production-grade example

Adds retry delivery, structured JSON logging, env-var config, timeout enforcement, and safe-default abort on failure.

python
# openai>=1.0.0, pydantic>=2.0, tenacity>=8.0
import os
import uuid
import time
import logging
from enum import Enum
from datetime import datetime, timezone
from pydantic import BaseModel
from tenacity import retry, stop_after_attempt, wait_exponential

log = logging.getLogger("agent.confirmation")
logging.basicConfig(level=logging.INFO, format='{"time":"%(asctime)s","level":"%(levelname)s","msg":%(message)s}')

class RiskTier(str, Enum):
    LOW = "low"
    MEDIUM = "medium"
    HIGH = "high"

class ConfirmationStatus(str, Enum):
    PENDING = "pending"
    APPROVED = "approved"
    REJECTED = "rejected"
    TIMED_OUT = "timed_out"

class ActionProposal(BaseModel):
    action_id: str
    description: str
    reasoning: str
    affected_resources: list[str]
    reversible: bool
    risk_tier: RiskTier
    requested_at: str

class ConfirmationResult(BaseModel):
    action_id: str
    status: ConfirmationStatus
    approved_by: str | None = None
    resolved_at: str | None = None

# Approval store -- replace with Redis or a DB in real deployments
_pending_approvals: dict[str, ConfirmationResult] = {}

TIMEOUT_SECONDS = int(os.environ.get("CONFIRMATION_TIMEOUT_SECONDS", "300"))
REQUIRED_APPROVERS = {"high": 2, "medium": 1, "low": 0}

def classify_action(description: str, reversible: bool, affected_count: int) -> RiskTier:
    if not reversible and affected_count > 100:
        return RiskTier.HIGH
    if not reversible or affected_count > 10:
        return RiskTier.MEDIUM
    return RiskTier.LOW

@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def _deliver_notification(proposal: ActionProposal) -> None:
    """Deliver via your real channel: Slack webhook, PagerDuty, etc."""
    webhook_url = os.environ.get("APPROVAL_WEBHOOK_URL")
    if not webhook_url:
        raise EnvironmentError("APPROVAL_WEBHOOK_URL not set")
    # import httpx; httpx.post(webhook_url, json=proposal.model_dump(), timeout=10)
    log.info(f'"event":"notification_sent","action_id":"{proposal.action_id}","tier":"{proposal.risk_tier}"')

def request_confirmation(proposal: ActionProposal) -> ConfirmationResult:
    result = ConfirmationResult(action_id=proposal.action_id, status=ConfirmationStatus.PENDING)
    _pending_approvals[proposal.action_id] = result

    if proposal.risk_tier == RiskTier.LOW:
        result.status = ConfirmationStatus.APPROVED
        result.approved_by = "auto"
        result.resolved_at = datetime.now(timezone.utc).isoformat()
        log.info(f'"event":"auto_approved","action_id":"{proposal.action_id}"')
        return result

    try:
        _deliver_notification(proposal)
    except Exception as exc:
        log.error(f'"event":"notification_failed","action_id":"{proposal.action_id}","error":"{exc}"')
        result.status = ConfirmationStatus.REJECTED
        result.resolved_at = datetime.now(timezone.utc).isoformat()
        return result

    deadline = time.monotonic() + TIMEOUT_SECONDS
    poll_interval = 5
    while time.monotonic() < deadline:
        current = _pending_approvals.get(proposal.action_id)
        if current and current.status != ConfirmationStatus.PENDING:
            log.info(f'"event":"resolved","action_id":"{proposal.action_id}","status":"{current.status}"')
            return current
        time.sleep(poll_interval)

    result.status = ConfirmationStatus.TIMED_OUT
    result.resolved_at = datetime.now(timezone.utc).isoformat()
    _pending_approvals[proposal.action_id] = result
    log.warning(f'"event":"timed_out","action_id":"{proposal.action_id}","timeout_seconds":{TIMEOUT_SECONDS}')
    return result

def execute_action_with_confirmation(
    description: str,
    reasoning: str,
    affected_resources: list[str],
    reversible: bool,
    action_fn,
):
    proposal = ActionProposal(
        action_id=str(uuid.uuid4()),
        description=description,
        reasoning=reasoning,
        affected_resources=affected_resources,
        reversible=reversible,
        risk_tier=classify_action(description, reversible, len(affected_resources)),
        requested_at=datetime.now(timezone.utc).isoformat(),
    )
    result = request_confirmation(proposal)
    if result.status == ConfirmationStatus.APPROVED:
        log.info(f'"event":"executing","action_id":"{proposal.action_id}"')
        return action_fn()
    log.warning(f'"event":"action_blocked","action_id":"{proposal.action_id}","reason":"{result.status}"')
    return None

How this code works

This code establishes a "confirmation loop" to ensure human oversight for high-stakes AI agent operations. Its primary job is to pause potentially impactful actions, request approval from a human, and only proceed if approved within a given timeframe, enhancing agent safety and guardrails.

The system first defines clear states using Enum for RiskTier (e.g., HIGH, LOW) and ConfirmationStatus (e.g., PENDING, APPROVED). Action details are structured with ActionProposal and ConfirmationResult using Pydantic's BaseModel for data validation. The classify_action function automatically determines an action's risk. The request_confirmation function then initiates the loop: low-risk actions are auto_approved. For higher risks, it attempts to send a notification via _deliver_notification. A subtle but important detail is the @retry decorator on _deliver_notification; this ensures the notification is retried up to three times if the initial attempt fails, improving reliability before an action is rejected due to a transient communication error. Finally, it enters a while loop, continuously checking the _pending_approvals store until the action is approved, rejected, or TIMED_OUT after a specified period.

Practice & master

Try the exercise, check your understanding, then mark this lesson mastered to track your path to pro.

Exercise

Build a confirmation loop for a hypothetical agent that can send bulk email campaigns. Implement risk classification based on recipient count and whether the email list was manually uploaded vs. automatically generated. Route appropriately to auto-approve, single-approval, or multi-party approval. Enforce a 60-second timeout with a TIMED_OUT default.

python
import time
import uuid
from enum import Enum
from pydantic import BaseModel

class RiskTier(str, Enum):
    LOW = "low"
    MEDIUM = "medium"
    HIGH = "high"

class EmailCampaignProposal(BaseModel):
    action_id: str
    subject: str
    recipient_count: int
    list_source: str  # "manual" or "auto-generated"
    risk_tier: RiskTier

def classify_email_action(recipient_count: int, list_source: str) -> RiskTier:
    # TODO: return HIGH if auto-generated and >1000 recipients
    # TODO: return MEDIUM if manual and >500, or auto-generated and <=1000
    # TODO: return LOW otherwise
    pass

def request_confirmation(proposal: EmailCampaignProposal, timeout_seconds: int = 60) -> str:
    # TODO: auto-approve LOW tier
    # TODO: for MEDIUM/HIGH, print proposal and wait for input
    # TODO: enforce timeout and return "TIMED_OUT" if exceeded
    # TODO: return "APPROVED", "REJECTED", or "TIMED_OUT"
    pass

if __name__ == "__main__":
    proposal = EmailCampaignProposal(
        action_id=str(uuid.uuid4()),
        subject="Q4 Newsletter",
        recipient_count=1500,
        list_source="auto-generated",
        risk_tier=classify_email_action(1500, "auto-generated"),
    )
    result = request_confirmation(proposal)
    print(f"Result: {result}")

Quick check

  1. An agent's confirmation loop defaults to APPROVED when the polling service returns a 503 error. What is wrong with this design?

  2. Why should an ActionProposal include the agent's reasoning chain, not just the proposed action description?

  3. You observe that on-call engineers are approving HIGH tier agent actions in under 8 seconds on average. What does this most likely indicate?

Self-check: Without looking at your notes, describe the four components of a production confirmation loop, explain what the safe default should be on timeout and why, and give one concrete signal that indicates your confirmation loop has become rubber-stamp theater rather than a real safety gate.