An agent ran at 3am and emailed 400 customers that their payments had failed. The payments had not failed. The engineer discovered the damage at 9am, six hours later, when the support queue hit 200 tickets.
The automation did exactly what it was built to do. It reconciled the day’s payments against a processor, identified charges that appeared to have failed, and sent polite “your payment didn’t go through” emails. The problem was a race condition: the reconciliation job read the ledger before a batch of successful payments had finished persisting. The gap looked like failures. The agent sent the emails.
This is not a story about buggy code. This is a story about the infrastructure gap between supervised and unattended agent execution.
The Two-Part Structure of Unattended Work
Every recurring automation has two distinct phases:
- Labor: Pull data, match rows, reconcile numbers, draft messages, assemble reports.
- Consequence: Send the email, charge the card, update the CRM, post the invoice.
The labor phase is safe to automate. Mistakes are cheap. The consequence phase touches real people and real systems. Mistakes are expensive.
The failure mode is treating both phases the same way. You automate the labor, then you automate the consequence, and you schedule both to run at 3am when nobody is watching.
Approval Gate Patterns for Irreversible Actions
An approval gate is a state boundary that requires human confirmation before an agent proceeds from labor to consequence. Here are three patterns that work in production.
Pattern 1: Batch Review Queue
The agent prepares the work but does not execute it. Instead, it writes the batch to a review table with a status of pending_approval. A human reviews the batch in the morning and clicks “approve” or “reject.”
Implementation shape:
class EmailBatchApproval:
def __init__(self, db, email_service):
self.db = db
self.email_service = email_service
def prepare_batch(self, failed_payments):
# Agent runs at 3am, prepares emails
batch_id = self.db.insert_batch({
'status': 'pending_approval',
'created_at': datetime.utcnow(),
'email_count': len(failed_payments),
'payload': json.dumps(failed_payments)
})
# Alert human reviewer
self.send_slack_notification(
f"Payment failure batch ready: {len(failed_payments)} emails. "
f"Review at /admin/batches/{batch_id}"
)
return batch_id
def execute_batch(self, batch_id, approved_by):
# Human runs this in the morning
batch = self.db.get_batch(batch_id)
if batch['status'] != 'pending_approval':
raise ValueError("Batch already processed")
emails = json.loads(batch['payload'])
results = self.email_service.send_bulk(emails)
self.db.update_batch(batch_id, {
'status': 'executed',
'approved_by': approved_by,
'executed_at': datetime.utcnow(),
'results': json.dumps(results)
})
Trade-off: Adds latency. Customers with real payment failures wait until morning for notification. For many use cases, this is acceptable. For time-sensitive actions (fraud alerts, system outages), it is not.
Pattern 2: Confidence Threshold with Automatic Bypass
The agent calculates a confidence score for each action. High-confidence actions execute immediately. Low-confidence actions go to the review queue.
Confidence signals:
- Data freshness (how recently was the source data updated?)
- Historical accuracy (what percentage of similar batches were correct?)
- Anomaly detection (is this batch size or timing unusual?)
Example logic:
def should_auto_execute(batch):
confidence_score = 0
# Data is fresh (updated in last 5 minutes)
if batch['data_age_seconds'] < 300:
confidence_score += 40
# Batch size is normal (within 2 std dev of mean)
if abs(batch['size'] - historical_mean) < 2 * historical_stddev:
confidence_score += 30
# No recent rollbacks in this workflow
if days_since_last_rollback > 30:
confidence_score += 30
return confidence_score >= 80
Trade-off: Requires instrumentation to calculate confidence. Requires historical data to set thresholds. Adds complexity but preserves low latency for the common case.
Pattern 3: Dry Run with Diff Review
The agent executes a dry run and writes the diff (what would change) to a review dashboard. A human reviews the diff and approves the real run.
What the diff includes:
- Number of emails to send
- Sample of 5 random email bodies
- Comparison to last 7 days of batches (size, timing, content)
- List of any customers who received a similar email in the last 30 days
Trade-off: Requires building a diff viewer. Requires the agent to support dry-run mode. Adds the most visibility but also the most implementation work.
Rate Limits and Circuit Breakers for Agent Tool Calls
Traditional API rate limits protect the server. Agent rate limits protect the world from the agent.
Tool-Level Rate Limits
Each tool the agent can call gets its own rate limit. The limit is not about server capacity. The limit is about blast radius.
| Tool | Rate Limit | Rationale |
|---|---|---|
send_email | 50/hour | Prevents mass email incidents |
charge_card | 10/hour | Prevents billing disasters |
update_crm | 200/hour | CRM can handle load, but limits data corruption scope |
query_database | 1000/hour | Read-only, low risk, high limit |
post_to_slack | 5/minute | Prevents notification spam |
Circuit Breaker on Error Rate
If the agent’s error rate crosses a threshold, the circuit breaker trips and all tool calls fail fast until a human investigates.
Implementation:
class AgentCircuitBreaker:
def __init__(self, error_threshold=0.1, window_size=100):
self.error_threshold = error_threshold
self.window_size = window_size
self.recent_calls = deque(maxlen=window_size)
self.is_open = False
def record_call(self, success):
self.recent_calls.append(success)
if len(self.recent_calls) >= self.window_size:
error_rate = 1 - (sum(self.recent_calls) / len(self.recent_calls))
if error_rate > self.error_threshold:
self.is_open = True
self.alert_human("Circuit breaker tripped: error rate {:.1%}".format(error_rate))
def allow_call(self):
if self.is_open:
raise CircuitBreakerOpen("Agent circuit breaker is open. Human review required.")
return True
When to trip the breaker:
- Error rate exceeds 10% over last 100 calls
- Any single tool call fails 3 times in a row
- Agent attempts an action that was manually rolled back in the last 24 hours
Rollback Architecture for Multi-System Actions
An agent that sends 400 emails has touched 400 customer records, 400 email service API calls, and possibly 400 CRM updates. Rolling back requires a transaction log that spans all three systems.
Event Sourcing for Agent Actions
Every agent action writes an event to an append-only log before executing. The event includes enough information to reverse the action.
Event schema:
{
"event_id": "evt_abc123",
"agent_id": "payment_reconciliation_agent",
"action": "send_email",
"timestamp": "2026-10-10T03:14:22Z",
"parameters": {
"to": "customer@example.com",
"template": "payment_failed",
"customer_id": "cust_xyz789"
},
"result": {
"email_id": "msg_def456",
"status": "sent"
},
"rollback_instructions": {
"send_apology_email": true,
"template": "payment_failed_correction",
"mark_original_as_error": true
}
}
Rollback Execution
A rollback tool reads the event log, filters to the bad batch, and executes the inverse action for each event.
For the 400-email incident:
- Query event log for all
send_emailactions between 3:00am and 3:15am - For each email sent:
- Send apology email explaining the error
- Mark original email in email service as “sent in error”
- Add note to customer record in CRM
- Log rollback action to same event log
- Generate rollback report: how many apologies sent, how many failed, how many customers contacted support before rollback completed
Rollback is not undo. You cannot unsend an email. You can only send another email that corrects the record. The event log makes this possible.
Observability for the Six-Hour Detection Gap
The agent sent emails at 3am. The engineer discovered the problem at 9am. Six hours is too long.
What to Monitor
Batch anomaly detection:
- Batch size compared to 7-day moving average
- Batch timing (did it run earlier or later than usual?)
- Batch composition (percentage of customers who received similar email in last 30 days)
Downstream signal monitoring:
- Support ticket creation rate (spike indicates customer confusion)
- Email bounce rate (spike indicates bad recipient list)
- Customer login rate (spike indicates customers checking their accounts)
Agent health metrics:
- Tool call success rate per tool
- Execution duration compared to baseline
- Data freshness at execution time
Alert Thresholds
| Metric | Threshold | Action |
|---|---|---|
| Batch size > 2x average | Immediate | Page on-call, halt execution |
| Support tickets > 1.5x average | 15 minutes | Slack alert to support team |
| Email bounce rate > 5% | Immediate | Halt email sending, investigate list |
| Agent error rate > 10% | Immediate | Trip circuit breaker, page on-call |
The 15-Minute Rule
If an unattended agent takes an action that affects more than 50 customers, a human should know within 15 minutes. This requires:
- Real-time event streaming from agent to monitoring system
- Anomaly detection that runs continuously, not on a cron schedule
- Alert routing that wakes someone up at 3am if necessary
When Approval Gates Are the Wrong Tool
Approval gates add latency and human cost. They are not always the right answer.
Use approval gates when:
- Action is irreversible (send email, charge card, delete data)
- Blast radius is large (affects more than 10 customers)
- Error cost is high (regulatory penalty, customer churn, brand damage)
- Execution frequency is low (daily or less)
Skip approval gates when:
- Action is reversible (update cache, generate report, log event)
- Blast radius is small (affects single customer or internal system)
- Error cost is low (retry fixes it, no customer impact)
- Execution frequency is high (every minute or more)
For the payment reconciliation case: The action (send email) is irreversible. The blast radius (400 customers) is large. The error cost (support load, customer confusion) is high. The execution frequency (once per day) is low. Approval gate is the right tool.
Technical Verdict
Use unattended agent execution when:
- You have event sourcing for all agent actions
- You have real-time anomaly detection with sub-15-minute alerting
- You have approval gates for irreversible actions affecting more than 10 entities
- You have tested your rollback procedure on production-like data
Avoid unattended agent execution when:
- You cannot afford the cost of a mistake (financial services, healthcare, legal)
- You do not have 24/7 on-call coverage to respond to 3am alerts
- Your agent’s error rate is above 1% and you do not know why
- You have not practiced rollback in the last 30 days
The gap between supervised and unattended execution is not a code problem. It is an infrastructure problem. You need approval gates, rate limits, circuit breakers, event logs, anomaly detection, and rollback procedures. You need them before you schedule the cron job, not after you wake up to 200 support tickets.