Most text-to-SQL benchmarks test whether an agent can generate a single SELECT statement. Argo-Bench asks a harder question: can an agent navigate 235 tables, reconstruct hidden business facts, run statistical analyses, and execute actions that change the state of a simulated enterprise?
The answer is no. Frontier models score above 95 on only 34.8% of tasks and average 59.5 points. The gap reveals what breaks when you move from query generation to multi-stage data workflows.
The Problem with Existing Benchmarks
Text-to-SQL benchmarks like Spider and BIRD evaluate query generation in isolation. You get a natural language question, a schema, and a database. The agent writes SQL. The grader compares output to a reference answer.
This setup has three structural flaws:
- Answer keys are frequently wrong. Audits of popular benchmarks find incorrect ground truth, making it unclear whether a failing agent is broken or correct.
- Single-table bias. Public datasets fit business events into one table. Real enterprise warehouses spread a single transaction across dozens of normalized tables.
- No action execution. Agents generate queries but never act on results. Real workflows require banning accounts, allocating budgets, or issuing refunds based on what the query reveals.
Argo-Bench addresses all three by grading actions in a simulator, not SQL strings against answer keys.
Architecture: Simulated World Plus Hidden Ground Truth
Argo-Bench models a food delivery platform in New York City with 81 million orders in 2024. The simulation includes:
- Grounded economics (pricing, fees, discounts)
- Fraud patterns (account takeovers, promo abuse)
- Marketplace incentives (courier bonuses, restaurant promotions)
The simulator exports data to a 235-table ERP warehouse modeled on Oracle E-Business Suite. The warehouse contains 7.5 billion rows.
The critical design choice: the simulator’s ground-truth state is withheld from the warehouse the agent sees. Tasks require reconstructing facts by navigating incomplete or denormalized data before acting.
For example, identifying fraudulent accounts requires joining order history, payment methods, device fingerprints, and geolocation logs across multiple tables. The agent must infer fraud patterns, not just query a fraud_flag column.
Task Structure and Grading
Each of the 210 tasks follows this flow:
- Natural language task description (e.g., “Ban accounts that placed more than 10 orders from different zip codes in 24 hours”)
- Agent explores the warehouse (queries, statistical analysis, reasoning across tables)
- Agent files an action (ban list, budget allocation, refund batch)
- Grader scores the action by its consequences in the simulator
The grader does not compare SQL. It evaluates whether the agent’s action produces the correct outcome in the hidden simulation state.
Every task includes an executable reference solution that demonstrates solvability using only the warehouse. This proves the task is not impossible and provides a correctness baseline.
What Breaks in Multi-Table Workflows
The benchmark exposes four failure modes:
| Failure Mode | Example | Root Cause |
|---|---|---|
| Schema navigation | Agent queries wrong table or misses join path | 235 tables exceed context window or reasoning depth |
| Statistical reasoning | Agent calculates median incorrectly or misinterprets percentile | LLMs struggle with multi-step numerical logic |
| State reconstruction | Agent assumes data completeness when records are missing | No explicit signal that warehouse is incomplete |
| Action formulation | Agent outputs correct analysis but files malformed action | Disconnect between reasoning and execution interface |
The strongest models fail most often on state reconstruction. They treat the warehouse as authoritative when it is actually a partial, denormalized view of the simulator’s hidden state.
Observability and Debugging
Because the grader scores actions in a simulator, you get deterministic feedback. If an agent bans the wrong accounts, you can replay the simulation, inspect the agent’s queries, and trace where reasoning diverged from ground truth.
This is harder with traditional benchmarks. If an agent’s SQL output does not match the answer key, you cannot tell whether the agent is wrong or the answer key is wrong without manual review.
Argo-Bench provides:
- Execution logs for every query the agent ran
- Intermediate state snapshots showing what the agent knew at each step
- Simulator diffs comparing the agent’s action to the reference solution’s action
This makes it possible to debug multi-step reasoning failures, not just query syntax errors.
State Management Requirements
Agents need to maintain context across multiple queries and analyses. A typical task requires:
- Exploring 10-15 tables to understand schema relationships
- Running 3-5 exploratory queries to validate assumptions
- Performing statistical aggregations (percentiles, moving averages, cohort analysis)
- Formulating an action based on the results
This implies:
- Persistent memory to track which tables have been explored
- Intermediate result storage for multi-step calculations
- Hypothesis tracking to avoid re-querying the same data
- Action staging to validate the action before submission
Most agents tested in the paper use stateless prompting with full conversation history. This works for simple tasks but fails when reasoning depth exceeds context window limits.
Security and Isolation Boundaries
The benchmark runs agents against a read-only warehouse. Agents cannot modify data, only file actions through a controlled API.
This separation is critical for production deployment. Enterprise data agents should never have write access to the warehouse. Instead, they should:
- Query the warehouse read-only
- Propose actions through a staging API
- Wait for human or automated approval
- Execute actions in a separate transaction layer
Argo-Bench enforces this boundary by design. The agent sees the warehouse but acts through the simulator’s API, which validates and scores each action.
Code Example: Reference Solution Pattern
Here is a simplified reference solution for a fraud detection task:
# Step 1: Identify accounts with suspicious order patterns
suspicious_accounts = db.query("""
SELECT customer_id, COUNT(DISTINCT zip_code) as zip_count
FROM orders
WHERE order_date >= CURRENT_DATE - INTERVAL '24 hours'
GROUP BY customer_id
HAVING COUNT(DISTINCT zip_code) > 10
""")
# Step 2: Cross-reference with payment anomalies
fraud_candidates = db.query("""
SELECT DISTINCT o.customer_id
FROM orders o
JOIN payment_methods pm ON o.payment_id = pm.id
WHERE o.customer_id IN ({})
AND pm.card_country != o.delivery_country
""".format(','.join(str(id) for id in suspicious_accounts['customer_id'])))
# Step 3: File ban action
action = {
"type": "ban_accounts",
"account_ids": fraud_candidates['customer_id'].tolist(),
"reason": "Multi-zip + cross-border payment anomaly"
}
submit_action(action)
The reference solution demonstrates the workflow: explore, reason, act. The grader scores whether the banned accounts match the simulator’s ground-truth fraud list.
Deployment Considerations
Running Argo-Bench requires:
- Warehouse infrastructure (235 tables, 7.5 billion rows)
- Simulator runtime (to score actions and provide ground truth)
- Agent execution environment (isolated from production data)
The authors provide:
- Parquet exports of the warehouse (downloadable)
- Simulator code (open source)
- Grading harness (deterministic scoring)
For production use, you would replace the simulated warehouse with your actual ERP schema and build a custom grader that validates actions against business rules instead of a simulator.
Likely Failure Modes
Agents will fail in production when:
- Schema drift changes table relationships without updating the agent’s schema knowledge
- Data quality issues introduce nulls or duplicates that break statistical assumptions
- Action APIs change and the agent continues using deprecated endpoints
- Human approval delays cause the agent’s analysis to become stale before the action executes
Argo-Bench does not model these failure modes. It assumes a static schema, clean data, and immediate action execution. Real deployments need monitoring for schema changes, data quality checks, and staleness detection.
Technical Verdict
Use Argo-Bench when:
- You are building data agents that must navigate complex ERP schemas
- You need to evaluate multi-step reasoning, not just query generation
- You want deterministic grading based on action outcomes, not SQL string matching
- You are designing eval harnesses for enterprise automation workflows
Avoid Argo-Bench when:
- You only need to test single-query text-to-SQL generation
- Your data fits in a few tables and does not require multi-step reasoning
- You cannot run a 7.5 billion row warehouse locally or in CI
- You need benchmarks that model schema drift, data quality issues, or approval workflows
The benchmark is most useful for teams building agents that orchestrate multi-stage analytics pipelines. It exposes the gap between generating SQL and executing data-driven actions at enterprise scale.
Source Links
- Argo-Bench Paper (ArXiv)
- Code Repository (see paper for link)
- Dataset Download (see paper for link)