mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

AI Agents

OpenAI and Ironclad: Turning Contract Workflows Into Agent Evals

How OpenAI and Ironclad built a training loop that uses real contract workflows as both agent training data and reproducible evaluation benchmarks.

Source: openai.com
OpenAI and Ironclad: Turning Contract Workflows Into Agent Evals

OpenAI and Ironclad just published a case study on using production contract workflows as both training data and evaluation benchmarks for computer-use agents. This is not a demo. It is a partnership where a SaaS company opens its workflow engine to become an agent training ground.

The plumbing question is simple: how do you turn a multi-step contract approval flow into a reproducible eval without leaking customer data, and how do you keep that eval valid when the underlying product changes?

Why Contract Workflows Matter for Agent Evals

Most computer-use benchmarks are synthetic. They simulate browser tasks or API calls in controlled environments. Ironclad’s contract lifecycle management platform offers something different: real multi-step workflows with approval chains, redlining, negotiation loops, and conditional branching.

These workflows are stateful. A contract might move from draft to legal review, back to sales for edits, then to finance for approval. Each step has different permissions, different UI surfaces, and different failure modes. If an agent can navigate this, it can navigate most SaaS tools.

The partnership gives OpenAI access to:

  • Real workflow graphs with conditional logic
  • Multi-role approval chains
  • Document versioning and redlining
  • Integration points with external systems
  • Actual failure cases from production use

The Eval Infrastructure Challenge

Turning a production workflow into an agent eval requires solving three problems:

1. Data sanitization
You cannot train on customer contracts. Ironclad must generate synthetic contracts that preserve workflow complexity without exposing real terms, parties, or negotiation history. This means:

  • Template-based contract generation with realistic clauses
  • Synthetic approval chains that mirror real org structures
  • Redline patterns extracted from anonymized edits

2. Reproducibility
A contract workflow is not deterministic. Different users take different paths. To make this an eval, you need:

  • Snapshot environments with fixed UI state
  • Versioned workflow definitions
  • Deterministic approval logic (no human-in-the-loop during eval runs)

3. Version drift
Ironclad ships product updates. When the UI changes, the eval breaks. The infrastructure must:

  • Track product version alongside eval version
  • Maintain backward-compatible eval snapshots
  • Flag when UI changes invalidate existing test cases

Architecture: Workflow Capture to Agent Execution

Here is the likely flow:

┌─────────────────┐
│ Ironclad Prod   │
│ Workflow Engine │
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Workflow Logger │  ← Captures state transitions, UI events, API calls
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Sanitization    │  ← Strips PII, generates synthetic contracts
│ Pipeline        │
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Eval Snapshot   │  ← Versioned environment + workflow definition
│ Generator       │
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Agent Executor  │  ← OpenAI agent runs against snapshot
└────────┬────────┘
         │
         ▼
┌─────────────────┐
│ Failure Labeler │  ← Human annotators classify agent errors
└─────────────────┘

The workflow logger is the critical piece. It must capture:

  • DOM snapshots at each step
  • API request/response pairs
  • User intent signals (what the human was trying to do)
  • Success criteria (what “done” looks like for each task)

Failure Mode Labeling

When an agent fails on a contract task, someone has to classify why. This is not automated. The labeling taxonomy likely includes:

Failure TypeExampleFix Strategy
Navigation errorClicked wrong buttonImprove UI element detection
State misreadThought contract was approved when pendingBetter state extraction from DOM
Logic errorSkipped required approval stepRefine workflow graph understanding
TimeoutTook too long on redline comparisonOptimize document parsing
Permission boundaryTried to approve without authorityTeach role-based access model

Ironclad employees likely do the initial labeling. OpenAI uses those labels to retrain. The loop tightens over time.

Code Snippet: Workflow Snapshot Schema

A reproducible eval needs a snapshot format. Here is a plausible schema:

from dataclasses import dataclass
from typing import List, Dict, Any

@dataclass
class WorkflowSnapshot:
    version: str  # Ironclad product version
    workflow_id: str
    initial_state: Dict[str, Any]  # Contract metadata, approvers, etc.
    steps: List[Dict[str, Any]]  # Ordered list of expected actions
    success_criteria: Dict[str, Any]  # What "done" looks like
    ui_snapshots: List[str]  # DOM or screenshot hashes per step
    
    def validate_agent_run(self, agent_trace: List[Dict]) -> bool:
        """
        Compare agent actions against expected workflow steps.
        Returns True if agent reached success criteria.
        """
        for i, expected_step in enumerate(self.steps):
            if i >= len(agent_trace):
                return False  # Agent stopped early
            
            actual = agent_trace[i]
            if not self._step_matches(expected_step, actual):
                return False
        
        return self._check_success(agent_trace[-1])
    
    def _step_matches(self, expected: Dict, actual: Dict) -> bool:
        # Fuzzy match on action type and target element
        return (
            expected["action_type"] == actual["action_type"] and
            expected["target_element"] in actual["dom_path"]
        )
    
    def _check_success(self, final_state: Dict) -> bool:
        # Check if contract reached expected status
        return final_state.get("contract_status") == self.success_criteria.get("status")

This snapshot is versioned. When Ironclad ships a UI update, old snapshots remain valid for regression testing. New snapshots get generated for the updated product.

Observability: What Gets Logged

To debug agent failures, you need full trace data:

  • Agent reasoning log: Why it chose each action
  • DOM diff per step: What changed in the UI
  • API call log: Backend state transitions
  • Screenshot timeline: Visual proof of what the agent saw
  • Latency breakdown: Time spent on vision, planning, execution

The observability stack must correlate these streams. If an agent clicks the wrong button, you need to see:

  1. What the DOM looked like
  2. What the vision model extracted
  3. What the planner decided
  4. What the executor sent to the browser

Without this correlation, failure labeling becomes guesswork.

Version Control for Evals

When Ironclad changes a workflow (new approval step, different UI layout), the eval suite must adapt. Two strategies:

Pinned snapshots
Keep old product versions running in isolated environments. Agents train against v1.2, v1.3, v1.4 simultaneously. This catches regressions but requires infrastructure to run multiple product versions.

Adaptive evals
Update eval snapshots when the product changes. Mark old snapshots as deprecated. This keeps evals current but loses historical comparison.

Most teams use a hybrid: pin critical workflows, adapt the rest.

Security Boundaries

Ironclad cannot give OpenAI direct access to production. The sanitization pipeline must run inside Ironclad’s VPC. The output (synthetic contracts and workflow snapshots) gets exported to OpenAI’s training environment.

Key boundaries:

  • Data exfiltration: No customer data leaves Ironclad
  • Workflow isolation: Eval runs cannot touch production contracts
  • Access control: OpenAI agents run with limited permissions, cannot escalate

The sanitization pipeline is the trust boundary. If it leaks PII, the partnership fails.

Trade-Offs: Production Workflows vs. Synthetic Benchmarks

DimensionProduction WorkflowsSynthetic Benchmarks
RealismHigh (real complexity)Low (simplified tasks)
Data privacyHard (requires sanitization)Easy (no real data)
ReproducibilityHard (version drift)Easy (static)
Failure diversityHigh (real edge cases)Low (designed cases)
Setup costHigh (partnership required)Low (build in-house)

Production workflows expose real failure modes. Synthetic benchmarks are easier to control. The best eval suites use both.

Technical Verdict

Use this approach when:

  • You need agents to handle real SaaS complexity, not toy tasks
  • You have a SaaS partner willing to expose workflow internals
  • You can build a sanitization pipeline that strips PII reliably
  • You have labeling capacity to classify agent failures
  • You need eval data that evolves with product updates

Avoid this approach when:

  • You are early in agent development and need fast iteration
  • You cannot guarantee data privacy in the sanitization step
  • Your SaaS partner ships breaking changes weekly
  • You lack infrastructure to version-control product snapshots
  • Synthetic evals give you enough signal

The OpenAI-Ironclad partnership shows that production workflows can become agent training grounds. The plumbing is non-trivial: sanitization, versioning, failure labeling, and observability all need custom infrastructure. But the payoff is agents that handle real work, not just demos.