mech.app

The mech.app newsletter

Agentic AI, minus the noise.

Get practical field notes on AI agents, automation, developer tools and security delivered to your inbox.

No spam. Unsubscribe anytime.

Dev Tools

OpenSRE: Building AI SRE Agents with 60+ Tool Integrations and Training Environments

How OpenSRE orchestrates observability tools, defines safe remediation workflows, and simulates incidents for agent training without touching production.

Source: github.com
OpenSRE: Building AI SRE Agents with 60+ Tool Integrations and Training Environments

OpenSRE is an open-source framework for building AI agents that answer production questions and execute remediation tasks across your observability stack. It connects 60+ tools (Datadog, Grafana, Slack, PagerDuty) into a unified agent runtime and includes a training environment for simulating incidents without touching live systems.

The project (11,359 stars, trending #9 on GitHub Python) addresses a specific gap: most agent frameworks treat SRE tooling as an afterthought, bolting on integrations without workflow boundaries or safety controls. OpenSRE inverts this by making tool orchestration, permission scoping, and evaluation infrastructure first-class concerns.

Architecture: Workflow Definitions Over Brittle Integrations

OpenSRE avoids the typical integration mess by separating tool connectors from workflow logic. Each integration is a thin adapter that exposes a standard interface. Workflows define how agents combine tools, what permissions they need, and what guardrails apply.

Core components:

  • Tool Registry: Catalog of available integrations with capability metadata (read-only vs. write, rate limits, required credentials)
  • Workflow Engine: Orchestrates multi-step agent tasks, maintains state across tool calls, enforces permission boundaries
  • Training Environment: Simulates production incidents using snapshots of real observability data, allows agents to practice without blast radius
  • Evaluation Loop: Measures agent performance on incident resolution time, false positive rate, and tool usage efficiency

The workflow engine is the critical piece. It tracks which tools an agent has called, what data it has seen, and what actions it is allowed to take next. This prevents agents from escalating privileges or executing destructive operations outside their defined scope.

# Example workflow definition (simplified)
workflow = Workflow(
    name="investigate_high_latency",
    triggers=["alert:latency_p99_exceeded"],
    tools=[
        Tool("datadog.metrics", permissions=["read"]),
        Tool("grafana.dashboards", permissions=["read"]),
        Tool("slack.notify", permissions=["write"], rate_limit="5/min"),
        Tool("kubernetes.restart_pod", permissions=["write"], requires_approval=True)
    ],
    max_steps=10,
    timeout_minutes=15
)

This structure makes it explicit what an agent can do. The requires_approval=True flag on destructive operations creates a human-in-the-loop gate before execution.

Training and Evaluation: Incident Simulation Without Production Risk

Most agent frameworks skip training infrastructure entirely. OpenSRE treats it as a core requirement. The training environment ingests snapshots of observability data (metrics, logs, traces) from past incidents and replays them in a sandboxed runtime.

How incident simulation works:

  1. Snapshot Capture: Export metrics, logs, and alert states from a real incident window
  2. Environment Replay: Load snapshot into isolated environment with mock tool endpoints
  3. Agent Execution: Run agent workflows against replayed data, track tool calls and decisions
  4. Evaluation: Compare agent actions to known-good remediation steps, measure time to resolution

The evaluation loop generates metrics on agent performance:

  • Resolution accuracy: Did the agent identify the correct root cause?
  • Tool efficiency: How many unnecessary tool calls were made?
  • Safety violations: Did the agent attempt actions outside its permission scope?

This feedback loop is what separates a toy demo from a production-ready SRE agent. You can iterate on workflow definitions and agent prompts without risking live infrastructure.

State Management: Maintaining Context Across Multi-Step Investigations

Incident investigations span hours or days. An agent needs to remember what it has already checked, what hypotheses it has ruled out, and what remediation steps are in progress.

OpenSRE uses a persistent session store to track agent state:

  • Investigation graph: Nodes represent hypotheses, edges represent tool calls that tested them
  • Tool call history: Ordered log of every API request, response, and decision point
  • Approval queue: Pending actions that require human confirmation before execution

When an agent resumes work on an incident, it rehydrates state from the session store and continues where it left off. This avoids redundant queries and prevents agents from forgetting context when they hit timeouts or rate limits.

Permission Scoping and Blast Radius Control

Giving an agent write access to production systems is the hard part. OpenSRE handles this with layered permission controls:

Control LayerMechanismExample
Tool-levelEach integration declares required permissionskubernetes.restart_pod requires write:pods
Workflow-levelWorkflows specify which tools are allowedOnly investigate_high_latency can restart pods
Approval gatesDestructive actions block until human confirmsSlack message with approve/deny buttons
Rate limitsPer-tool throttling prevents runaway executionMax 5 pod restarts per 10 minutes
Audit logImmutable record of every agent actionQueryable log for post-incident review

The approval gate mechanism is critical. When an agent wants to execute a destructive operation, it pauses execution and sends a notification to Slack (or PagerDuty, or email) with context about why the action is needed. A human reviews the request and approves or denies it. The agent resumes only after approval.

Observability of Agents: Tracing Agent Decisions

OpenSRE instruments the agent runtime itself. Every tool call, decision point, and state transition is traced and exported to your observability stack.

What gets traced:

  • Tool call spans: Duration, input parameters, response payload, error states
  • Decision logs: Which hypothesis the agent is testing, what data informed the decision
  • Workflow transitions: When the agent moves from investigation to remediation
  • Failure modes: Timeout errors, permission denials, rate limit hits

This tracing layer is essential for debugging agent behavior. When an agent makes a bad decision, you can replay the exact sequence of tool calls and see where it went wrong.

The framework exports traces in OpenTelemetry format, so they flow into your existing observability pipeline (Jaeger, Honeycomb, Datadog APM).

Deployment Shape and Failure Modes

OpenSRE runs as a long-lived service, not a serverless function. Agents need persistent state and the ability to resume work after interruptions.

Typical deployment:

  • Agent runtime: Python service running in Kubernetes, manages workflow execution and state persistence
  • Tool adapters: Separate processes or sidecars for each integration, isolates credential management
  • State store: PostgreSQL or Redis for session persistence, supports multi-agent concurrency
  • Message queue: RabbitMQ or Kafka for async tool calls, decouples agent logic from slow APIs

Common failure modes:

  • Tool API outages: Agent retries with exponential backoff, falls back to read-only mode if write APIs are down
  • Permission drift: Tool credentials expire or permissions change, agent logs error and notifies on-call
  • State corruption: Session store becomes inconsistent, agent resets to last known-good checkpoint
  • Runaway execution: Agent hits rate limit or max step count, workflow terminates and alerts human

The framework includes circuit breakers for each tool integration. If a tool API starts returning errors at high rate, the circuit breaker opens and the agent stops calling that tool until the error rate drops.

Integration Catalog: What 60+ Tools Actually Means

The 60+ tool count includes both observability platforms and incident management systems. Here is a breakdown of major categories:

  • Metrics and monitoring: Datadog, Grafana, Prometheus, New Relic, Dynatrace
  • Logging: Splunk, Elasticsearch, Loki, CloudWatch Logs
  • Tracing: Jaeger, Zipkin, Honeycomb, Lightstep
  • Incident management: PagerDuty, Opsgenie, VictorOps, Incident.io
  • Communication: Slack, Microsoft Teams, Discord, email
  • Infrastructure: Kubernetes, AWS, GCP, Azure, Terraform
  • Databases: PostgreSQL, MySQL, MongoDB, Redis (for querying slow query logs, connection counts)

Each integration is a Python module that implements a standard interface. Adding a new tool means writing an adapter class and registering it in the tool catalog.

When to Use OpenSRE vs. Vendor Platforms

OpenSRE makes sense when you need full control over agent behavior and want to avoid vendor lock-in. It is a framework, not a managed service. You own the deployment, the training data, and the workflow definitions.

Use OpenSRE when:

  • You already run a diverse observability stack and need a unified agent layer
  • You want to train agents on your own incident data without sending it to a third party
  • You need custom workflows that vendor platforms do not support
  • You have the engineering capacity to operate the agent runtime and debug failures

Avoid OpenSRE when:

  • You want a turnkey solution with zero operational overhead
  • Your team is small and cannot dedicate resources to agent training and evaluation
  • You need enterprise support and SLAs
  • Your observability stack is simple enough that manual runbooks still work

The framework is in public alpha (v0.1), so expect rough edges. The core workflow engine and tool registry are stable, but the training environment is still evolving.

Technical Verdict

OpenSRE is the first open-source SRE agent framework that treats training and evaluation as first-class concerns. The workflow engine, permission scoping, and incident simulation infrastructure are production-grade. The integration catalog is broad enough to cover most observability stacks.

The main trade-off is operational complexity. You are running a stateful agent runtime, managing tool credentials, and maintaining a training environment. This is not a weekend side project. It is infrastructure that requires ongoing investment.

If you are building an internal platform team and want AI agents that answer production questions on your own terms, OpenSRE is the most complete open-source option available. If you need something that works out of the box with minimal setup, look at vendor platforms like Blameless or Rootly.

The project is worth watching. The focus on training infrastructure and permission boundaries addresses real gaps in the agent ecosystem. As the framework matures, expect to see more teams using it as the foundation for custom SRE automation.