OpenSRE is an open-source framework for building AI agents that answer production questions and execute remediation tasks across your observability stack. It connects 60+ tools (Datadog, Grafana, Slack, PagerDuty) into a unified agent runtime and includes a training environment for simulating incidents without touching live systems.
The project (11,359 stars, trending #9 on GitHub Python) addresses a specific gap: most agent frameworks treat SRE tooling as an afterthought, bolting on integrations without workflow boundaries or safety controls. OpenSRE inverts this by making tool orchestration, permission scoping, and evaluation infrastructure first-class concerns.
Architecture: Workflow Definitions Over Brittle Integrations
OpenSRE avoids the typical integration mess by separating tool connectors from workflow logic. Each integration is a thin adapter that exposes a standard interface. Workflows define how agents combine tools, what permissions they need, and what guardrails apply.
Core components:
- Tool Registry: Catalog of available integrations with capability metadata (read-only vs. write, rate limits, required credentials)
- Workflow Engine: Orchestrates multi-step agent tasks, maintains state across tool calls, enforces permission boundaries
- Training Environment: Simulates production incidents using snapshots of real observability data, allows agents to practice without blast radius
- Evaluation Loop: Measures agent performance on incident resolution time, false positive rate, and tool usage efficiency
The workflow engine is the critical piece. It tracks which tools an agent has called, what data it has seen, and what actions it is allowed to take next. This prevents agents from escalating privileges or executing destructive operations outside their defined scope.
# Example workflow definition (simplified)
workflow = Workflow(
name="investigate_high_latency",
triggers=["alert:latency_p99_exceeded"],
tools=[
Tool("datadog.metrics", permissions=["read"]),
Tool("grafana.dashboards", permissions=["read"]),
Tool("slack.notify", permissions=["write"], rate_limit="5/min"),
Tool("kubernetes.restart_pod", permissions=["write"], requires_approval=True)
],
max_steps=10,
timeout_minutes=15
)
This structure makes it explicit what an agent can do. The requires_approval=True flag on destructive operations creates a human-in-the-loop gate before execution.
Training and Evaluation: Incident Simulation Without Production Risk
Most agent frameworks skip training infrastructure entirely. OpenSRE treats it as a core requirement. The training environment ingests snapshots of observability data (metrics, logs, traces) from past incidents and replays them in a sandboxed runtime.
How incident simulation works:
- Snapshot Capture: Export metrics, logs, and alert states from a real incident window
- Environment Replay: Load snapshot into isolated environment with mock tool endpoints
- Agent Execution: Run agent workflows against replayed data, track tool calls and decisions
- Evaluation: Compare agent actions to known-good remediation steps, measure time to resolution
The evaluation loop generates metrics on agent performance:
- Resolution accuracy: Did the agent identify the correct root cause?
- Tool efficiency: How many unnecessary tool calls were made?
- Safety violations: Did the agent attempt actions outside its permission scope?
This feedback loop is what separates a toy demo from a production-ready SRE agent. You can iterate on workflow definitions and agent prompts without risking live infrastructure.
State Management: Maintaining Context Across Multi-Step Investigations
Incident investigations span hours or days. An agent needs to remember what it has already checked, what hypotheses it has ruled out, and what remediation steps are in progress.
OpenSRE uses a persistent session store to track agent state:
- Investigation graph: Nodes represent hypotheses, edges represent tool calls that tested them
- Tool call history: Ordered log of every API request, response, and decision point
- Approval queue: Pending actions that require human confirmation before execution
When an agent resumes work on an incident, it rehydrates state from the session store and continues where it left off. This avoids redundant queries and prevents agents from forgetting context when they hit timeouts or rate limits.
Permission Scoping and Blast Radius Control
Giving an agent write access to production systems is the hard part. OpenSRE handles this with layered permission controls:
| Control Layer | Mechanism | Example |
|---|---|---|
| Tool-level | Each integration declares required permissions | kubernetes.restart_pod requires write:pods |
| Workflow-level | Workflows specify which tools are allowed | Only investigate_high_latency can restart pods |
| Approval gates | Destructive actions block until human confirms | Slack message with approve/deny buttons |
| Rate limits | Per-tool throttling prevents runaway execution | Max 5 pod restarts per 10 minutes |
| Audit log | Immutable record of every agent action | Queryable log for post-incident review |
The approval gate mechanism is critical. When an agent wants to execute a destructive operation, it pauses execution and sends a notification to Slack (or PagerDuty, or email) with context about why the action is needed. A human reviews the request and approves or denies it. The agent resumes only after approval.
Observability of Agents: Tracing Agent Decisions
OpenSRE instruments the agent runtime itself. Every tool call, decision point, and state transition is traced and exported to your observability stack.
What gets traced:
- Tool call spans: Duration, input parameters, response payload, error states
- Decision logs: Which hypothesis the agent is testing, what data informed the decision
- Workflow transitions: When the agent moves from investigation to remediation
- Failure modes: Timeout errors, permission denials, rate limit hits
This tracing layer is essential for debugging agent behavior. When an agent makes a bad decision, you can replay the exact sequence of tool calls and see where it went wrong.
The framework exports traces in OpenTelemetry format, so they flow into your existing observability pipeline (Jaeger, Honeycomb, Datadog APM).
Deployment Shape and Failure Modes
OpenSRE runs as a long-lived service, not a serverless function. Agents need persistent state and the ability to resume work after interruptions.
Typical deployment:
- Agent runtime: Python service running in Kubernetes, manages workflow execution and state persistence
- Tool adapters: Separate processes or sidecars for each integration, isolates credential management
- State store: PostgreSQL or Redis for session persistence, supports multi-agent concurrency
- Message queue: RabbitMQ or Kafka for async tool calls, decouples agent logic from slow APIs
Common failure modes:
- Tool API outages: Agent retries with exponential backoff, falls back to read-only mode if write APIs are down
- Permission drift: Tool credentials expire or permissions change, agent logs error and notifies on-call
- State corruption: Session store becomes inconsistent, agent resets to last known-good checkpoint
- Runaway execution: Agent hits rate limit or max step count, workflow terminates and alerts human
The framework includes circuit breakers for each tool integration. If a tool API starts returning errors at high rate, the circuit breaker opens and the agent stops calling that tool until the error rate drops.
Integration Catalog: What 60+ Tools Actually Means
The 60+ tool count includes both observability platforms and incident management systems. Here is a breakdown of major categories:
- Metrics and monitoring: Datadog, Grafana, Prometheus, New Relic, Dynatrace
- Logging: Splunk, Elasticsearch, Loki, CloudWatch Logs
- Tracing: Jaeger, Zipkin, Honeycomb, Lightstep
- Incident management: PagerDuty, Opsgenie, VictorOps, Incident.io
- Communication: Slack, Microsoft Teams, Discord, email
- Infrastructure: Kubernetes, AWS, GCP, Azure, Terraform
- Databases: PostgreSQL, MySQL, MongoDB, Redis (for querying slow query logs, connection counts)
Each integration is a Python module that implements a standard interface. Adding a new tool means writing an adapter class and registering it in the tool catalog.
When to Use OpenSRE vs. Vendor Platforms
OpenSRE makes sense when you need full control over agent behavior and want to avoid vendor lock-in. It is a framework, not a managed service. You own the deployment, the training data, and the workflow definitions.
Use OpenSRE when:
- You already run a diverse observability stack and need a unified agent layer
- You want to train agents on your own incident data without sending it to a third party
- You need custom workflows that vendor platforms do not support
- You have the engineering capacity to operate the agent runtime and debug failures
Avoid OpenSRE when:
- You want a turnkey solution with zero operational overhead
- Your team is small and cannot dedicate resources to agent training and evaluation
- You need enterprise support and SLAs
- Your observability stack is simple enough that manual runbooks still work
The framework is in public alpha (v0.1), so expect rough edges. The core workflow engine and tool registry are stable, but the training environment is still evolving.
Technical Verdict
OpenSRE is the first open-source SRE agent framework that treats training and evaluation as first-class concerns. The workflow engine, permission scoping, and incident simulation infrastructure are production-grade. The integration catalog is broad enough to cover most observability stacks.
The main trade-off is operational complexity. You are running a stateful agent runtime, managing tool credentials, and maintaining a training environment. This is not a weekend side project. It is infrastructure that requires ongoing investment.
If you are building an internal platform team and want AI agents that answer production questions on your own terms, OpenSRE is the most complete open-source option available. If you need something that works out of the box with minimal setup, look at vendor platforms like Blameless or Rootly.
The project is worth watching. The focus on training infrastructure and permission boundaries addresses real gaps in the agent ecosystem. As the framework matures, expect to see more teams using it as the foundation for custom SRE automation.