Introduction

OpenSRE is an open-source AI SRE platform that investigates production incidents autonomously. When an alert fires, OpenSRE's AI agents gather context from your observability stack, reason about root causes, and produce a detailed incident report — the way an experienced SRE would, but faster and around the clock.

Who is OpenSRE for?

OpenSRE is built for teams that are tired of manual, repetitive incident investigation:

  • Platform engineers and SREs who spend too much time on routine investigations
  • On-call engineers who need help at 3 AM when cognitive load is highest
  • Engineering managers who want to reduce MTTR and reliance on tribal knowledge
  • DevOps teams building their observability practice

Key Capabilities

Autonomous Incident Investigation

When you start an investigation — from Slack, Microsoft Teams, or the web console — a root investigator agent (built on the Claude Agent SDK) plans the work, loads integration skills on demand, and can dispatch specialist subagents in parallel for independent sub-questions (checking Kubernetes, querying metrics, reading logs and traces). Progress streams back to the client in real time as the agent thinks, calls tools, and produces its final report.

Episodic Memory System

OpenSRE remembers past investigations. After every conversation, it extracts structured metadata — issue type, affected components, root cause, resolution — and stores it as an episode in Neo4j. Recall is agent-driven: once the agent has concrete symptoms (an error, a failing service, a stack trace), it searches memory for similar past episodes and can pull in a synthesized strategy playbook when enough similar cases exist. Nothing is pre-loaded blindly on a vague alert — see Episodic Memory for how this actually works end to end.

Knowledge Graph

OpenSRE keeps service topology — dependencies and blast radius — in the same Neo4j instance as episodic memory. The agent queries it on demand once an affected service is known, to reason about what else might be impacted. Investigations continue normally if the graph is empty or unavailable; it's optional context, not a hard dependency.

51 Investigation Skills

OpenSRE ships with 51 built-in investigation skills covering Kubernetes and cloud infrastructure, observability (Prometheus, Grafana, Datadog, Elastic, Splunk, and more), incident and alerting tools, databases, version control, and ticketing. Skills are loaded progressively — the agent sees lightweight metadata for all of them, then loads the full methodology for the ones it actually decides to use.

Integrations

Works with: Kubernetes, Prometheus, Grafana, Datadog, Elasticsearch, Splunk, New Relic, Coralogix, Jaeger, Honeycomb, Sentry, PagerDuty, Opsgenie, GitHub, GitLab, Bitbucket, Jira, Linear, Confluence, and more — see Integrations for the full list.

Architecture at a Glance

OpenSRE's agent runtime is built on the Claude Agent SDK — there's no separate graph-orchestration layer. Alerts and questions enter via Slack, Microsoft Teams, or the web console; sre-agent runs the investigation and streams results back over Server-Sent Events.

Slack / Teams  →  bot  →  sre-agent (Claude Agent SDK)
Web UI ──────────────→        │
                          ┌────┴────┐
                          │    │    │
                       Memory Skills KG
                       (Neo4j)     (Neo4j)

For a deep dive into the architecture, see Architecture.

Open Source

OpenSRE is released under the Apache 2.0 license. Self-host it in your own infrastructure. Your data, your control.

Get started →