Architecture

OpenSRE investigates incidents with an AI agent runtime built on the Claude Agent SDK — not a hand-rolled graph orchestration layer. A root investigator agent plans the work, loads integration skills on demand, can dispatch specialist subagents for independent sub-questions, and streams progress back to clients in real time. Episodic memory and a service topology knowledge graph both live in Neo4j; configuration and full investigation traces live in PostgreSQL.

System Overview

     Web UI              Slack Bot              Teams Bot
     (Next.js)         (Socket Mode)         (Microsoft Teams)
         \                  |                    /
          \                 |                   /
           v                v                  v
         ┌─────────────────────────────────────────┐
         │                sre-agent                 │
         └──────────────────┬──────────────────────┘
                             │
         ┌───────────────────┼──────────────┬─────────────────┐
         v                   v              v                 v
    ┌──────────┐       ┌──────────┐  ┌──────────┐   ┌─────────────┐
    │ config-  │       │ Postgres │  │  Neo4j   │   │ LiteLLM     │
    │ service  │       │ agent    │  │ memory + │   │ (optional   │
    │          │       │ runs +   │  │ knowledge│   │  proxy)     │
    │          │       │ config   │  │ graph    │   │             │
    └──────────┘       └──────────┘  └──────────┘   └─────────────┘

Entry points

  • Web UI — admin console and investigation UI; streams SSE from sre-agent. Optional Microsoft Entra SSO (Entra SSO).
  • Slack bot — Socket Mode; @mentions and threads; streams the same SSE protocol.
  • Teams bot — Microsoft Teams channels or DMs; streams the same SSE protocol.
  • REST API — POST /investigate and related thread endpoints on sre-agent, used by the clients above.

Investigation Flow

  1. A user starts an investigation from the web console, Slack, Teams, or the REST API.
  2. sre-agent creates or resumes a thread and opens an InteractiveAgentSession (Claude Agent SDK).
  3. The root agent (typically investigator, falling back to planner) plans the work, loads skills on demand, and dispatches specialists via the SDK Task tool.
  4. After concrete symptoms are known — error text, a failing service, a stack trace — the agent may invoke memory-search and infrastructure-neo4j. There is no automatic pre-injection of past episodes or topology into the first prompt; recall is agent-driven.
  5. Tool calls, thoughts, subagent progress, and the final answer stream to the client as Server-Sent Events.
  6. On turn completion, a lifecycle hook upserts a Neo4j :Episode (one per conversation) and persists the full tool trace to PostgreSQL for replay in the web UI.

There is no LangGraph, no fixed pipeline of nodes. The SDK session itself is the runtime.

The Agent SDK Runtime

The investigation engine is InteractiveAgentSession, which wraps the SDK's ClaudeSDKClient and maps SDK messages onto OpenSRE's own SSE event protocol.

Root agent and subagents

  • The root agent is resolved from team config (investigator preferred, then planner). Its system prompt comes from a nested config path (agents.{agent_id}.prompt.system) so it can be customized per team without touching other teams.
  • Specialist subagents are registered as SDK AgentDefinition objects, built from the team's configured agent topology. The root delegates to them via the SDK Task tool — optionally in the background, so multiple specialists can work concurrently.
  • Skills live as SKILL.md files plus optional scripts. The SDK Skill tool loads lightweight metadata for every skill up front (roughly 100 tokens each) and the full skill content only when the agent actually chooses to use it. Skill scripts then run via Bash under that loaded skill.

Message stream drain

The session follows the SDK-recommended pattern rather than the simpler single-response API:

  1. Start client.query() with a streaming user-message generator, running concurrently with the receive loop.
  2. Drain the response with receive_messages() — not receive_response() alone.

receive_messages() doesn't stop automatically at the first result. That matters because a turn isn't necessarily over just because the SDK reports one — background subagents may still be running.

Background subagents

Subagents started in the background emit their own lifecycle messages on the SDK stream. While any are still outstanding:

  • Interim root text is emitted to the client as a thought event, not a final answer.
  • Clients receive a background_waiting event with the pending task IDs.
  • The parent agent automatically continues once every outstanding task clears — no follow-up prompt required from the user.

Mid-run message queue

You can add context to a running investigation without waiting for it to finish:

  • API: POST /threads/{thread_id}/queue-message
  • SSE: a message_queued event when queued text is merged into the agent's turn

Queued messages are debounced for a fraction of a second and merged into a single numbered guidance block, so several quick follow-ups don't each interrupt the agent separately. Any remaining queue is flushed at the turn boundary so nothing gets silently dropped.

SSE event types (high level)

TypePurpose
thoughtAgent reasoning / interim narration
tool_start / tool_endSkill, Bash, Task, and other tool lifecycle
task_started / task_notificationBackground subagent lifecycle
background_waitingParent waiting on outstanding background tasks
message_queuedMid-run user message accepted or consumed
questionAgent asking a clarifying question
resultTerminal answer for the turn
errorFailure or timeout

Investigations run under a wall-clock timeout independent of the SDK's own turn-count limit — the two are separate caps and raising one doesn't raise the other.

Deployment Modes

Default (make dev / self-host)Sandbox mode
Serverserver_simple.pyserver.py
Agent processSame process as the APIPer-thread isolated pod
IsolationTrusted local / single-tenant useFilesystem and network isolation
SkillsCopied to a per-thread workspaceBaked into the sandbox image

The default mode is what make dev runs and what typical self-hosting uses. Sandbox mode is a separate, production-oriented stack for stronger isolation between concurrent investigations and isn't required to run OpenSRE.

Episodic Memory

OpenSRE stores a condensed summary of every conversation as a Neo4j :Episode, and recalls similar past episodes on demand via semantic vector search — the agent decides when to search, based on evidence gathered so far, not a fixed rule fired on every alert. See Episodic Memory for the full model, including how strategy playbooks get synthesized from repeated episodes.

Knowledge Graph

Neo4j also holds service topology — dependencies and blast radius — in the same database as episodic memory, queried agent-driven via a dedicated skill once an affected service is known. See Knowledge Graph for what's tracked and how it's populated today.

Skills

Skills are the primary integration surface: a SKILL.md methodology doc plus optional scripts, organized by domain (Kubernetes and cloud, observability, databases, incident tooling, version control, and more). See Investigation Skills for the full catalog and how to add your own.

Configuration

config-service is the control plane:

  • Hierarchical org → team configuration with deep merge (dicts merge, lists replace).
  • Team tokens for runtime auth; admin tokens for the configuration UI/API.
  • Agent prompts, model choice, subagent topology, and skill enablement, all team-scoped.
  • Integration credentials, pulled from environment variables via ${VAR} substitution — never hardcoded in config.

See Configuration for the full reference.

Local Development Stack

cp .env.example .env   # set ANTHROPIC_API_KEY
make dev               # core stack
make dev-slack         # core + Slack bot
make dev-teams         # core + Teams bot
ServicePortRole
Web UI3002Next.js console
sre-agent8001Investigation API + SSE
config-service8081Config, tokens, agent-run storage
PostgreSQL5433Config DB + agent runs
Neo4j Browser / Bolt7475 / 7688Memory + topology graph
Slack bot—Optional (--profile slack / make dev-slack)
Teams bot3978Optional (--profile teams / make dev-teams)
LiteLLM4001Optional (--profile litellm)

By default the agent calls Anthropic directly (ANTHROPIC_API_KEY). To use OpenRouter or another provider, start the LiteLLM profile and set ANTHROPIC_BASE_URL to the proxy.