Episodic Memory System

OpenSRE remembers past investigations and lets the agent recall them when they're relevant. After every conversation, it extracts structured metadata from the outcome and stores it as an episode in Neo4j. When a similar incident comes up later, the agent — not a fixed pipeline stage — decides to search that memory and pull in what worked before, the way an experienced SRE reaches for a pattern they've seen.

Why Episodic Memory Matters

Without it, every investigation starts from zero: no knowledge of past outages, no patterns to recognize, no proven approach to try first. With it, OpenSRE builds up:

  • Recognized patterns — similar issue types and affected components surfaced from past episodes
  • Recalled root causes — what was actually wrong last time, with the evidence that supported it
  • Synthesized strategies — a reusable playbook once enough similar episodes exist, not written by hand

Two Stores, Two Jobs

RecordContentsStore
Agent run traceEvery tool call and streamed output for one turn — the forensic record, shown in the web UI's Investigations viewPostgreSQL (config-service)
Investigation episodeCondensed, LLM-extracted summary: issue type, components, root cause, skills used, resolutionNeo4j :Episode

An episode is a derived summary, never a raw tool dump. It links back to the full trace by correlation_id (the conversation/thread) and to the latest turn's agent_run_id for deep-linking into the trace viewer.

Recall Is Agent-Driven, Not Automatic

This is the part most people get wrong about how OpenSRE's memory works: nothing is pre-loaded into the first prompt. At session start, the root agent's system prompt gets a short guidance block instructing it to search memory — but only after it has concrete evidence (an error, a failing service, a stack trace), not on a vague initial alert alone.

Session init
 └─ system prompt += memory guidance
      "After concrete evidence, search memory for similar
       past investigations via the memory-search skill."

Agent turn
 ├─ gathers evidence (logs, metrics, deploy history, …) via other skills
 ├─ decides it has enough to search
 └─ invokes memory-search with symptom + component + system
      e.g. "Maven OOM surveys-module test failure Jenkins"

A lightweight topology hint (see Knowledge Graph) can still be attached to the user prompt when a service name is detected — that's separate from episodic recall and carries no past-investigation content.

There's also a legacy pre-injection path, off by default, that would prepend similar episodes and a strategy block directly into the first user prompt before the agent does anything. It exists for rollback/comparison purposes; agent-driven recall via memory-search is the supported model.

What memory-search Returns

{
  "success": true,
  "result": {
    "episodes": [
      {
        "issue_type": "connection-pool-exhaustion",
        "root_cause": "...",
        "resolved": true,
        "summary": "...",
        "skills_used": ["infrastructure-kubernetes", "observability-datadog"],
        "score": 0.87
      }
    ],
    "strategy": "## Common Root Causes\n..."
  }
}
  • episodes are ranked by semantic similarity; the agent is guided to prefer resolved episodes with a matching root cause and useful skills_used.
  • strategy is a synthesized markdown playbook — present only when enough similar episodes exist (see below).
  • A failed or empty search returns {"success": false} and is non-blocking — the investigation just continues without it.

How an Episode Is Written

One :Episode node per conversation (correlation_id), upserted at the end of every turn — the latest turn's conclusion wins.

FieldRole
issue_type, issue_descriptionClassification, in your terms — not tied to a fixed alert taxonomy
componentsTyped affected subjects, e.g. {type: "service", name: "payments-api"}
skills_usedAccumulated across every turn in the conversation
key_findingsCompact tool-derived findings, replaced each turn
root_cause, summary, resolvedThe narrative outcome — resolved means a root cause was identified with evidence, not that the environment was fixed
effectiveness_scoreHeuristic: ~0.8 when resolved with a root cause, ~0.4 when resolved only, ~0.1 otherwise — a ranking tie-break, not a validated quality metric
embedding384-dimensional vector for semantic search

Example: turn 1 stores root_cause: "Redis cache"; turn 2 corrects it to "missing DB index"; turn 3 sets resolved: true. The episode reflects the latest conclusion — wrong intermediate guesses don't stick around.

If extraction fails to produce a usable summary, the episode is still listed (with a "couldn't summarize" indicator and a retry action in the web UI) but excluded from semantic search until it's retried successfully.

Embeddings

Embeddings are computed with fastembed (ONNX, self-hosted — no external embedding API), by default BAAI/bge-small-en-v1.5 at 384 dimensions, over the combined issue_type + issue_description + summary + root_cause text. Retrieval uses the same vector space via neo4j-graphrag's semantic search, scoped to your org and team.

Retrieval Strategies

This is the layer most memory systems skip: OpenSRE doesn't just recall individual episodes — it synthesizes them into a reusable strategy once there's enough evidence that a pattern is real.

  • When it's generated — lazily, inside a memory-search call, whenever that search returns two or more similar episodes for the same issue shape and component. There's no separate cron job or background synthesis step; a strategy is only ever built in response to an actual search that qualifies.
  • What it is — a cached, LLM-written markdown playbook keyed by (org, team, issue_type, component), roughly structured as: common root causes, recommended investigation steps, key skills/commands that worked, and anti-patterns to avoid.
  • Provenance — every strategy links back to the episodes it was derived from, so you can always see which past incidents shaped a given playbook.
  • Staying current — a qualifying search always re-synthesizes rather than serving a stale cached copy, so a strategy reflects the latest set of similar episodes, not just whichever were around when it was first written. A strategy manually edited in the web UI can be overwritten the next time a qualifying search re-generates it — treat manual edits as provisional unless you also intend to stop new episodes from qualifying.
  • How the agent uses it — a strategy only ever reaches the agent through the memory-search skill's JSON response (the strategy field above), or through a person reading it in the Memory hub. It's never silently injected into a system prompt on its own.

Viewing Memory

The web console's Memory hub covers:

  • Episodes — browse past investigation summaries, with severity/resolution context and a link back to the full trace.
  • Search — run the same semantic search the agent uses, by hand.
  • Strategies — browse synthesized playbooks, see which episodes derived them, and edit or delete a playbook directly.