What is Episodic Memory in SRE?
Episodic memory allows AI SRE agents to learn from past investigations. Instead of starting from scratch, an agent can recall a prior investigation with similar symptoms, the evidence it gathered, and the root cause it identified. This accumulated operational history helps the agent investigate recurring problems.
The Problem: Stateless AI Isn't Enough
Most AI tools are stateless. They have knowledge from their training data, but they don't remember your specific systems, your specific incidents, or what's worked for you in the past.
Imagine a senior SRE on their first day at a new company. They know SRE practices well, but they don't know your systems. Now imagine that same SRE after 2 years on-call. They've seen hundreds of incidents, they recognize patterns, they know the quirks of each service. Episodic memory is how we give that accumulated expertise to an AI agent.
How Episodic Memory Works in OpenSRE
Step 1: Investigation Completes
After every investigation, OpenSRE produces a structured record of the reported problem, evidence gathered, affected components, and the best-supported root cause when one was identified.
Step 2: Metadata Extraction
An LLM extracts structured metadata from the investigation outcome:
- Summary: A concise description of the investigation
- Issue type: A short classification of the problem shape, not only an alert type
- Issue description: What was reported or asked
- Affected components: Services, deployments, jobs, databases, hosts, or other involved subjects
- Severity:
critical,warning,info, or not applicable - Root cause: The current best-supported root cause, when identified
- Diagnostic status: Whether a root cause was identified with enough supporting evidence
Diagnostic status does not mean the production issue was fixed or remediated.
Step 3: Episode Storage
OpenSRE stores one investigation episode per conversation in Neo4j. Each episode captures the investigation summary, evidence-backed root cause when one was identified, affected services, and outcome. Unresolved investigations remain useful because they preserve approaches and evidence that did not produce a confirmed answer.
Step 4: Semantic Retrieval
When an investigation encounters concrete symptoms, the agent can search past episodes for semantically related incidents. OpenSRE uses vector retrieval over the stored episodes, then returns the most relevant investigation history as evidence for the current investigation.
Step 5: Agent-Guided Recall
Memory is not injected blindly into every prompt. The investigator uses the memory-search capability after it has enough concrete symptoms to form a useful query. This reduces irrelevant recall and keeps the current evidence primary.
Step 6: Strategy Synthesis
When a memory search returns at least two relevant episodes, OpenSRE can synthesize a concise investigation strategy from their root causes, successful skills, and ineffective approaches. The generated strategy is cached in Neo4j and remains traceable to its source episodes.
Episodic Memory vs Generic RAG
Generic retrieval-augmented generation usually searches documentation: runbooks, service docs, tickets, or knowledge-base pages. Episodic memory searches what actually happened during prior investigations.
The distinction matters during incidents:
- Documentation says what should happen. Episodes preserve what did happen.
- Runbooks describe intended steps. Episodes retain the evidence, failed hypotheses, and confirmed root cause from a real investigation.
- A knowledge base is curated. Episodic memory grows from operational work.
The two approaches complement each other. OpenSRE can use documentation for procedural knowledge while episodic memory supplies environment-specific history.
The Difference From Fine-Tuning
Episodic memory is not fine-tuning. Fine-tuning bakes knowledge into the model weights — it's expensive, requires ML expertise, and goes stale as your systems evolve.
Episodic memory is a retrieval system. Episodes are stored in Neo4j and retrieved by the agent when concrete symptoms warrant a memory search. New investigations extend that operational history without retraining the model.
What Gets Better Over Time
The compound effect of episodic memory:
| Memory state | What changes |
|---|---|
| First episodes | OpenSRE begins building environment-specific investigation history |
| Similar episodes | Relevant past evidence can surface for recurring symptoms |
| Two or more relevant search results | OpenSRE can synthesize a reusable investigation strategy |
| Repeated patterns | Investigators gain richer evidence about effective and ineffective approaches |
Viewing Memory in OpenSRE
The OpenSRE web console includes Memory. View stored episodes, filter by service or severity, see which strategies have been generated, and review the context provided to investigations.
Explore OpenSRE on GitHub → | Read the Memory docs → | Run OpenSRE →