Investigation Skills

Skills are the primary integration surface for OpenSRE's agent: each one is a SKILL.md methodology doc plus optional Python scripts that let the agent query one part of your infrastructure — from checking Kubernetes pod health to reading Jira tickets to running a Cypher query against the topology graph.

How Skills Work

OpenSRE's agent uses the Claude Agent SDK's Skill tool. At session start, the agent sees lightweight metadata for every enabled skill (roughly 100 tokens each) — just enough to know what each one is for. When the agent decides a skill is relevant, the SDK loads that skill's full SKILL.md into context. Any scripts the skill defines then run via Bash, scoped to that skill's directory.

This progressive loading is what keeps the agent's context window manageable with 51 skills available — it only pays the token cost for the ones it actually uses in a given investigation.

Skill Categories

CategoryExamples
Core methodologyinvestigate, remediation, alerting-context
Observabilityobservability-coralogix, observability-grafana, observability-elasticsearch, observability-datadog, observability-splunk, observability-newrelic, observability-honeycomb, observability-jaeger, observability-sentry, observability-loki, observability-victorialogs, metrics-analysis, metrics-victoriametrics
Infrastructure & cloudinfrastructure-kubernetes, infrastructure-aws, infrastructure-azure, infrastructure-gcp, infrastructure-docker, infrastructure-neo4j (topology)
Alerting & on-callalerting-opsgenie
Incident managementincident-incidentio, incident-blameless, incident-firehydrant, incident-comms
Databasesdatabase-postgresql, database-mysql, database-snowflake, database-bigquery
Docs & knowledgeknowledge-base, knowledge-raptor, docs-notion, docs-google
Memorymemory-search
Code & version controlvcs-github, vcs-gitlab, vcs-bitbucket, vcs-sourcegraph
Ticketing & projectproject-jira, project-linear, project-clickup
Other integrationsplatform-jenkins, platform-argocd, platform-vercel, streaming-kafka, analytics-amplitude, runtime-config-flagd

See the full, current list under sre-agent/.claude/skills/ in the repo — this table covers every category, not necessarily every skill.

Adding Custom Skills

You can add investigation skills for your own stack. Skills are directories under sre-agent/.claude/skills/:

sre-agent/.claude/skills/
  my-custom-skill/
    SKILL.md       # What the skill does, and how to use it
    scripts/       # Executable scripts the skill invokes via Bash

SKILL.md describes the skill's purpose and usage to the agent. After adding a skill, regenerate the config-service skills catalog so it shows up in team/agent skill toggles (config_service/scripts/gen_skills_catalog.py).

Skills Configuration

Enable or disable skills at the team level with a wildcard plus a disable-list, or restrict a single agent explicitly:

skills:
  enabled:
    - '*'
  disabled:
    - observability-datadog
    - database-mysql
{
  "agents": {
    "my-agent-id": {
      "skills": {
        "infrastructure-kubernetes": true,
        "observability-datadog": false
      }
    }
  }
}

Disabled skills are removed from that thread's skill workspace before the agent's session starts — the agent never even sees their metadata. See Configuration for the full config reference.