RUNBOOK.md
Agent operations runbook
Give operators a tested path from alert to diagnosis, containment, rollback, recovery, communication, and follow-up using the signals and failure modes of the complete agent workflow.
Fillable structure
Replace prompts with project evidence.
Remove any field that does not apply. A smaller maintained file is more useful than generic documentation that agents and reviewers cannot trust.
- 01
Service ownership and dependencies
Make the operating boundary and escalation path available before an incident.
Owners: List primary, secondary, security, data, product, and vendor contacts
Dependencies: List models, tools, queues, databases, auth, providers, and critical limits
Environments: List production, staging, regions, release identifiers, and access procedure
- 02
Signals and diagnosis
Connect every alert to a workflow stage, user impact, and known failure mode.
Service indicators: Define completion, correctness, safety, latency, cost, and availability signals
Dashboards and traces: Link workflow, model, tool, queue, authorization, and business views
Diagnostic sequence: List the fastest read-only checks and evidence to preserve
- 03
Containment and recovery
Provide exact reversible actions with authority and stop conditions.
Contain: Define feature disablement, permission revocation, queue pause, isolation, and traffic controls
Rollback: Define code, configuration, prompt, model, schema, and data rollback
Recover: Define validation, replay, reconciliation, re-enable order, and customer remediation
- 04
Communication and learning
Keep stakeholders informed and convert failures into maintained controls.
Updates: Define internal, user, vendor, legal, and regulatory communication owners and cadence
Closure evidence: State the signals and reviewers required to close the incident
Follow-up: Add eval cases, tasks, architecture decisions, documentation, and owner deadlines
Related field guides
Next template
DECISIONS.md · Architecture decision log