Operations Map
Operations Map
Use this section after the first local run works. It groups the material needed to deploy, monitor, troubleshoot, and harden a sandbox fleet.
Operator Lanes
| Lane | What it covers | Read |
|---|---|---|
| Production Setup | Host prerequisites, service setup, deployment layout, and AIWG-connected mode. | Deployment |
| Day-2 Procedures | Server lifecycle, runtime management, task operations, HITL, and incident routines. | Operations |
| Monitoring | Prometheus, Grafana, metrics naming, SLOs, alerts, and dashboard wiring. | Monitoring |
| Troubleshooting | Common install, runtime, task, agent, and AIWG integration failures. | Troubleshooting |
Operations Flow
1. Deployment - install and configure the host. 2. Operations - run the service day to day. 3. Monitoring - instrument the fleet. 4. Reliability Map - define SLOs and failure handling. 5. Troubleshooting - diagnose and recover.
Specialized Ops
- Crash Loop Detection - runtime crash detection and
operator unblock.
- Telemetry - metrics and textfile collector pipeline.
- Transport Audit - operator-facing event/log streams.
- macOS Host Runtime and Keychain - opt-in
launchd host execution and fail-closed local CA Keychain storage.
- Observability Design - deeper observability
architecture and implementation checklist.