Observability and reliability
Instrument a service with OpenTelemetry-spec traces, metrics, and logs, understand why agent traces are kept as evidence, and reason about reliability through SLIs, SLOs, error budgets, and the policy that turns a spent budget into a release freeze.
Units5
Duration26 min
Levelintermediate
By the end of this module, you'll be able to:
- Describe Cogitave's telemetry model - OpenTelemetry-spec traces, metrics, and logs under one correlation, bound to the Core model and queryable from MCP.
- Explain what AI/agent observability adds - the agent trace, token and cost, eval drift, and inference records kept as responsible-AI evidence.
- Define an SLI, an SLO, and an error budget, and compute a budget from an SLO and a window.
- State the error-budget policy - the rule that converts a spent budget into a per-service feature freeze - and why the SLA is always looser than the SLO.
- Read the observability and reliability standards as the source of truth and follow them for the detail.