Why Monitoring AI Applications Is Different
Running AI-powered applications in production isn’t just “software with a model.” The system behaves stochastically, depends on volatile data, and interacts with third-party model APIs that can throttle, drift, or fail without warning. Traditional observability—logs, metrics, and traces—still matters, but you must also monitor model quality, data integrity, hallucination/safety issues, and real-time costs. The goal is simple: zero-downtime operations where incidents are detected early, contained fast, and resolved safely without user impact.
This guide walks through the core components of a modern monitoring stack for AI apps, with actionable steps to set up logging, error tracking, and incident response. You’ll also learn how to manage AI-specific risks like model drift, token explosions, and prompt sensitivity while preserving privacy and controlling costs.
Define What “Healthy” Means: SLOs, SLIs, and Error Budgets
Before you collect data, define what “good” looks like. Monitoring is only useful when tied to explicit expectations.
- Service level indicators (SLIs): Quantifiable measures of system behavior.
- Availability: Successful request rate.
- Latency: p50/p90/p95/p99 latency for inference, streaming start time, and time-to-first-token.
- Quality: Accuracy, RMSE, BLEU/ROUGE, or LLM-specific acceptance rate, groundedness/hallucination rate, safety violation rate, response format validity.
- Cost: Cost per request, tokens per request, per-user budget usage.
- Infrastructure: GPU/CPU utilization, VRAM, CUDA error rate.
- Service level objectives (SLOs): Targets for SLIs, e.g., “p95 latency < 800 ms, 30-day rolling availability ≥ 99.9%.”
- Error budgets: How much unreliability you can tolerate before halting risky changes. Example: with 99.9% SLO, monthly error budget is ~43 minutes of aggregate SLI breach.
Tie alerts to SLO burn rate rather than static thresholds to reduce noise. Keep SLOs per surface: API inference, batch retraining pipelines, streaming outputs, and admin tooling may each need different targets.
Architecture for Observability: Logs, Metrics, Traces, and Events
Aim for full-fidelity context you can query during incidents.
- Logs: Structured, centralized, and privacy-aware.
- Metrics: Time-series counters and gauges for fast alerting.
- Traces: End-to-end request path from gateway to model and downstream services.
- Events: Discrete annotations like “model v3.4 deployed” or “LLM provider outage,” used for correlation.
A common open stack:
- OpenTelemetry (OTel) for instrumentation (logs, metrics, traces).
- Prometheus for metrics, Grafana for dashboards/alerts.
- Loki or ELK (Elasticsearch/Logstash/Kibana) for logs.
- Jaeger or Tempo for tracing.
- Sentry or equivalent for error tracking.
- ML monitoring: Evidently AI, WhyLabs, Arize, Fiddler, or LLM-focused tools like Langfuse/Gantry/Humanloop.
- Model lifecycle: MLflow or Weights & Biases for versioning and evaluations.
Use a consistent correlation ID strategy:
- Generate a request_id at the edge.
- Propagate trace_id/span_id via HTTP headers (W3C Trace Context).
- Include request_id + trace_id in every log and span.
Structured Logging That Protects Privacy and Unlocks Insight
Logs should be machine-parsable, low-noise, and built for aggregation. For AI apps, add AI-aware fields and privacy controls.
What to Log (and What Not To)
- Always:
- request_id, trace_id, user_id_hash
- endpoint, http_method, http_status, error_code
- model_name, model_version, provider, region
- params: temperature, top_p, max_tokens
- input_length, output_length, token_count, cache_hit
- latency_ms (end-to-end and per-stage)
- cost_usd_estimate
- safety_flags (toxicity, PII, policy_violation)
- retries, backoff_ms, circuit_breaker_state
- infra: node_id, gpu_model, vram_used_mb
- Sometimes:
- prompt_hash or prompt_template_id (avoid raw prompts unless explicitly allowed)
- output_schema_valid (true/false), validation_errors_count
- feature_stats_id for tabular/vision models (link to feature store)
- Avoid:
- Raw PII. If needed for debugging, enable a short-lived, encrypted, access-controlled debug mode with automatic redaction and retention limits.
Example: Python Structured JSON Logs with Redaction
import json, time, uuid, os
from datetime import datetime
REDACT_KEYS = {"email", "name", "ssn", "phone", "address"}
def redact(payload: dict):
def _redact_value(v):
if isinstance(v, str) and len(v) > 0:
return "[REDACTED]"
return v
return {
k: _redact_value(v) if k.lower() in REDACT_KEYS else v
for k, v in payload.items()
}
def log(event: dict):
event["ts"] = datetime.utcnow().isoformat() + "Z"
event["service"] = "inference-api"
event["env"] = os.getenv("ENV", "prod")
print(json.dumps(event, separators=(",", ":")))
def log_inference_start(ctx, req):
ctx["request_id"] = ctx.get("request_id") or str(uuid.uuid4())
log({
"level": "INFO",
"event": "inference_start",
"request_id": ctx["request_id"],
"trace_id": ctx.get("trace_id"),
"user_id_hash": req.get("user_id_hash"),
"model_name": ctx["model_name"],
"model_version": ctx["model_version"],
"params": {"temperature": req.get("temperature", 0.2), "max_tokens": req.get("max_tokens", 512)},
"input_length": len(req.get("prompt", "")),
"prompt_hash": hash(req.get("prompt_template_id", "default")),
})
def log_inference_end(ctx, req, res, started_at):
log({
"level": "INFO",
"event": "inference_end",
"request_id": ctx["request_id"],
"latency_ms": int((time.time() - started_at) * 1000),
"output_length": len(res.get("output", "")),
"token_count": res.get("token_count", 0),
"cache_hit": res.get("cache_hit", False),
"cost_usd_estimate": res.get("cost_usd_estimate", 0.0),
"safety_flags": res.get("safety_flags", []),
"http_status": 200,
})
def log_error(ctx, err, http_status=500, code="ERR_INFER"):
log({
"level": "ERROR",
"event": "inference_error",
"request_id": ctx.get("request_id"),
"trace_id": ctx.get("trace_id"),
"error_code": code,
"msg": str(err),
"http_status": http_status,
})
Key practices:
- Use structured JSON logs.
- Keep levels consistent (DEBUG/INFO/WARN/ERROR).
- Anonymize user identifiers and redact PII fields.
- Log start and end events to measure latency and cost.
- Guard logs behind role-based access control (RBAC); encrypt at rest and in transit.
- Implement log sampling for high-volume events but never sample ERROR logs.
Error Tracking: Capture, Classify, and Act
Centralize errors with stack traces, request context, and breadcrumbs. Tools like Sentry, Rollbar, or OpenTelemetry’s exception signals work well.
Common AI-specific error classes:
- Upstream provider issues: rate limits, 429/503s, regional outages, degraded latency.
- Model-time failures: CUDA OOM, kernel errors, device lost, operator mismatch after driver upgrades.
- Data/feature issues: missing features, schema drift, null spikes, out-of-range distributions.
- Business logic errors: response schema invalid, hallucination threshold exceeded, unsafe content.
- Infra/network: DNS failures, TLS handshake errors, connection pools exhausted.
Mitigations you should implement:
- Retries with jittered exponential backoff and caps.
- Circuit breakers to shed load when upstream is unhealthy.
- Idempotency keys for write operations to avoid duplication.
- Dead-letter queues (DLQ) for failed async tasks with replay tooling.
- Fallback strategies: alternate models/providers, cached answers, or degraded modes.
Example: Resilient LLM Call with Backoff and Circuit Breaker (Python)
import random, time
from datetime import datetime, timedelta
class CircuitBreaker:
def __init__(self, failures=5, reset_seconds=30):
self.failures = failures
self.reset_seconds = reset_seconds
self.state = "CLOSED"
self.count = 0
self.next_try = datetime.utcnow()
def allow(self):
if self.state == "OPEN" and datetime.utcnow() >= self.next_try:
self.state = "HALF_OPEN"
return self.state != "OPEN"
def record_success(self):
self.state = "CLOSED"
self.count = 0
def record_failure(self):
self.count += 1
if self.count >= self.failures:
self.state = "OPEN"
self.next_try = datetime.utcnow() + timedelta(seconds=self.reset_seconds)
def call_llm_with_retry(client, prompt, params, breaker: CircuitBreaker, max_attempts=5):
attempt = 0
while attempt < max_attempts:
attempt += 1
if not breaker.allow():
raise RuntimeError("UPSTREAM_UNAVAILABLE")
try:
res = client.generate(prompt=prompt, **params)
breaker.record_success()
return res
except Exception as e:
breaker.record_failure()
sleep = min(2 ** attempt + random.random(), 8.0)
time.sleep(sleep)
if attempt == max_attempts:
raise e
Instrument your retries and breaker state in logs and metrics. Alert on:
- Retry rate spikes
- Circuit breaker opens
- DLQ size growth
- Provider error code rates
Tracing and Correlation: See the Whole Request Path
Distributed tracing is vital when your request touches a feature store, vector database, content filters, model servers, and third-party LLMs.
Best practices:
- Create spans for pre-processing, feature retrieval, embedding, reranking, inference, post-processing, and safety checks.
- Add span attributes: model_name, model_version, temperature, token_count, cache_hit, provider_region.
- Record events on spans for retries, fallback decisions, and guardrail triggers.
- Propagate trace context across process boundaries and external calls via HTTP headers (traceparent).
Sampling strategies:
- Always sample ERROR traces.
- Sample a percentage of healthy requests, increasing during deployments.
- Dynamically elevate sampling for outlier latency or cost.
Metrics and Dashboards: The Golden Signals Plus AI-Specific KPIs
Start with the golden signals:
- Traffic: Requests per second, concurrent sessions, queue depth.
- Latency: p50/p95/p99 for total and per-stage.
- Errors: 4xx/5xx rates, timeout rate.
- Saturation: CPU/GPU/memory utilization, VRAM fragmentation, connection pool usage.
Add AI-specific metrics:
- Tokens/request, tokens/minute per user/org, cost/request, cost/minute.
- Cache metrics: cache hit ratio for embeddings or LLM responses.
- Safety/quality: groundedness score, policy violation rate, format validation pass rate.
- Model health: accuracy, F1, calibration error, drift scores (PSI/KL divergence), feature null rates.
- Provider health: 429/5xx rates by region, p95 latency by provider.
A minimal Grafana dashboard should include:
- Request volume and p95 latency by route and model version.
- Error rate with breakdown by error_code/provider.
- GPU/CPU/VRAM utilization and CUDA error counts.
- Token and cost charts with per-tenant breakdown.
- Drift widgets: PSI over time for key features; LLM safety violations per 1k requests.
- Annotations for deploys, model promotions, and provider incidents.
AI-Specific Monitoring: Data, Drift, and Output Quality
Traditional uptime isn’t enough. Your app can be “up” but wrong, unsafe, or expensive.
- Data quality:
- Schema checks: enforce input/output schemas with versioning.
- Constraints: ranges, uniqueness, null thresholds, categorical domains.
- Distribution drift: PSI for tabular; KL divergence or Earth Mover’s Distance for continuous features.
- Model quality:
- Continuous evaluation: sample production traffic to a labeled set or synthetic evaluators.
- Shadow deployments: run new model in parallel on mirrored traffic; compare metrics.
- LLM quality and safety:
- Response validators: JSON schema validation or pydantic for structured outputs.
- Hallucination detection: self-consistency checks, retrieval grounding verification (cite presence and overlap), and content moderation scores.
- Guardrails: block/transform unsafe outputs; log violation types.
Example: Response Schema Validation (Python + Pydantic)
from pydantic import BaseModel, ValidationError
class ProductSummary(BaseModel):
title: str
price: float
tags: list[str]
def validate_output(raw_text):
# Assume the LLM returns a JSON block
import json
try:
data = json.loads(raw_text)
ProductSummary(**data)
return True, None
except (json.JSONDecodeError, ValidationError) as e:
return False, str(e)
Alert on validation failure rate spikes and route invalid responses to a fallback formatter or a safer, smaller model specialized in structure repair.
Setting Up a Zero-Downtime Delivery Pipeline
Zero downtime is a product of progressive delivery and graceful degradation.
- Blue/green or canary deployments:
- Shift 1-5% of traffic to new versions; compare SLOs and rollback if burn rate exceeds thresholds.
- For model servers, use versioned endpoints: /v3.4 and route via gateway.
- Shadow traffic:
- Mirror requests to candidate models; do not serve responses. Compare latency, cost, and quality offline.
- Feature flags:
- Toggle behaviors (e.g., enable RAG, enable cache) without redeploying.
- Multi-provider fallback:
- Orchestrate across vendors; auto-failover on SLA breach or region outage.
- Caching:
- Prompt/embedding cache with TTL; prewarm hot keys on deploys.
Example: Canary Alert Rule (YAML-ish)
alert: CanaryModelSLOBurn
expr: (rate(http_requests_total{version="v-canary",status=~"5.."}[5m]) / rate(http_requests_total{version="v-canary"}[5m])) > 0.02
for: 10m
labels:
severity: critical
annotations:
summary: "Canary error rate > 2% for 10m"
description: "Rolling back canary model due to elevated errors."
Wire alerts to ChatOps and PagerDuty/On-Call with clear runbook links.
Incident Response: Prepare, Detect, Contain, Resolve, Learn
A strong incident program keeps users safe while you fix the problem.
- Preparation:
- On-call rotation with a staffed primary and backup.
- Runbooks per service: detection signals, quick checks, rollback commands, verification steps.
- Playbooks for common scenarios: provider outage, GPU OOM cascade, schema drift, cost surge, safety violation spike.
- Game days: simulate partial provider outages and region failovers.
- Detection:
- SLO burn-rate alerts; symptom-based alerts (latency, error rate, DLQ size).
- Quality alerts: spike in validation failures, hallucination rate, or policy violations.
- Cost alerts: tokens/minute exceeds budget by tenant; unexpected cache miss rate.
- Containment:
- Trigger circuit breakers for affected routes.
- Reduce generation parameters (max_tokens, temperature) to cap cost/latency.
- Switch to backup provider or a cheaper, smaller model for non-critical paths.
- Increase cache TTL and rate-limit heavy users.
- Resolution:
- Roll back recent deploy or model version.
- Hotfix validated bug; run playback tests before re-enabling traffic.
- Backfill missing features or rebuild corrupted index.
- Post-incident:
- Blameless postmortem with timeline, impact, root cause, contributing factors, and action items.
- Add tests/monitors that would have prevented or reduced impact.
- Update runbooks; close the loop with the team.
Example Scenario: Provider Degradation at Peak
- 18:00: p95 latency alert for /generate rises past 2s; 429 rate spikes in one region.
- 18:02: Circuit breaker opens for provider A us-east; traffic auto-fails over to us-west and provider B.
- 18:04: Canary model stays below SLO, but cost per request up 30% due to provider B pricing.
- 18:05: Activate “cost-cap mode”: lower max_tokens from 1024 to 512, enable aggressive cache.
- 18:10: Latency stabilizes; users unaffected; announce degraded vendor status in status page.
- 18:30: Provider A resolves; gradually shift back 25% traffic every 5 minutes.
- Postmortem: Add region health checks to pre-emptively shift traffic and a cost-based alert to flag rapid provider switches.
Guardrails and Safety: Real-Time Protections
Safety is both a product feature and a compliance necessity.
- Pre- and post-generation filters for toxicity, PII leakage, and policy categories.
- Retrieval grounding: require citations and measure overlap between response and retrieved context; block if below threshold.
- Rate limits per user/tenant; audit logs for admin overrides.
- Human-in-the-loop review queues for high-risk outputs, with sampling of borderline cases.
Log safety outcomes and build dashboards that surface:
- Violations per category, per endpoint.
- False positives/negatives from human review.
- Time to remediation for unsafe outputs.
Cost and Performance Controls
Costs can explode silently in AI workloads. Monitor and enforce budgets.
- Track tokens and cost at request, user, and tenant levels.
- Enforce per-tenant budgets and daily caps; return graceful degradation responses if exceeded.
- Measure cache effectiveness; consider input normalization to improve hit rate.
- Autoscaling:
- Horizontal Pod Autoscalers (HPA) based on concurrency/latency.
- GPU autoscaling with queue depth signals.
- Optimize parameters:
- Lower temperature and max_tokens where quality isn’t impacted.
- Use response truncation with graceful ending and a “continue” follow-up pattern.
- Use smaller models for simple requests; route complex cases to larger models.
Data and Privacy Compliance
Logging and monitoring must respect user privacy and regulations.
- Minimize stored data; hash user IDs; redact PII by default.
- Encryption at rest and in transit; secrets rotation and KMS integration.
- Access control: audit who can view raw logs; separate prod and dev accounts.
- Retention policies: tiered storage; short retention for verbose logs; longer for aggregated metrics.
- Regional data residency and cross-border restrictions as needed.
Putting It All Together: A Reference Setup
-
Instrumentation
- Use OpenTelemetry SDKs in all services. Include request_id and trace context propagation.
- Emit JSON logs with standardized fields; redact sensitive data.
-
Collection and Storage
- Metrics: Prometheus scraping; remote write to long-term store if needed.
- Logs: Fluent Bit/Vector sidecars to Loki/ELK.
- Traces: OTLP to Jaeger/Tempo with sampling rules.
-
Visualization
- Grafana dashboards for golden signals, model quality, costs, and safety.
- Sentry for error grouping and release tracking.
-
AI Monitoring
- Evidently/WhyLabs/Arize for drift and model KPIs; alerts on PSI thresholds.
- Langfuse/Gantry/Humanloop for LLM prompt/response traces and evaluation.
-
Delivery
- Canary releases with automated rollback on SLO burn.
- Shadow traffic for model candidates; continuous evaluation pipeline.
-
Resilience
- Circuit breakers, retries with backoff, DLQs for async jobs.
- Multi-provider abstraction with health checks and cost awareness.
-
Incident Management
- PagerDuty/On-Call integrated with SLO-based alerts.
- Runbooks linked from alerts; ChatOps for status updates and rollbacks.
Practical Checklists
Logging Checklist
- JSON structured logs with consistent schema
- Correlation IDs: request_id + trace_id
- AI fields: model_name/version, params, tokens, cost, safety flags
- Latency and per-stage timings
- Redaction of PII; RBAC; encrypted storage
- Error logs with stack traces and error codes
- Sampling for INFO; never sample ERROR
Error Tracking Checklist
- Centralized error tool (Sentry/OTel) with environment and release tags
- Retry/backoff + circuit breakers instrumented
- DLQs and replay tooling
- Idempotency keys for writes
- Provider-specific error code classification
Metrics and Alerts Checklist
- Golden signals per service
- AI KPIs: quality, safety, drift, cost, cache hit rate
- GPU/infra health: VRAM, CUDA errors
- SLO-based burn-rate alerts; deduplicated, routed to on-call
- Deploy annotations on dashboards
Incident Response Checklist
- On-call rotation and escalation policies
- Runbooks/playbooks for common failures
- Rapid rollback capability and proven canary process
- ChatOps + status page integration
- Blameless postmortems with action items
Actionable 30-Day Plan
Week 1:
- Define SLIs/SLOs for each surface; create error budgets.
- Standardize logging schema; implement redaction and correlation IDs.
- Basic Grafana dashboards for golden signals.
Week 2:
- Add tracing with OpenTelemetry across services.
- Integrate Sentry; classify top 10 error signatures; implement retries and circuit breakers.
- Add cost and token metrics.
Week 3:
- Introduce model/LLM-specific metrics: safety flags, validation failures, drift checks with Evidently/WhyLabs.
- Build canary pipeline and feature flags; set up shadow traffic.
Week 4:
- Write runbooks and incident playbooks; run a game day simulating provider outage.
- Implement budget enforcement (tenant caps) and cache pre-warming.
- Tune alert thresholds based on SLO burn rates; dedupe noisy alerts.
Common Pitfalls and How to Avoid Them
- Over-logging raw prompts/outputs: Redact or hash. Provide secure, time-limited debug toggles instead.
- Alert fatigue: Tie alerts to SLOs; use multi-signal conditions and escalation.
- Ignoring cost signals: Track tokens and per-tenant budgets; alert on cost anomalies.
- One-provider dependency: Implement multi-region, multi-provider fallback now, not after the first outage.
- No model rollback path: Version every model artifact; enable instant rollback and shadow evaluation.
- Unvalidated outputs: Enforce schemas and guardrails; treat validation failure rate as a first-class SLI.
Conclusion
Zero-downtime AI operations are achievable with a deliberate observability strategy tuned to the realities of stochastic models, volatile data, and third-party dependencies. Start by defining SLOs that reflect user experience, instrument your system with structured logs, metrics, and traces, and add AI-native monitoring for drift, safety, and cost. Back it with resilient patterns—retries, circuit breakers, multi-provider failover—and a mature incident response process with runbooks and canaries. The result is not just higher uptime, but predictable quality, controlled costs, and the confidence to ship AI features faster.