technology

Monitor and Aggregate Logs Effectively: Using ELK and Sentry for Proactive System Management

Learn to master proactive system management with ELK and Sentry, honing your log monitoring and aggregation skills for improved system reliability and performance.

October 9, 2025
log-monitoring ELK-stack Sentry system-management log-aggregation proactive-management IT-operations
16 min read

Why Proactive Log Monitoring Matters

When systems fail, logs tell the story—if you’ve designed them to. Proactive system management means spotting early warning signs, isolating failures quickly, and continuously learning from incidents. To do that at scale, you need two capabilities:

  • A way to aggregate, parse, query, and visualize ALL your logs.
  • A way to capture and triage application errors with context, alerts, and ownership.

The Elastic Stack (ELK: Elasticsearch, Logstash, Kibana—often with Beats) delivers the first. Sentry delivers the second. Used together, they provide a powerful, complementary toolkit for maximizing reliability, reducing mean time to detect (MTTD) and mean time to resolve (MTTR), and guiding performance improvements.

This guide shows you how to design a log strategy, implement ELK and Sentry, connect them, and operate them confidently at scale—complete with examples, configurations, and practices to avoid alert fatigue and runaway costs.

ELK and Sentry at a Glance

ELK for Log Aggregation and Analytics

  • Elasticsearch stores and indexes event data for fast search and aggregation.
  • Logstash transforms and enriches logs (parse, normalize, add context).
  • Kibana visualizes data, builds dashboards, and sets alerts.
  • Beats (Filebeat, Metricbeat, etc.) ship data from hosts and services.

Best use cases:

  • Centralized log collection from multiple services and environments.
  • Exploratory analysis, distributed debugging, trend analysis.
  • Operational dashboards and long-term storage.

Sentry for Error Monitoring and Performance

  • Captures exceptions, stack traces, breadcrumbs, and performance spans.
  • De-duplicates issues, manages ownership, and provides triage workflow.
  • Alerts on error spikes, regressions, releases, and performance thresholds.

Best use cases:

  • Deep application awareness: exceptions, transactions, user impacts.
  • Proactive code-quality feedback, release health, and regression detection.
  • Assignable issues with context for engineering teams.

Design a Logging Strategy First

Before tooling, decide what you want from your logs:

  • What questions should logs answer? (Who, what, where, when, why)
  • Which systems are covered? (Services, jobs, infra, edge, clients)
  • What’s the data model? (Standard fields, schemas, relationships)
  • How will you alert? (Indicators, thresholds, burning rates, ownership)
  • How long will you keep data? (Retention, archiving, compliance)

Structured Logging: Make Logs Queryable

Switch from free-form text to structured JSON logs. Benefits:

  • Reliable parsing, consistent fields.
  • Easier correlation across services.
  • Lower ingestion headaches and fewer parsing rules.

Adopt a common schema like Elastic Common Schema (ECS) for fields:

  • event., log., host., http., url., user., trace., transaction., error., service., cloud.*

Example JSON log:

{
  "timestamp": "2025-03-10T13:03:21.456Z",
  "log.level": "error",
  "message": "Checkout failed",
  "service.name": "checkout-api",
  "event.dataset": "app",
  "http.method": "POST",
  "http.response.status_code": 500,
  "user.id": "u_28391",
  "order.id": "o_98172",
  "trace.id": "6b1f5d2f7fe6a9a2",
  "transaction.id": "a39e4dfe2f2d3a71",
  "sentry.event_id": "b5e3c1d27a1d4a3f9de0d1e1f57e2c9a",
  "error.type": "PaymentDeclined",
  "error.message": "Payment processor timeout"
}

Correlation IDs and Context Propagation

Propagate a correlation ID through every request across services:

  • trace.id: A unique ID to correlate logs across services.
  • transaction.id: A span per request in a single service hop.

Include these IDs in both logs and Sentry. This lets you jump from a Sentry issue to the exact logs of the request in Kibana (and vice versa).

Tips:

  • Generate a trace ID at the edge (API gateway or first service).
  • Pass via headers (e.g., traceparent for W3C Trace Context).
  • Log and tag it in all downstream services.

Sampling and Retention

  • Logs: Consider dynamic sampling of verbose logs in production (reduce debug-level noise). Keep structured “event” logs at full fidelity.
  • Sentry: Use dynamic sampling (by release, environment, user segment) to control volume without losing critical signals.
  • Retention: Apply index lifecycle management (hot/warm/cold/frozen tiers). Keep high-fidelity recent data; downsample or archive older data.

A Reference Architecture to Start With

[Apps/Services] --JSON logs--> [Filebeat/FluentBit] --> [Logstash] --> [Elasticsearch] --> [Kibana]
        |                                  |
        |--Sentry SDK (errors, traces)-----|
  • Apps emit structured logs and are instrumented with Sentry SDKs.
  • Filebeat (or Fluent Bit) ships logs off the host or container.
  • Logstash parses, normalizes to ECS, enriches (geo, cloud, k8s).
  • Elasticsearch stores, indexes, and manages lifecycle.
  • Kibana is your analysis and alerting UI.
  • Sentry captures exceptions and transaction performance; Sentry event IDs land in logs.

Step-by-Step Implementation

1) Instrument Apps for Structured Logging

Use your language’s structured logger (e.g., pino for Node, structlog or loguru for Python, logrus for Go). Ensure every log is JSON and includes standard fields.

Example in Node.js (pino + Sentry):

// npm i @sentry/node pino pino-http

const Sentry = require('@sentry/node');
const pino = require('pino');
const pinoHttp = require('pino-http');
const { v4: uuid } = require('uuid');

Sentry.init({
  dsn: process.env.SENTRY_DSN,
  environment: process.env.NODE_ENV || 'dev',
  tracesSampleRate: 0.2 // adjust with dynamic sampling later
});

const logger = pino({ level: process.env.LOG_LEVEL || 'info' });

const httpLogger = pinoHttp({
  logger,
  customProps: (req, res) => ({
    'service.name': 'checkout-api',
    'trace.id': req.headers['x-trace-id'] || uuid(),
    'user.id': req.user?.id,
    'http.method': req.method,
    'http.response.status_code': res.statusCode
  })
});

// Example route
app.post('/checkout', httpLogger, async (req, res) => {
  const traceId = req.log.bindings()['trace.id'];

  try {
    // ... do work
    res.status(200).send({ ok: true });
  } catch (err) {
    const eventId = Sentry.captureException(err, scope => {
      scope.setTag('trace.id', traceId);
      scope.setContext('order', { id: req.body.orderId });
      return scope;
    });
    req.log.error({ 'sentry.event_id': eventId, 'trace.id': traceId, 'order.id': req.body.orderId, err }, 'Checkout failed');
    res.status(500).send({ error: 'internal_error', traceId, eventId });
  }
});

This example:

  • Captures exceptions in Sentry with tags and context.
  • Logs the Sentry event_id and the trace.id so you can correlate.

2) Ship Logs Reliably with Filebeat

Install Filebeat on hosts or as a DaemonSet on Kubernetes to tail JSON logs and forward to Logstash.

filebeat.yml:

filebeat.inputs:
  - type: filestream
    id: app-logs
    paths:
      - /var/log/apps/*.json
    parsers:
      - ndjson:
          add_error_key: true
          expand_keys: true

processors:
  - add_host_metadata: ~
  - add_cloud_metadata: ~
  - add_docker_metadata: ~
  - add_kubernetes_metadata: ~
  - decode_json_fields:
      fields: ["message"]
      target: ""
      overwrite_keys: true
      when:
        regexp:
          message: '^\{.*\}$'
  - drop_fields:
      fields: ["agent.version", "ecs.version"] # keep essential fields only if needed

output.logstash:
  hosts: ["logstash:5044"]
  ssl.enabled: false

Notes:

  • If your logs are already JSON with flattened fields, you may not need decode_json_fields.
  • For Kubernetes, consider Filebeat’s autodiscover with hints.

3) Parse and Enrich in Logstash

Use Logstash pipelines to normalize fields to ECS, parse message fallback, and enrich.

logstash.conf:

input {
  beats { port => 5044 }
}

filter {
  # Ensure timestamp alignment
  if [timestamp] {
    date { match => ["timestamp", "ISO8601"] }
  }

  # Parse exceptions if present
  if [err] and [err][stack] {
    mutate {
      add_field => { "[error][stack_trace]" => "%{[err][stack]}" }
    }
  }

  # Example: parse free-form fallback
  if [message] and ![message].is_a?(Hash) {
    grok {
      match => { "message" => "%{TIMESTAMP_ISO8601:timestamp} %{LOGLEVEL:log.level} %{DATA:service.name} - %{GREEDYDATA:message}" }
    }
  }

  # Normalize to ECS
  mutate {
    rename => { "level" => "[log][level]" }
    rename => { "status" => "[http][response][status_code]" }
  }

  # Derive fields
  if [http][response][status_code] >= 500 {
    mutate { add_tag => ["error"] }
  }

  # Drop secrets
  mutate {
    remove_field => ["password", "token", "credit_card"]
  }
}

output {
  elasticsearch {
    hosts => ["http://elasticsearch:9200"]
    index => "logs-app-%{+YYYY.MM.dd}"
    ilm_enabled => true
    ilm_rollover_alias => "logs-app"
    ilm_policy => "logs-app-policy"
  }
  stdout { codec => rubydebug }
}

4) Set Up Index Lifecycle Management (ILM)

Control retention and cost with hot-warm-cold tiers.

ILM policy example:

{
  "policy": {
    "phases": {
      "hot": { "actions": { "rollover": { "max_size": "50gb", "max_age": "7d" } } },
      "warm": { "min_age": "7d", "actions": { "forcemerge": { "max_num_segments": 1 }, "shrink": { "number_of_shards": 1 } } },
      "cold": { "min_age": "30d", "actions": { "searchable_snapshot": { "snapshot_repository": "s3_repo" } } },
      "delete": { "min_age": "90d", "actions": { "delete": {} } }
    }
  }
}

Tips:

  • Start with fewer primary shards (e.g., 1–3) for low to moderate volumes. Too many shards waste memory.
  • Compress older segments via forcemerge in warm phase.
  • Use snapshots to S3 or equivalent for long-term archiving.

5) Visualize and Explore in Kibana

Create dashboards that answer operational questions fast:

  • Top errors by service.name, error.type.
  • Error rate over time (count of logs with log.level:error or status >= 500 / total requests).
  • Latency distributions (if you log timing).
  • 95th percentile per endpoint.
  • User impact: errors by user.id or cohort.

Handy KQL queries:

  • All errors for a specific request:
    • trace.id: “6b1f5d2f7fe6a9a2”
  • Recent server errors with checkout keyword:
    • log.level: “error” and service.name: “checkout-api” and message: “checkout”
  • Sentry-linked logs in production:
    • sentry.event_id: “*” and environment: “prod”

6) Capture Exceptions and Transactions in Sentry

Set up Sentry SDK to capture errors, performance, and release context.

Python example (FastAPI):

# pip install sentry-sdk fastapi uvicorn

import sentry_sdk
from fastapi import FastAPI, Request
import uuid
import logging

sentry_sdk.init(
    dsn=os.getenv("SENTRY_DSN"),
    environment=os.getenv("ENV", "dev"),
    traces_sample_rate=0.2
)

logger = logging.getLogger("app")
app = FastAPI()

@app.middleware("http")
async def add_trace_id(request: Request, call_next):
    trace_id = request.headers.get("x-trace-id", str(uuid.uuid4()))
    response = await call_next(request)
    response.headers["x-trace-id"] = trace_id
    request.state.trace_id = trace_id
    return response

@app.get("/pay")
def pay(request: Request):
    try:
        raise RuntimeError("Payment gateway timeout")
    except Exception as e:
        with sentry_sdk.push_scope() as scope:
            scope.set_tag("trace.id", request.state.trace_id)
            event_id = sentry_sdk.capture_exception(e)
        logger.error("Payment failed", extra={"trace.id": request.state.trace_id, "sentry.event_id": event_id})
        return {"ok": False, "trace_id": request.state.trace_id, "sentry_event": event_id}

Sentry tips:

  • Set release tags on deploy for regression tracking.
  • Use environments (dev/stage/prod) to segment alerts.
  • Define issue ownership rules (e.g., by service or path) so alerts reach the right team.

7) Alerting and Anomaly Detection

Set alerts that catch real problems without flooding inboxes.

In Kibana:

  • Threshold alerts: error logs rate > X/min for Y minutes.
  • Ratios: 5xx responses / total requests > 2% for 10 minutes.
  • Unique user errors spike: cardinality of user.id with error tag > baseline.

Example approach:

  • Create a Kibana alert on index “logs-app-*”:
    • Condition: count of logs where log.level: “error” AND service.name: “checkout-api” over 5 minutes > 100.
    • Action: Send to Slack; include a link to a pre-filtered dashboard.

In Sentry:

  • Alert on new issue frequency: > 10 events in 5 minutes.
  • Alert on error rate per release spike: error rate increased by 3x vs baseline.
  • Performance: Apdex or p95 latency threshold breach.

Guardrails:

  • Use notification rules per environment.
  • Mute known noisy issues with filters or sampling.
  • Bundle alerts with runbook links.

Correlating Sentry Issues with ELK Logs

The most powerful workflow is moving seamlessly between Sentry and Kibana.

Pattern:

  • Every request has trace.id.
  • Sentry events include trace.id via tags/context.
  • Logs include sentry.event_id when exceptions occur.
  • Both Sentry and Kibana display these fields, enabling cross-links.

Practical steps:

Result:

  • On-call sees a Sentry alert with a one-click link to all related logs for the same trace.
  • From Kibana, engineers jump straight to the exact Sentry issue.

Scaling and Cost Optimization

As volume grows, control cost and maintain performance.

Elasticsearch:

  • Right-size shards: Start with 1–3 primary shards; increase only if needed. Target shard size 20–50 GB.
  • Data tiers: Hot (ingest/search), Warm (search-heavy, less ingest), Cold/Frozen (cheap, search ok).
  • ILM tuning: Shorten hot retention if needed; compress aggressively in warm.
  • Avoid high-cardinality fields in aggregations (e.g., user.id) unless necessary; use rollups or sampling for analytics.
  • Use ingest pipelines for lightweight transforms instead of heavy Logstash filters when feasible.

Sentry:

  • Dynamic sampling by transaction name, user cohort, or environment.
  • Filter noisy errors (e.g., client aborts, bot traffic).
  • Set rate limits per project to prevent bursts from overwhelming budgets.
  • Grouping rules to deduplicate related issues effectively.

Logs:

  • Do not log entire payloads—log metadata and hashes. Redact or tokenize PII.
  • Limit debug logs in production; enable on-demand via feature flags.

Security and Compliance

  • PII scrubbing:
    • In Logstash: remove or hash known sensitive fields.
    • In Sentry: configure server-side data scrubbing (headers, query strings, request bodies).
  • Access controls:
    • Elasticsearch/Kibana: role-based access; separate indices for sensitive services.
    • Sentry projects per team; API tokens scoped to specific projects.
  • Encryption: TLS in transit; disk encryption at rest.
  • Retention policies aligned with compliance (e.g., delete logs after 90 days).
  • Audit trails: log access to dashboards and issue changes.

Cloud-Native and Kubernetes

Kubernetes setups benefit from standard patterns:

  • Shipping logs:
    • Filebeat or Fluent Bit as a DaemonSet, capturing container stdout/stderr.
    • Autodiscover to enrich with k8s metadata: namespace, pod, container.
  • Config via hints:
    • Annotate pods with parsing hints to reduce centralized parsing logic.
  • ECK (Elastic Cloud on Kubernetes):
    • Manage Elasticsearch, Kibana, and Beats via CRDs.
  • Sentry:
    • Use sidecar or environment injection for DSN and release; include commit SHA in release tags.
  • Network:
    • Use headless services for Logstash; cluster-local endpoints for reliability.

Common Pitfalls and How to Avoid Them

  • Unstructured logs:
    • Fix: enforce structured JSON at the source; add CI checks.
  • Too many shards:
    • Fix: start small, monitor shard size, roll indices with ILM; shard only when necessary.
  • Cardinality explosion:
    • Fix: avoid high-cardinality fields in aggregations; store as keyword but aggregate carefully; use runtime fields sparingly.
  • Alert fatigue:
    • Fix: set SLO-based alerts (burn rates), deduplicate, use minimum durations, and route to owners; mute noisy rules.
  • Parsing at scale in Logstash:
    • Fix: move lightweight transforms to ingest pipelines; scale Logstash horizontally; benchmark grok patterns.
  • PII leakage:
    • Fix: scrub at source and at pipeline; implement data classification; test with synthetic data.

From Signals to Actions: SLOs, Burn Rates, and Runbooks

Define SLIs and SLOs so alerts reflect user experience, not just server-side noise.

SLIs to consider:

  • Error rate: 5xx responses / total requests.
  • Latency: p95 response time per key endpoint.
  • Availability: successful checks over time window.
  • Client error impact: Sentry issue rate by user or release.

SLO examples:

  • Checkout API error rate < 1% over 30 days.
  • p95 latency < 300 ms for /checkout in business hours.

Burn rate alerts:

  • Short-window burn (high sensitivity): 2h window; trigger if burn rate > 14x (rapid degradation).
  • Long-window burn (low sensitivity): 24h window; trigger if burn rate > 6x (sustained degradation).

Runbooks:

  • Link Kibana searches and Sentry dashboards for the service.
  • Steps to collect context: latest deploys, feature flags, dependency health.
  • Known remediations: rollback, scaling, cache bypass, circuit breaker.
  • On-call rotation and escalation rules.

Integrate with Tracing and APM

While logs and Sentry cover a lot, full observability includes metrics and traces.

  • OpenTelemetry:
    • Instrument apps to emit spans and metrics.
    • Propagate W3C trace context so logs, spans, and Sentry events share trace.id.
  • Elastic APM or Sentry Performance:
    • Use one or both for deep transaction visibility.
    • If using Elastic APM, map transaction.id and trace.id into logs and Sentry tags.
  • Correlation UX:
    • Kibana APM traces link to logs by trace.id.
    • Sentry Performance shows traces with spans; tags carry trace.id for cross-tool jumps.

30-60-90 Day Rollout Plan

First 30 days:

  • Define schema, correlation strategy, and retention.
  • Instrument 1–2 critical services with structured logging and Sentry.
  • Deploy Filebeat + Logstash + Elasticsearch + Kibana (or Elastic Cloud).
  • Basic Kibana dashboard for errors and latency.
  • Sentry alerts for new issues and spikes.

Days 31–60:

  • Expand to all user-facing services.
  • Implement ILM and hot/warm tiers; set up snapshots.
  • Add anomaly/ratio alerts in Kibana; performance alerts in Sentry.
  • Ownership rules: map services to teams; document runbooks.
  • Start dynamic sampling in Sentry for cost control.

Days 61–90:

  • Introduce OpenTelemetry for tracing where missing.
  • Correlate Sentry issues to Kibana with links; add URL formatters.
  • Harden security: RBAC, scrubbing, encryption, retention.
  • Cost optimization: shard tuning, field selection, compress warm data.
  • Post-incident reviews incorporate log and Sentry insights; refine dashboards and alerts.

Practical Checklists

Instrumentation checklist:

  • JSON structured logs across all services.
  • trace.id and transaction.id added to every request.
  • sentry.event_id logged on exceptions.
  • PII scrubbed at source; secrets never logged.

ELK checklist:

  • Filebeat/Fluent Bit shipping with k8s metadata.
  • Logstash or ingest pipelines normalize fields to ECS.
  • ILM policy with hot/warm/cold/delete phases.
  • Dashboards for top errors, latency, and user impact.
  • Alerts on error rate and burn rates.

Sentry checklist:

  • SDK initialized with environment and release tags.
  • Ownership rules per service/path.
  • Alerts for new issue spikes and performance regressions.
  • Dynamic sampling configured for production.

Operations checklist:

  • RBAC for Kibana and Sentry.
  • Snapshots to object storage; tested restore process.
  • Cost monitoring: shard sizes, index counts, event volume.
  • Runbooks linked from alerts; on-call schedule set.

Actionable Tips You Can Implement Today

  • Add a trace.id header at your API gateway and log it end-to-end.
  • Inject sentry.event_id into logs on every captured exception.
  • Create a Kibana saved search filtered on trace.id and add the URL to your Sentry alert template.
  • Enforce JSON logging in CI by failing builds that emit plain text logs in tests.
  • Turn on ILM with a conservative hot window (e.g., 7 days) and verify search speed before widening.
  • Start Sentry with a low tracesSampleRate, then raise for key transactions.
  • Add a single “Error Overview” dashboard showing:
    • Error rate by service
    • Top 5 error types
    • Error count by release
    • Recent high-severity traces with links to Sentry

Bringing It All Together

ELK and Sentry complement each other: ELK gives you unified visibility across systems, while Sentry focuses on application health, code-level context, and triage. Mastering both—plus correlation via trace IDs and shared context—transforms your operations from reactive firefighting to proactive reliability engineering.

Design your schema, instrument with discipline, set smart alerts, and tune for scale. With the right foundations, you’ll see issues earlier, resolve them faster, and build the observability muscle that keeps systems resilient as they grow.

Share this article
Last updated: October 9, 2025

Related technology Posts

Discover more startup know-how and business insights

How to Resolve Specific Safari Bugs: A Detailed Troubleshoot...

Discover effective solutions for resolving specific Safari bugs in 2024 with our...

Effective Memory Management Solutions: Addressing Out-of-Mem...

Discover modern strategies to tackle out-of-memory errors and enhance your syste...

How to Resolve Preflight Request Failures: Troubleshooting C...

Master CORS troubleshooting in 2024 by understanding and resolving preflight req...

CDN Configuration Errors: Troubleshooting Guide with Cloudfl...

Master the art of troubleshooting CDN configuration errors with Cloudflare and A...

Need Expert Help?

Get professional consulting for startup and business growth.
We help you build scalable solutions that lead to business results.