Back to ProjectsMonitoring & SRE

Full-Stack Observability Platform

Centralised metrics, logs, and distributed traces stack. Custom SLO dashboards, automated incident response, and 13-month data retention via Thanos.

MTTR Reduction
82%

MTTR Reduction

Metrics/sec
500K+

Metrics/sec

Dashboards
40+

Dashboards

Log Retention
13 mo

Log Retention

Project Overview

Designed and implemented a unified observability platform covering metrics (Prometheus + Thanos for long-term storage), logs (Fluent Bit → OpenSearch), and distributed traces (Jaeger / OpenTelemetry). Built 40+ custom Grafana dashboards tracking business KPIs alongside infrastructure health. Created automated PagerDuty alerting rules aligned with SLO/SLA thresholds with runbooks stored as code. MTTR dropped from 45 minutes to under 8 minutes within the first month of rollout.

Key Achievements

  • Reduced Mean Time to Resolution (MTTR) by 82% in first month post-rollout
  • Ingests 500K+ metrics/sec and 2 TB+ logs/day with 13-month retention via Thanos
  • 40+ custom dashboards covering business, infrastructure, and application layers
  • SLO-based alerting with automated runbook execution for 12 common failure patterns