Back to ProjectsMonitoring & SRE
Full-Stack Observability Platform
Centralised metrics, logs, and distributed traces stack. Custom SLO dashboards, automated incident response, and 13-month data retention via Thanos.
- MTTR Reduction
- 82%
- Metrics/sec
- 500K+
- Dashboards
- 40+
- Log Retention
- 13 mo
MTTR Reduction
Metrics/sec
Dashboards
Log Retention
Project Overview
Designed and implemented a unified observability platform covering metrics (Prometheus + Thanos for long-term storage), logs (Fluent Bit → OpenSearch), and distributed traces (Jaeger / OpenTelemetry). Built 40+ custom Grafana dashboards tracking business KPIs alongside infrastructure health. Created automated PagerDuty alerting rules aligned with SLO/SLA thresholds with runbooks stored as code. MTTR dropped from 45 minutes to under 8 minutes within the first month of rollout.
Key Achievements
- Reduced Mean Time to Resolution (MTTR) by 82% in first month post-rollout
- Ingests 500K+ metrics/sec and 2 TB+ logs/day with 13-month retention via Thanos
- 40+ custom dashboards covering business, infrastructure, and application layers
- SLO-based alerting with automated runbook execution for 12 common failure patterns