Platform Reliability Engineer
AWS EKS | Terraform | Kubernetes | Aurora | Redis | OpenSearch | Prometheus | Grafana
I operate production SaaS infrastructure on AWS at scale.
My focus is Kubernetes reliability, cloud cost reduction, incident response, performance systems, and production-grade observability.
I Build and maintain systems across EKS, Aurora, Redis, OpenSearch, and GitHub Actions.
| Area | Outcome |
|---|---|
| Cloud cost optimization | Worked with a Qubole team to reduce AWS spend from $375k → $70k through rightsizing, idle resource cleanup, and architecture changes |
| Availability | Maintained 95% +uptime on production SaaS workloads |
| Data platform | Operated Aurora, Redis, and OpenSearch under production load |
- Stabilize — stop the bleed, rollback or scale, protect data
- Communicate — clear status, ETA, single owner
- Diagnose — metrics, logs, traces, recent changes
- Fix — minimal safe change with verification
- Prevent — action items, runbook updates, monitoring gaps closed
- LinkedIn: linkedin.com/in/philip-kumah-junior
Remote · UTC+0 · Open to Senior SRE / Platform Reliability / AWS Infrastructure roles



