I keep cloud systems observable, resilient, and dependable at 3 a.m. Strong today in observability, CI/CD, and Azure/AWS infrastructure. Building depth in Kubernetes, Go, and OpenTelemetry. Looking for an SRE seat on a product team.
In one minute | Skills | How I work | Practice | Contact
- What I do: Reliability engineering for cloud systems. Monitoring that explains failures, pipelines that ship safely, automation that removes human error from operations.
- Where I'm strong: Observability (Dynatrace, Prometheus, Grafana), CI/CD (Azure DevOps, GitHub Actions), infrastructure as code (Terraform, Ansible) on Azure and AWS.
- What I'm building next: Kubernetes depth, Go services, OpenTelemetry instrumentation, SLO and error-budget practice.
- What I want next: A Site Reliability or Platform Engineering role on a product team where reliability is measured, not promised.
| Domain | Tools |
|---|---|
| Observability | Prometheus, Grafana, Dynatrace, LogicMonitor, Azure Monitor, OpenTelemetry |
| Containers and Orchestration | Docker, Kubernetes, Helm |
| Infrastructure as Code | Terraform, Bicep, Ansible |
| CI/CD | Azure DevOps, GitHub Actions |
| Cloud | Azure, AWS |
| Languages and Scripting | Go, Python, Bash, TypeScript, C# |
| Systems | Linux internals, networking fundamentals, distributed systems patterns |
- Define SLIs first, promise uptime later. If we can't measure it, we have no business promising it.
- Treat error budgets as agreements. When the budget burns down, features wait and reliability wins.
- Automate repeated work. Doing a task twice by hand is fine. The third time it becomes a pipeline.
- Roll out gradually. Flags, canaries, staged rollouts. Blast radius is chosen in advance, not discovered during an incident.
- Write blameless postmortems. Incidents come from system gaps, not bad people. Fix the gap.
- Keep toil under control. Operational grunt work should stay well under half the job. The rest goes to engineering.
- Running OpenTelemetry's Astronomy Shop microservices locally, injecting faults and tracing failures end to end so good instrumentation becomes routine
- Working through Google's SRE Book and SRE Workbook, turning each chapter into checklists I can apply on real systems
- Keeping the algorithmic edge sharp on Codeforces and LeetCode, since debugging distributed systems still comes back to reasoning about code
- CKA certification and hands-on chaos engineering are next in line
Currently: building reliability tooling that pairs monitoring data with AI assistance, so on-call people get fewer and better alerts instead of more noise.
The fastest way to reach me is LinkedIn.
- LinkedIn: linkedin.com/in/benito-j-d-095a63294
- Phone: +91 8870764795
- Portfolio: benify.in
- LeetCode: leetcode.com/u/Benito-JD
- Codeforces: codeforces.com/profile/Benito_JD



