▣ Featured article
View all writing →From Splunk to Kubernetes:
Rethinking Failure Detection
Eight years of SRE alert fatigue shaped how I think about failure detection, rollout safety and reconciliation.
I work at the intersection of Kubernetes networking, Linux infrastructure and SRE to build resilient platforms, improve developer experience and reduce operational toil.
Eight years of SRE alert fatigue shaped how I think about failure detection, rollout safety and reconciliation.
Reproducible hands-on environments for Kubernetes networking, failure analysis and automated validation.
Linux operations, production support, monitoring and infrastructure automation.
Incident response, observability and automation in banking production environments.
Kubernetes networking, Linux internals, Go tooling, Cilium and eBPF.
I’m open to conversations about Platform Engineering, Kubernetes, infrastructure and remote engineering roles.
↗Connect on LinkedIn