Responsibilities
- Design and implement frameworks and self-service tools that empower teams to take full ownership of their service reliability within a 'You Build It, You Run It' model.
- Lead the development and execution of AI-powered operations strategies to automate issue detection, resolution, and preemptive failure prevention.
- Promote a culture of reliability by integrating SRE principles across engineering teams through design reviews, production readiness assessments, and operational best practices.
- Take command during high-severity incidents, demonstrating operational excellence and ensuring post-incident reviews result in actionable, blame-free improvements.
- Build and maintain comprehensive observability systems using Prometheus, Grafana, OpenTelemetry, and continuous profiling to proactively identify and resolve performance issues.
- Provide technical mentorship and knowledge transfer to SRE and product engineers to strengthen reliability practices across the organization.
Work Arrangement
Remote (Worldwide) — New York City, U.S., U.K., Finland, India, Singapore, Canada, Ireland