Responsibilities
- Oversee the full lifecycle of Site Reliability Engineering programs, covering incident response, software releases, monitoring, automated processes, and infrastructure scaling to ensure system stability and efficiency
- Facilitate collaboration among software engineering, SRE, DevOps, and business teams to align goals, address obstacles, and monitor progress on technical milestones and interdependencies
- Develop and maintain status reports, dashboards, and performance metrics to communicate program health, key achievements, risks, and results to leadership and stakeholders
- Collaborate with incident leads during outages, support post-mortem analyses, drive root cause investigations, and ensure corrective actions are implemented to prevent recurrence
- Convey technical program updates, risks, and business implications clearly to non-technical audiences, ensuring alignment across organizational levels