Responsibilities
- Lead, hire, mentor, and grow a team of platform engineers focused on delivering secure, scalable, and resilient infrastructure and developer tools.
- Promote a team environment built on collaboration, inclusivity, and ongoing knowledge exchange.
- Design and support individual learning and development plans for team members.
- Develop and execute technical strategies for new systems and architectural initiatives.
- Ensure the team consistently meets high standards for production stability and developer experience.
- Serve as the primary liaison between the engineering team and internal stakeholders or cross-functional groups.
- Drive the adoption and evolution of Site Reliability Engineering practices across teams in partnership with service owners.
- Establish and advocate for SLO/SLI frameworks and use error budgets to guide trade-offs between reliability and feature development.
- Shape the vision for internal developer platforms, including self-service infrastructure, tooling, and deployment methods to increase team efficiency.
- Oversee team resourcing, capacity planning, and staffing alignment to meet current and future demands.
- Organize, staff, and manage on-call schedules for incident support.
- Oversee incident management processes, including resolution, root cause analysis, and implementation of preventive improvements.
Work Arrangement
Remote
Team
Cloud Engineering
Responsibilities
- Lead, hire, mentor, and grow a team of platform engineers focused on delivering secure, scalable, and resilient infrastructure and developer tools.
- Promote a team environment built on collaboration, inclusivity, and ongoing knowledge exchange.
- Design and support individual learning and development plans for team members.
- Develop and execute technical strategies for new systems and architectural initiatives.
- Ensure the team consistently meets high standards for production stability and developer experience.
- Serve as the primary liaison between the engineering team and internal stakeholders or cross-functional groups.
- Drive the adoption and evolution of Site Reliability Engineering practices across teams in partnership with service owners.
- Establish and advocate for SLO/SLI frameworks and use error budgets to guide trade-offs between reliability and feature development.
- Shape the vision for internal developer platforms, including self-service infrastructure, tooling, and deployment methods to increase team efficiency.
- Oversee team resourcing, capacity planning, and staffing alignment to meet current and future demands.
- Organize, staff, and manage on-call schedules for incident support.
- Oversee incident management processes, including resolution, root cause analysis, and implementation of preventive improvements.
Not specified