Responsibilities
- Build and sustain tooling that enables effective monitoring, alerting, and service reliability.
- Partner with engineering teams to integrate reliability principles into service design and promote proper automation use.
- Find and execute on ways to enhance the scalability, efficiency, and dependability of services.
- Collaborate with infrastructure teams to grasp operational challenges and build solutions that ease service management.
- Engage in incident response and post-incident reviews to resolve root causes and prevent recurrence.
- Assess emerging technologies and industry standards to refine SRE tools and incident handling processes.
- Develop deep expertise in critical system components, including services, infrastructure, products, tools, and workflows.
- Lead urgent incident resolution efforts and support junior engineers in mastering incident management.
Work Arrangement
Remote (Country) — Brazil