Responsibilities
- Deploy and oversee cloud infrastructure on AWS using modern automation practices and industry-leading tooling.
- Maintain and improve Kubernetes environments, CI/CD pipelines, and service mesh configurations to enable fast and stable software delivery.
- Create and manage scalable, secure, and resilient infrastructure for a multi-tenant SaaS product.
- Establish and track service level objectives, service level agreements, and error budgets to safeguard system availability.
- Develop infrastructure-as-code solutions using Terraform or equivalent tools for consistent and auditable deployments.
- Build automated internal tools to streamline provisioning, monitoring, and operational workflows.
- Ensure system health and performance through proactive monitoring of infrastructure and applications.
- Participate in on-call rotations, responding to incidents, diagnosing outages, and ensuring timely resolution and stakeholder communication.
- Lead incident response efforts during on-call shifts, coordinating actions across teams to meet service level targets.
- Improve incident management procedures to reduce detection and recovery times.
- Independently troubleshoot and resolve complex technical issues with minimal escalation.
- Collaborate with development teams to integrate observability, resilience, and reliability into service architecture.
- Automate operational procedures, health assessments, and alerting systems to minimize manual intervention.
- Enable safe and rapid releases through support for automated testing, canary rollouts, and rollback mechanisms.
- Support initiatives in security compliance, automated policy enforcement, and infrastructure cost efficiency.