Responsibilities
- Architect and implement secure, scalable, and highly available cloud environments on AWS with self-healing capabilities
- Use Terraform to define, provision, and manage AWS infrastructure as code
- Analyze and optimize cloud spending to improve cost efficiency
- Enforce system compliance and consistency using configuration management tools
- Lead the development and evolution of CI/CD workflows in collaboration with Development, QA, and operations teams
- Build and maintain automation tools for deployment, monitoring, and operational analysis of cloud systems
- Integrate automated workflows and infrastructure practices into development processes
- Administer CircleCI pipelines and assist developers with pipeline errors or infrastructure requests
- Implement monitoring, logging, and metrics solutions across AWS environments
- Configure and support Datadog monitoring, including alert integrations with Slack for Terraform events
- Enhance system observability and improve error reporting for network issues
- Manage version upgrades for MySQL databases and associated RDS services
- Conduct performance optimization and write SQL queries for infrastructure and application troubleshooting
- Operate and maintain DynamoDB instances within AWS
- Automate security policies, governance checks, and compliance validation in the cloud
- Serve as a point of contact with third-party vendors during service disruptions, including AWS, CircleCI, and Terraform
- Work with engineering teams to resolve production incidents, including off-hours support when required
- Mentor team members through informal knowledge sharing on AWS technologies
- Investigate and resolve escalated infrastructure performance problems
Other
Advanced English - Excellent verbal and written communication skills in English.