Responsibilities
- Create and manage Grafana dashboards to visualize system performance, latency, error rates, saturation, and platform health metrics.
- Set up and oversee observability tools like Prometheus, ensuring accurate tracking of critical service metrics and SRE Golden Signals.
- Build, validate, and update reusable Ansible playbooks for automating infrastructure deployment, configuration management, patching, and operational tasks.
- Enable enterprise-level automation using AWX by managing centralized job execution, templates, role-based access control, and workflow reuse.
- Manage Infrastructure as Code repositories in Git, applying version control best practices, code reviews, and CI/CD integration.
- Engage in Agile processes such as Sprint Planning, Daily Stand-ups, Backlog Refinement, and Retrospectives to support team execution and continuous improvement.
- Work closely with Product Owners, Scrum Masters, and engineering teams to convert business needs into technical user stories and advance automation and reliability efforts.