Responsibilities
- Manage, sustain, and optimize live cloud systems on AWS, Azure, GCP, Windows, or hybrid setups.
- Develop and oversee monitoring, log collection, metrics tracking, distributed tracing, dashboards, and alert systems for live services.
- Expand observability to cover infrastructure components, apps, databases, message queues, and network connections.
- Refine alert configurations to minimize distractions, enhance detection accuracy, and enable effective incident handling.
- Engage in resolving live incidents, diagnosing production issues, determining root causes, and implementing corrective actions.
- Develop automated solutions for infrastructure management, deployment processes, health verification, runbook execution, and recovery procedures.
- Collaborate with development teams to establish service level objectives, indicators, error budgeting, and operational preparedness criteria.
- Assist with cloud networking, domain name systems, TLS configuration, load distribution, identity and access management, storage, compute resources, and managed services.
- Enhance the dependability, uptime, responsiveness, and growth capacity of systems hosted in cloud environments.
- Sustain Infrastructure as Code frameworks and configuration control methods to ensure consistent, reproducible environments.
- Develop and update runbooks, operational guides, and escalation protocols.
- Detect potential production risks and lead efforts to resolve them through automation, architectural upgrades, and platform-wide standards.