Responsibilities
- Manage, support, and optimize cloud environments deployed on AWS, Azure, GCP, Windows, or mixed configurations.
- Develop and sustain monitoring systems, log aggregation, metric collection, distributed tracing, dashboards, and alerting mechanisms for live services.
- Expand visibility into system behavior across infrastructure layers, applications, databases, messaging queues, and network components.
- Refine alerting rules to minimize false positives, enhance detection accuracy, and support effective incident handling.
- Engage in incident management, diagnose live system issues, conduct root cause investigations, and implement corrective actions.
- Develop automated solutions for infrastructure management, deployment pipelines, health validation, operational runbooks, and failure recovery.
- Collaborate with development teams to establish service-level objectives, service-level indicators, error budget policies, and readiness criteria.
- Provide operational support for cloud networking, domain name systems, TLS certificates, load distribution, identity access management, storage, compute resources, and managed services.
- Enhance system dependability, uptime, responsiveness, and capacity to scale under load in cloud-deployed architectures.
- Sustain infrastructure-as-code frameworks and configuration control methods to enable consistent, reproducible environments.
- Develop and update operational guides, runbooks, and escalation protocols for incident management.
- Detect potential production vulnerabilities and lead mitigation efforts via automation, architectural refinements, and platform-wide standards.