Responsibilities
- Monitors system health, performance metrics, and alerts to identify and respond to incidents promptly and diagnoses issues, troubleshoots problems, and restores services in a timely manner.
- Implements incident response processes to minimize downtime and improve system availability.
- Designs, develops, and maintains automation tools, scripts, and processes to streamline system management tasks, deployments, and configuration changes.
- Implements infrastructure-as-code principles to ensure consistency and repeatability.
- Optimizes system resources, configurations, and processes to enhance performance, scalability, and efficiency.
- Uses monitoring tools and performance testing to identify bottlenecks and implement optimizations.
- Collaborates with teams to forecast system resource needs, plans for capacity growth, and ensures adequate scalability.
- Leads incident response efforts, coordinates with cross-functional teams, and drives the resolution of system issues.
- Performs thorough post-incident analysis to identify root causes and implements preventive measures to minimize future incidents.
- Identifies opportunities for automation and drives the implementation of self-healing, monitoring, and deployment of automation tools and frameworks.
- Continuously improves operational efficiency, system reliability, and availability through process enhancements and automation.
- Ensures consistency across environments, tracks changes, and enforces configuration standards.
- Works closely with development teams, operations teams, and other stakeholders to ensure effective collaboration, knowledge sharing, and alignment on reliability goals.
- Implements security best practices, works with security teams to assess and address vulnerabilities, and ensures compliance with security standards and regulations.
- Performs any other related task as required.
Requirements
- Seasoned technical expertise in Linux/Unix systems, networking, and system administration.
- Seasoned proficiency in scripting or programming languages, such as Python, Go, Java, or Ruby.
- Seasoned knowledge of cloud platforms (such as AWS, Azure, or Google Cloud) and associated services.
- Seasoned proven expertise in performance monitoring, optimization, and troubleshooting using tools such as Prometheus, Grafana, or New Relic.
- Seasoned expertise in incident management, root cause analysis, and post-incident reviews
- Excellent problem-solving and analytical skills, with a keen attention to detail.
- Excellent communication, collaboration, and leadership skills.
- Seasoned ability to optimize system performance, scalability, and reliability. experience with performance monitoring and tuning tools (for example, Prometheus, Grafana, or New Relic) to identify bottlenecks, analyze performance data, and implement optimization strategies.
- Seasoned understanding of security principles, best practices, and compliance requirements. experience in designing and implementing security controls, performing security assessments, and ensuring compliance with industry standards.
Work Arrangement
On-site
