Responsibilities
- Ensure high levels of reliability, uptime, performance, and scalability for the AWS-based SaaS platform, meeting customer-facing SLA and SLO commitments.
- Establish and manage Service Level Objectives and Indicators across production services, using error budgets to guide engineering priorities.
- Lead incident management efforts, including triage, resolution coordination, stakeholder communication, and post-mortem analysis with defined follow-up actions.
- Promote a proactive approach to system reliability by identifying potential failures early through load testing, chaos engineering, and failure analysis.
- Design and manage cloud infrastructure on AWS using infrastructure-as-code tools such as AWS CDK or Terraform to ensure consistency, auditability, and security.
- Improve and maintain CI/CD pipelines using GitHub Actions to support fast, safe, and repeatable deployments across environments.
- Administer and fine-tune core AWS services including ECS, EC2, Aurora RDS, DynamoDB, Lambda, S3, SQS, EventBridge, Cognito, Secrets Manager, and CloudFront.
- Implement comprehensive observability by centralizing logs, metrics, traces, and alerts via CloudWatch, Sentry, and similar tools for rapid issue detection and response.
- Oversee capacity planning, cloud cost optimization, and governance of cloud spending across the platform.
- Ensure infrastructure and operations comply with HIPAA, SOC 2, and HITRUST standards due to handling of protected health information.
- Collaborate with security teams on managing vulnerabilities, securing systems, managing secrets, and enforcing access controls.
- Maintain and routinely test disaster recovery and business continuity strategies, aligned with defined recovery time and point objectives.
- Support compliance audits by preparing documentation and evidence for certification requirements.
- Lead, mentor, and develop a mixed team of full-time SREs/DevOps engineers and offshore contractors, promoting ownership and continuous improvement.
- Manage distributed team operations by setting communication norms, documentation practices, and handover procedures for effective collaboration.
- Conduct regular one-on-ones, define performance goals, and support career development for team members.
- Sustain a balanced on-call schedule with proper tooling, runbooks, and escalation paths to prevent burnout and maintain team health.
- Partner with software engineering teams to integrate reliability practices into the development lifecycle, including readiness reviews and deployment standards.
- Work with product and engineering leaders to align feature development speed with operational risk and technical debt management.
- Influence platform architecture decisions by providing operational insights on new services and major technical changes.
- Advocate for automation to reduce manual work through scripting, tooling, and process optimization.
Benefits
- Flexible and collaborative remote work environment
- Unlimited paid time off
- Stock options offering equity participation
- Access to an expanding suite of innovative tools and technologies
- Competitive salary and comprehensive benefits
- Opportunity to contribute meaningfully to a pioneering digital health solution
Compensation
Competitive salary and benefits package
Work Arrangement
Primarily remote
Team
Distributed team with full-time and offshore members
Other
- Role is primarily remote with occasional travel requirements.
- Expect multiple video calls to get to know you.
- All equipment comes directly from us—no purchases required.
Not specified