Responsibilities
- Overseeing Kubernetes clusters deployed across multiple environments and geographic regions
- Taking full ownership of infrastructure defined and managed through code for all system resources
- Enhancing and maintaining continuous integration and delivery pipelines using GitOps practices
- Optimizing high-throughput data pipelines that handle billions of daily events via distributed queues and stream processing systems
- Designing and implementing comprehensive monitoring, alerting, and system observability frameworks
- Investigating and resolving production issues spanning multiple services and components
- Overseeing cloud spending and planning capacity to ensure efficient resource utilization
- Collaborating directly with a compact engineering team while maintaining end-to-end ownership of the entire infrastructure stack
Benefits
- Medical, dental, and vision insurance fully covered by the company for employees and their families
- 401(k) plan with company match up to 4%, fully vested from day one
- Paid parental leave of 20 weeks for primary caregivers and 12 weeks for secondary caregivers
- Monthly stipend of up to $85 to support internet and mobile expenses for remote work
- Company-provided short-term, long-term disability, and life insurance coverage
Work Arrangement
On-site
Team
Small engineering team with full infrastructure ownership model
Full ownership — small team, you own the entire infrastructure, not a slice of it
You will have complete responsibility for the infrastructure in a lean team environment, with no shared ownership boundaries
Real scaling challenges — bursty scraping workloads, cache invalidation, multi-region, millions of daily requests
The system faces demanding scalability requirements including unpredictable scraping loads, cache management, multi-region operations, and handling millions of requests each day
AI-native company — your infra directly powers AI agents used by leading companies in the space
Infrastructure work directly enables AI-powered agents used by top organizations, placing reliability and performance at the core of product success