Responsibilities
- Tackle complex technical challenges using a broad set of modern technologies.
- Use software engineering principles to automate operations and enhance system dependability, performance, and fault tolerance.
- Create infrastructure that supports fast development cycles, spanning cloud platforms and embedded systems on spacecraft.
- Influence infrastructure strategy and promote high operational standards in containerized and modern environments.
- Deploy, manage, and support critical systems that power spacecraft operations and enterprise functions.
- Develop and refine Infrastructure as Code frameworks using tools like Terraform.
- Set up and maintain observability solutions including metrics, logs, and distributed tracing, with effective alerting.
- Design and sustain CI/CD pipelines that enable fast, secure, and consistent software releases.
- Collaborate with engineering teams to build scalable, reliable systems and ensure access to necessary tools and infrastructure.
- Detect, assess, and eliminate system bottlenecks and reliability threats through optimization and long-term fixes.
- Address production outages, conduct root cause investigations, and lead improvements via post-incident reviews.
- Participate in an on-call rotation to ensure continuous system availability and responsiveness.
- Travel periodically to customer locations and company sites for deployment, testing, or troubleshooting of key systems.
Work Arrangement
Hybrid — El Segundo, California, Washington, DC, Huntsville, AL
Other
Must be willing to work extended hours and weekends as needed