Responsibilities
- Design and implement distributed data systems by defining APIs, data schemas, and replication strategies to ensure long-term durability and efficient querying of large-scale workflow history.
- Document architectural decisions and operational insights clearly to support reliable deployment and ongoing service management.
- Ensure system reliability and high performance by owning service-level objectives, developing chaos testing strategies, analyzing performance bottlenecks, and leading post-incident analysis.
- Provide technical leadership by decomposing roadmap initiatives, guiding mid-level engineers, and overseeing design documentation through the RFC process.
- Collaborate across teams, working closely with Server, Cloud, and Developer Experience groups to deliver end-to-end feature implementation.
Compensation
Competitive salary and equity package
Work Arrangement
Flexible work environment with remote options
Team
Part of a core infrastructure team focused on scalable data storage solutions
Responsibilities
- Design & build distributed data systems – craft APIs, schemas, and replication paths that keep petabytes of workflow history durable and query-able.
- Clearly document design choices and operational knowledge to successfully deploy and run service with those features.
- Drive reliability & performance – own SLOs, create chaos-test plans, profile hot paths, and lead incident reviews.
- Technical leadership – break down roadmap epics, mentor mid-level engineers, steward design docs through RFC.
- Cross-team collaboration – partner with the Server, Cloud, and DX teams to land features end-to-end.
Available for qualified candidates