Responsibilities
- Manage and run customer-facing infrastructure, including Refinery as a Service and private cloud deployments across multiple AWS accounts and regions.
- Develop and sustain Terraform modules, Helm charts, and automated deployment systems for EKS clusters, collector groups, and Refinery instances.
- Create monitoring, alerting, and observability solutions for managed services, using internal tools to monitor system health.
- Handle scaling, updates, and incident response for customer environments, including planning capacity and optimizing cloud costs.
- Develop self-service tooling for automated deployment and management of field-operated services.
- Act as the top-tier technical contact for complex customer issues, including production outages, collector configurations, Refinery performance, and architectural reviews.
- Troubleshoot deep technical problems across distributed systems, Kubernetes, AWS networking components, and polyglot service meshes.
- Work directly with customer engineering and platform teams to resolve live production incidents under pressure, especially when business impact is significant.
- Participate in an on-call schedule for managed services, offering Tier 2 support for infrastructure-related customer issues.
- Develop and maintain standard operating procedures, runbooks, and diagnostic frameworks to speed up issue resolution for support teams.
- Support and enhance OpenTelemetry components such as distributions, collectors, exporters, and instrumentation libraries used by customers.
- Engage with the OpenTelemetry community by joining SIGs, reviewing code, triaging issues, and promoting best practices.
- Produce reference designs, sample configurations, and integration guides that show effective instrumentation in common environments like Kubernetes and serverless.
- Identify shortcomings in open source tools that hinder customer success and either contribute fixes or build interim solutions.
- Add features and improvements to open source projects like Refinery and the Honeycomb Collector Distro to support managed service functionality.
- Join technical calls when customer engagements move beyond demos, helping validate architectures and resolve live environment issues to close deals.
- Collaborate with Solutions Architects on key accounts, leading discussions on infrastructure and data pipelines while they focus on product storytelling.
- Lead technical workshops for customers, including architecture reviews, SLO design, and instrumentation deep dives, especially in complex setups.
- Take hands-on leadership in proof-of-concepts and pilot projects, setting up collectors, configuring pipelines, and validating integrations in real customer environments.
- Establish feedback mechanisms from field operations to product and engineering teams, highlighting patterns that influence development priorities.
- Build internal tools and dashboards that improve operational efficiency, such as deployment trackers, rule editors, and monitoring interfaces.
- Work closely with Solutions Architecture, Customer Success, and Support to deliver technical depth on challenging accounts.
- Partner with Product and Engineering teams on critical bugs, missing features, and integration hurdles, providing real-world operational insights.
- Support field enablement by training internal teams on advanced diagnostics, collector setup, Refinery internals, and evolving reliability practices.
Work Arrangement
Remote (Worldwide)