Victoria, British Columbia, Canada Remote (Country)

impact.com is hiring a Site Reliability Engineer

Responsibilities

  • Design and manage the observability framework by implementing OpenTelemetry across Java and C# services to capture accurate traces, metrics, and logs.
  • Develop integration tests for third-party social APIs and establish monitoring and alerting infrastructure to maintain system reliability and uptime.
  • Create and refine Grafana dashboards and alerting mechanisms focused on Latency, Traffic, Errors, and Saturation, optimized for JVM and .NET platforms.
  • Lead root cause investigations for failures in distributed systems and implement fixes through code improvements or infrastructure changes.
  • Use distributed tracing to detect inefficiencies in inter-service communication and streamline data flow from external APIs to internal storage.
  • Troubleshoot problems across all layers, including containerized applications (Java/C#), network interactions, and cloud infrastructure performance.
  • Evaluate application usage trends to guide capacity planning, enabling efficient handling of traffic surges while maintaining system stability and cost control.
About company
impact.com

impact.com transforms the way enterprises manage and optimize all types of partnerships.

The unified platform manages creators, affiliates, and referrals — powered by over $100B in partnership data and AI that turns discovery into performance.

impact.com is the AI-native partnership platform powering global commerce, enabling brands to see which partners influence AI and LLM recommendations and activate them within a trusted ecosystem.

All jobs at impact.com Visit website
Job Details
Category infrastructure
Posted a month ago