Responsibilities
- Design and manage cloud-native systems on GKE, Kubernetes, networking, and databases to support a highly available, low-latency financial platform handling real-time transactions across Europe.
- Ensure the reliability and scalability of over 100 microservices by defining service-level objectives, managing autoscaling, and implementing resilience techniques such as circuit breakers and graceful degradation.
- Lead continuous deployment practices with fully automated pipelines, including rollout and rollback strategies and scalable deployment observability.
- Implement comprehensive observability using metrics, distributed tracing, and structured logs to reduce mean time to resolution and empower engineers to independently diagnose issues.
- Maintain and evolve infrastructure-as-code using Terraform and Helm, enforcing strict peer review, version control, and auditability with no manual interventions.
- Promote security and compliance through Zero Trust principles, Workload Identity, dynamic secret management with Vault, network policies, and readiness for ISO27001 and SOC2 audits.
- Lead incident response for networking, load balancing, Kubernetes, and cloud infrastructure, and drive follow-up improvements to prevent future occurrences.
- Elevate engineering standards by contributing to architecture decisions, reviewing platform changes, and mentoring early-senior engineers.
Compensation
Top-of-market compensation package, including equity
Work Arrangement
Hybrid — Berlin, with 5 offices across Europe
Not specified