Santa Clara, Canada On-site Employment

Gatik AI is hiring a Senior Cloud Infrastructure Engineer

Responsibilities

  • Architect and maintain mission-critical Kubernetes clusters optimized for heavy GPU/TPU workloads.
  • Implement and optimize Kubernetes-native GPU scheduling (NVIDIA GPU Operator) to ensure maximum hardware utilization.
  • Drive the "Everything as Code" philosophy using Terraform, Helm, and cloud-native tools.
  • Deploy Autonomous AI Agents (LangGraph, CrewAI) to monitor cluster health and enable automated triage of hardware failures and NCCL timeouts.
  • Build large-scale pipelines using Apache Airflow, Kafka, and Spark to process raw sensor data into training-ready formats.
  • Implement robust GitOps workflows using ArgoCD, Gitlab CI/CD to automate the deployment of both infrastructure and model artifacts.
  • Maintain deep visibility into infrastructure health and model serving performance using Prometheus, Grafana, and OpenTelemetry.
  • Develop agent-driven workflows to optimize the developer experience, such as automated PR reviewers for Terraform and AI agents that proactively suggest Kubernetes resource-limit adjustments based on model training telemetry.
  • Design and maintain MLFlow and feature store integrations to provide a robust system of record for every model iteration.
  • Build complex, automated model lifecycles using Airflow and Kubernetes to streamline the transition from training to simulation.
  • Support the deployment of models into simulation and production environments using Triton Inference Server, Ray Serve, and ONNX Runtime.
  • Enable researchers to scale models (VLA, World Models) across multi-node setups using PyTorch Distributed (TorchElastic), Ray Train, and Horovod.
  • Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for 3D Gaussian Splatting (3DGS) and large-scale training.
  • Partner with researchers to fine-tune performance across multi-node GPU clusters for FSDP and DeepSpeed workloads.

Requirements

  • 5+ years in Cloud Infrastructure, DevOps, or MLOps supporting high-scale compute environments.
  • Deep expertise in K8s, Helm, and container orchestration.
  • Strong background in Apache Airflow, Argo Workflows, MLFlow, and Terraform.
  • Practical experience supporting frameworks like Ray and PyTorch Distributed.
  • Proficiency in Python, Bash scripting, and a solid understanding of IAM/RBAC.

Nice to Have

  • Distributed Training Expertise: Deep understanding of FSDP, and DeepSpeed.
  • AI Agent Orchestration: Experience building Agentic Workflows (LangGraph, AutoGen) for infrastructure automation or data curation.
  • Advanced Protocols: Familiarity with Model Context Protocol (MCP) to connect AI agents with infrastructure tools.

Work Arrangement

On-site — Santa Clara, CA

Additional Information

  • This role is onsite 5 days a week at the Santa Clara, CA office.
Required Skills
KubernetesHelmPythonMCP
About company
Gatik AI
Gatik is the leader in autonomous middle-mile logistics, revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and launched the world’s first fully driverless commercial transportation service with Walmart.
All jobs at Gatik AI Visit website
Job Details
Category infrastructure
Posted 2 days ago