Santa Clara, Canada On-site Employment

Gatik AI is hiring a Senior AI Infrastructure Engineer

Responsibilities

  • Architect and deploy high-performance AI platforms that support large-scale autonomous driving models
  • Bridge research and production environments to ensure scalability, reliability, and efficiency in AI systems
  • Enable researchers to scale advanced models like Vision-Language Architectures and World Models using distributed frameworks such as PyTorch Distributed and Ray Train
  • Optimize multi-GPU cluster performance through efficient model and data parallelism strategies on H100 and A100 hardware
  • Tune low-level networking protocols including NCCL, InfiniBand, and RoCE v2 to reduce communication latency during training and 3DGS workloads
  • Improve GPU utilization and cost efficiency using Kubernetes-native scheduling tools like NVIDIA GPU Operator and KubeFlow
  • Enhance inference pipelines by deploying optimized models via TensorRT, ONNX Runtime, and Triton Inference Server for real-time and batch scenarios
  • Build self-monitoring systems using AI agents (e.g., LangGraph, CrewAI, AutoGen) to detect and respond to hardware failures and NCCL timeouts automatically
  • Develop automated DevOps workflows using AI agents that review infrastructure code and recommend Kubernetes resource tuning based on model needs
  • Support the creation of autonomous data curation systems where AI agents identify, label, and validate critical edge cases from raw sensor data
  • Automate the end-to-end machine learning lifecycle using tools like MLFlow, Argo Workflows, and Kubernetes
  • Integrate experiment tracking and feature stores to maintain a reliable record of all model versions and training runs
  • Implement safe deployment patterns such as A/B testing, shadow deployments, and automated rollback procedures
  • Enforce infrastructure consistency and reproducibility through Infrastructure-as-Code using Terraform and Helm
  • Collaborate on scaling ETL workflows with Apache Airflow, Kafka, and Spark for efficient data processing
  • Work with data engineering teams to build high-throughput data pipelines that move sensor data to cloud storage like S3, GCS, or Delta Lake
  • Define and monitor key performance indicators such as training speed, latency, throughput, and model drift
  • Ensure comprehensive system visibility using monitoring tools including Prometheus, Grafana, OpenTelemetry, and the ELK Stack
  • Establish deep observability across both infrastructure layers and ML metrics like convergence and throughput
  • Track and alert on AI-specific KPIs including inference latency, throughput, and feature drift

Work Arrangement

On-site — Santa Clara, CA

Other

This position requires full-time, in-person attendance from Monday to Friday at the Santa Clara, CA office.

About company
Gatik AI
Gatik is the leader in autonomous middle-mile logistics, revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and launched the world’s first fully driverless commercial transportation service with Walmart.
All jobs at Gatik AI Visit website
Job Details
Category infrastructure
Posted 23 days ago