London, United Kingdom Hybrid Employment

PhysicsX is hiring a Principal Machine Learning Infrastructure Engineer

Responsibilities

  • Architect and manage distributed training systems for neural operator models such as Transolver and Point Cloud Transformer using NVIDIA DGX B200 hardware.
  • Enhance training pipeline efficiency by improving throughput, fault tolerance, and cost-effectiveness through techniques like gradient accumulation and multi-node synchronization.
  • Develop robust experiment tracking and observability tools to provide real-time insights into training progress, hyperparameter tuning, and model metrics.
  • Address performance limitations in data loading for large-scale mesh-based datasets.
  • Streamline data pipeline I/O operations from cloud storage via prefetching, caching, and optimized data formats.
  • Handle integration of diverse data sources with varying formats, structures, and resolutions.
  • Create scalable serving systems for pre-trained large physics models supporting zero-shot inference and uncertainty estimation using Monte Carlo Dropout.
  • Develop automated pipelines for packaging models for external deployment.
  • Ensure deployed models operate consistently in client environments and support fine-tuning workflows.
  • Guarantee full reproducibility of model behavior from any saved checkpoint.
  • Enhance research team productivity with tools enabling rapid iteration, stable CI/CD pipelines, and effective debugging capabilities.
  • Partner with infrastructure teams across the organization to align on shared standards and reusable patterns.
About company
PhysicsX
A deep-tech company focusing on AI-driven simulation software for engineering and manufacturing across advanced industries, with roots in numerical physics and Formula One.
All jobs at PhysicsX Visit website
Job Details
Department Research
Category infrastructure
Posted 11 days ago