Vauxhall, London On-site Employment

Circadia Health is hiring a ML Ops Engineer

Responsibilities

  • Own and extend Circadia’s ML pipeline orchestration using Apache Airflow, including training, evaluation, and deployment workflows.
  • Build and maintain automated pipelines for model retraining, validation, and promotion across development, staging, and production environments.
  • Implement pipeline monitoring, alerting, and failure recovery to eliminate silent failures and ensure operational reliability.
  • Design pipeline architectures that support rapid experimentation while enforcing production-grade reproducibility.
  • Deploy and manage ML models on AWS infrastructure (e.g. AWS Batch for batch inference workloads).
  • Support deployment of models to edge devices, including Circadia’s clinical monitoring hardware, working with firmware and embedded engineering teams as needed.
  • Manage model versioning, promotion, and rollback workflows through the MLflow model registry.
  • Evaluate and implement strategies for safe model rollouts (e.g. shadow deployments, canary releases) as the platform matures.
  • Maintain and improve the MLflow-based experiment tracking and model registry infrastructure.
  • Establish conventions for experiment logging, artifact storage, model metadata, and lineage tracking.
  • Enable ML engineers to move seamlessly from experimentation to production deployment with minimal friction.
  • Implement and maintain training data versioning and dataset management practices to ensure reproducibility of model training runs.
  • Track dataset lineage, labeling provenance, and feature dependencies alongside model versions.
  • Collaborate with ML engineers and data engineers to formalise dataset release and validation workflows.
  • Build monitoring systems for model performance in production, including data drift detection, prediction quality tracking, and alerting on degradation.
  • Implement operational dashboards for pipeline health, compute utilisation, and deployment status.
  • Collaborate with data engineering to ensure upstream data quality and pipeline reliability for ML feature inputs.
  • Develop incident response procedures and runbooks for ML system failures.
  • Manage and optimise AWS compute resources (Batch, EC2, or similar) used for model training and inference.
  • Design infrastructure-as-code solutions for reproducible ML environments.
  • Drive cost optimisation across ML compute, storage, and data transfer.
  • Support Snowflake integrations for feature generation and training data pipelines.
  • Introduce and champion ML engineering best practices including CI/CD for models, automated testing for ML pipelines, and reproducible training workflows.
  • Build internal tooling and templates that accelerate the ML development-to-production cycle.
  • Document operational processes, architecture decisions, and onboarding materials for the ML platform.
  • Participate in architecture discussions and technical planning to ensure ML systems scale with Circadia’s growth.
  • Ensure all ML pipelines and infrastructure meet healthcare security and privacy requirements, including HIPAA and SOC 2.
  • Apply best practices for handling Protected Health Information (PHI) in training data, model artifacts, and inference outputs.
  • Maintain audit trails for model decisions, data access, and deployment history.
About company
Circadia Health

We equip providers with patient insights and clinical support that transform senior care. Our continuous, contactless early detection platform empowers timely interventions that lower readmissions. Harness the power of predictive care and improve outcomes for everyone.

Our mission is to build tools that reduce the risks and costs these incidents place on the entire healthcare ecosystem and those who interact with it. In doing so, we are democratizing tools that improve clinical decision-making.

Instead of supporting a reactionary model of care, Circadia makes predictive care possible with the world’s first FDA-cleared early detection platform. Our goal is simple: reduce rehospitalizations by helping providers see changes in their patient’s health before they worsen and disrupt treatment.

All jobs at Circadia Health Visit website
Job Details
Department Engineering
Category infrastructure
Posted 6 months ago