← All Positions

Machine Learning Infrastructure Engineer

fast-scaling AI startup San Francisco, CA Direct Hire Other $200,000 - $300,000

Job Summary

We are building large physics foundation models trained on weather data to advance causal intelligence and predictive AI. This role focuses on designing and optimizing the distributed training and inference infrastructure that powers these massive models. You will tackle complex challenges in scaling ML systems for petabyte-scale datasets and frontier foundation models across physics-related domains.

Essential Functions

  • Design, deploy, and maintain large distributed ML training and inference clusters
  • Develop efficient, scalable end-to-end pipelines to manage petabyte-scale datasets and model training throughout the entire ML lifecycle
  • Research and test various training approaches including parallelization techniques and numerical precision trade-offs across different model scales
  • Analyze, profile and debug low-level GPU operations to optimize performance
  • Stay up-to-date on research to bring new ideas to work

Required Qualifications

  • 3-10 years of experience building large-scale ML infrastructure for core foundation models
  • Experience building ML infrastructure for core foundation models, not fine-tuning
  • Background working at a science or physical AI company such as self-driving, robotics, or biology
  • Experience as a generalist working across the ML lifecycle
  • Deep expertise in optimizing large-scale training and inference workloads
  • Proficiency with distributed training frameworks such as FSDP or DeepSpeed
  • Experience with low-level GPU performance optimization and debugging including CUDA and JAX
  • Demonstrated high intentionality in career choices and mission-driven focus

Preferred Qualifications

  • Background in physics, robotics, biology, or AI at the frontier of these fields
  • Experience with multimodal data and large GPU infrastructures
  • Strong grasp of state-of-the-art techniques for optimizing training and inference workloads
  • Familiarity with monitoring, logging, observability, and version control best practices for ML systems

Technical Skills

  • Distributed training frameworks (FSDP, DeepSpeed)
  • NVIDIA GPU chips and low-level optimization
  • CUDA and JAX
  • Python and C++
  • Linux operating systems
  • Containerization and orchestration (Kubernetes, Docker)
  • Cloud platforms (GCP, AWS, Azure)
  • FPGA and control systems
  • Scalable model serving and deployment architectures

Compensation & Benefits

Base salary range of 200000 to 300000 depending on experience and qualifications. Competitive compensation for mission-driven engineers who want to work on frontier AI infrastructure.

Apply for This Position

📄 Drag and drop or browse PDF, DOC, or DOCX (10 MB max)