Skip to content
← All jobs

Machine Learning Infrastructure Engineer

Together AI

Remote (US)RemoteFull-timeSenior4+ years$185,000 - $250,000

You'll live close to the hardware — kernel-level profiling, multi-node training frameworks, and the kind of debugging where the bug only appears at scale.

Requirements

  • 4+ years working on large-scale training or inference infrastructure
  • Comfortable in CUDA and profiling GPU utilization at the kernel level
  • Experience with multi-node training frameworks (DeepSpeed, Megatron, or similar)

Responsibilities

  • Push training throughput higher across our GPU cluster
  • Debug numerically subtle failures across thousands of GPUs
  • Build tooling that makes the cluster legible to researchers

Skills

  • Python
  • CUDA
  • PyTorch
  • Distributed systems
  • Kubernetes