← All jobs
Machine Learning Infrastructure Engineer
Together AI
Remote (US)RemoteFull-timeSenior4+ years$185,000 - $250,000
You'll live close to the hardware — kernel-level profiling, multi-node training frameworks, and the kind of debugging where the bug only appears at scale.
Requirements
- 4+ years working on large-scale training or inference infrastructure
- Comfortable in CUDA and profiling GPU utilization at the kernel level
- Experience with multi-node training frameworks (DeepSpeed, Megatron, or similar)
Responsibilities
- Push training throughput higher across our GPU cluster
- Debug numerically subtle failures across thousands of GPUs
- Build tooling that makes the cluster legible to researchers
Skills
- Python
- CUDA
- PyTorch
- Distributed systems
- Kubernetes