← All jobs
Cloud Engineer, Platform Reliability
Datadog
New York, NYHybridFull-timeMid-level4+ years$164,000 - $212,000
You'll work on the reliability of the platform that other companies depend on to know when THEIR systems are broken — the bar is high on purpose.
Requirements
- 4+ years operating cloud infrastructure at meaningful scale
- Deep familiarity with observability tooling — ideally as a heavy user, not just an operator
- Comfortable writing Go for internal tooling
Responsibilities
- Improve ingestion pipeline reliability as query volume grows
- Build internal tooling the SRE team uses to triage incidents faster
- Own capacity planning for a core storage tier
Skills
- Go
- AWS
- Kubernetes
- Terraform
- Prometheus