Senior ML Storage Infrastructure Engineer
Zoox
Software Engineering, Other Engineering, Data Science
Seattle, WA, USA
USD 176k-288k / year + Equity
Posted 6+ months ago
Zoox is looking for a software engineer to work on our custom High-Performance Computing infrastructure and its supporting ecosystem of tools and services. This infrastructure is central to machine learning workflows across all Zoox software divisions, from data engineering to computer vision perception to simulation and more. You will take on a breadth of end-to-end responsibilities including distributed system design, algorithmic job scheduling, and adaptive cloud scaling in support of all of Zoox’s computational needs.
In this role, you will:
- Design, build, and optimize a petabyte-scale, in-house HPC storage infrastructure, ensuring high performance and reliability for our machine learning workloads across both cloud and on-premise data centers.
- Drive GPU efficiency by strategically collocating storage and compute, architecting a storage layer that keeps tens of thousands of GPUs fully utilized and prevents bottlenecks.
- Drive key initiatives in training and storage optimization by partnering with ML practitioners, applying your deep understanding of frameworks such as PyTorch and TensorFlow to meet their evolving demands.
- Investigate and adopt new distributed system paradigms and cutting-edge technologies to ensure our infrastructure can scale to meet ever-growing computational and storage demands.
- Create production-grade web service APIs, SDKs, and other essential tools to deliver a world-class developer experience for all software teams at Zoox.
Qualifications:
- Experience designing and building high-performance, distributed storage systems (object/file) for large-scale, GPU-bound workloads.
- Proficiency in Python, Java, or similar languages for developing data-intensive, high-performance applications.
- Hands-on experience with cloud platforms (AWS, GCP, Azure), using their storage, GPU, and observability services to provide usage showback for ML practitioners.
- Bachelor's degree in Computer Science or a related field with a strong foundation in data structures and systems design.
Bonus Qualification:
- Experience with parallel filesystems (e.g., Lustre, FSx) and their integration with container orchestrators via Kubernetes CSI drivers.
- Deep knowledge of ML frameworks like PyTorch and TensorFlow, and workload schedulers such as SLURM or Kubernetes.
- Familiarity with emerging AI paradigms, including agentic systems, and observability tools like OpenTelemetry.
About Zoox
Zoox is developing the first ground-up, fully autonomous vehicle fleet and the supporting ecosystem required to bring this technology to market. Sitting at the intersection of robotics, machine learning, and design, Zoox aims to provide the next generation of mobility-as-a-service in urban environments. We’re looking for top talent that shares our passion and wants to be part of a fast-moving and highly execution-oriented team.
Accommodations
If you need an accommodation to participate in the application or interview process please reach out to accommodations@zoox.com or your assigned recruiter.
A Final Note:
You do not need to match every listed expectation to apply for this position. Here at Zoox, we know that diverse perspectives foster the innovation we need to be successful, and we are committed to building a team that encompasses a variety of backgrounds, experiences, and skills.