Senior Member of Technical Staff, Infrastructure

Remote·Posted today
pythonawsterraform
About the Role Join the engineering team at an AI/ML research company to build and scale infrastructure for research workloads. You will own cloud compute, distributed systems, storage, and experiment orchestration, helping improve the platform's reliability, performance, and resource efficiency. What You'll Do Build and manage cloud compute infrastructure that supports research workloads at scale. Design and optimize distributed workflow scheduling for performance and cost efficiency. Develop recovery safeguards for worker interruptions, application crashes, and partial results while preventing duplicate work. Improve monitoring, logging, and diagnostic tools to help identify and resolve system failures. Partner with researchers and engineers to develop capabilities and shape infrastructure architecture. What We're Looking For 5 to 10+ years of relevant experience operating production infrastructure or services. Strong Python and Linux skills, including scripting, process management, and debugging. Experience with AWS or another major cloud platform, including compute, storage, networking, and access management. Understanding of distributed systems concepts such as queues, timeouts, retries, and idempotency. Experience investigating production issues and balancing reliability, complexity, and cost. Familiarity with Terraform or similar infrastructure-as-code tools and monitoring platforms such as CloudWatch, Prometheus, or Grafana. Experience with virtualization, VM images and snapshots, desktop application automation, batch scheduling, experiment orchestration, or sandboxing untrusted code is a plus. Compensation & Benefits Competitive compensation, equity, health, dental, and vision coverage, plus visa sponsorship and relocation support. Location This is a full-time, on-site role in Zürich, Switzerland.