Senior MTS, Infrastructure

Remote·Posted today
infrastructurepythonterraform
About the Role Join an engineering team building and scaling cloud infrastructure for AI research workloads. You will own key systems across compute, distributed workflows, data storage, and experiment orchestration, helping improve platform reliability, performance, and efficiency. What You'll Do Build and operate cloud compute infrastructure that supports research workloads at scale. Design and optimize distributed workflow scheduling for performance and cost efficiency. Develop recovery safeguards for worker interruptions, application crashes, and partial results while preventing duplicate work. Improve monitoring, logging, and diagnostic tools to help identify and resolve system failures. Partner with researchers and engineers to develop platform capabilities and shape infrastructure architecture. What We're Looking For Typically 5 to 10 or more years of relevant experience in infrastructure or software engineering. Strong Python and Linux skills, including scripting, process management, and debugging. Experience operating services on a major cloud platform, including compute, storage, networking, and access management. Understanding of distributed systems concepts such as queues, timeouts, retries, and idempotency. Experience investigating production issues and balancing reliability, complexity, and cost. Self-directed, collaborative approach, with sound judgment and strong follow-through. Experience with KVM/QEMU, VM images or snapshots, desktop application automation, Terraform, monitoring tools, batch scheduling, experiment orchestration, or sandboxing untrusted code is beneficial. Compensation & Benefits Equity is offered, and visa sponsorship is available. Location On-site in Zürich, Switzerland.