Software Engineer, Infrastructure & Performance
Remote·Posted today
aideveloper-toolsinfrastructuretypescriptgorustkubernetesterraform
About the role Dedalus Labs builds persistent computers for AI agents. Our flagship product, Dedalus Machines, gives agents an isolated environment where they can run software, keep files and state, and work over time. We're looking for an infrastructure engineer who's bothered by a slow build, an unexplained latency spike, or a workflow that makes engineers do the same work twice. You want to understand where the time went. Then you want to fix it, measure the improvement, and make sure everyone who uses that system benefits from it. At a small startup, those details affect how quickly the whole company can move. A build that wastes ten minutes wastes them repeatedly. A confusing API slows every product built on it. A deployment that needs someone to supervise it interrupts work elsewhere. You'll seek out these problems across our systems and codebase, including the ones people have learned to tolerate. Your work can bring product releases and customer commitments forward by weeks, and over time, months. You'll work directly with our CTO and engineers, with substantial freedom to investigate, experiment, and choose an approach. We want someone who takes pride in the craft of software engineering and cares about the systems they'll leave for the next person. That means useful abstractions, clear failure behavior, tests that establish the right guarantees, and documentation that explains the decisions. What you'll work on Make the development loop faster. Profile builds, tests, CI queues, artifact transfers, and local workflows. Find repeated work, unnecessary dependencies, poor scheduling, and cache misses. Follow a bottleneck into a compiler, linker, filesystem, network, or upstream dependency when the evidence takes you there. Establish a baseline, make the change, and verify the result on a representative workload. You'll use and improve our open-source tooling: Bessemer (BSMR) , our build system, and Hollywood , which generates GitHub Actions from typed TypeScript definitions and runs actions locally. This includes working on the tools themselves when their implementation limits what we can do. Experience with Buck2, Bazel, Blaze, remote execution, or build cache design is particularly relevant. Build platforms other engineers can build on. Design APIs, controllers, CLIs, and services that support multiple internal products and consumers. Think through resource lifecycles, concurrency, cancellation, permissions, and failure reporting. Work with the engineers using those interfaces, understand their constraints, and make common operations straightforward without hiding important system behavior. Make delivery reliable and understandable. Improve CI/CD, GitOps, environment provisioning, staging, production rollout, and recovery. Connect the source change, build inputs, artifact, and running software so engineers can trace a release and diagnose a failure. Instrument the path with useful metrics, logs, and traces. Use incidents and recurring manual work to identify what needs a software fix. Work with hardware we control. We have an in-office homelab where we test and debug on real machines and networks we control. Depending on your strengths, you can help assemble, provision, and maintain machines, configure networking, and build test infrastructure so engineers can reproduce failures and measure changes under conditions we understand. You'd be a good fit if A problem keeps your attention when the first few explanations turn out to be wrong. You read source, build a smaller reproduction, ask a better question, or collect another trace. You ask for help when it will move the investigation forward and keep responsibility for the outcome. You notice slowness and repeated effort even outside your immediate assignment. You investigate proactively and can explain why fixing it matters to the people using the system. You enjoy the last 20% of performance work, including the part that takes 80% of the effort. You can also judge when that effort is worth spending and when a deadline calls for a smaller, complete improvement. You're detail oriented and a little perfectionist. You care about API names, error messages, correctness, performance, and the next engineer's ability to understand your work. You can make a decision, finish it, and ship. You respect the craft and stay open to better ways of practicing it. You use modern AI coding tools to accelerate exploration, implementation, testing, and review. You understand the resulting code and take responsibility for its behavior. You work well with ambiguity. You can turn an incomplete goal into a concrete problem, agree on the constraints, and choose a useful next step. You enjoy the freedom to tinker and can turn an experiment into something the team depends on. Experience that helps You should have strong software engineering skills and experience building or operating infrastructure that other people rely on. We're particularly interested in depth in Go, Rust, C++, or another systems language, practical Linux debugging, and thoughtful API design. You'll also work with TypeScript in our tooling. Build systems, CI/CD, GitOps, Kubernetes, observability, and infrastructure as code are relevant areas of experience. Familiarity with Terraform is useful. Tell us where you've gone deep and what you learned from operating the system after you built it. Hands-on hardware and networking experience is a plus. You may have assembled and maintained machines, installed storage or network cards, provisioned Linux, or configured switches, routing, and VLANs. We're interested in people who understand how the parts fit together and can trace a problem across application code, the operating system, storage, and the network. Deeper experience with Ethernet, SFP+/QSFP optics and direct-attach cables, RDMA, or RoCE is useful additional depth. This role may be a poor fit if You need a complete specification and a fixed twelve-month roadmap before you can make progress. You prefer an assignment limited to one tool or layer and find it frustrating to follow a problem across software, infrastructure, and hardware. You routinely accept slow or manual workflows because they've always worked that way. You keep polishing after the agreed deadline or declare a performance improvement before measuring it. Location and compensation This is a full-time, in-person role in San Francisco. The annual base salary range is $180,000-$250,000 USD, depending on relevant experience and role scope. Relocation assistance and visa sponsorship are available. Show us your work Tell us about a system you improved because something about it bothered you. What did you notice? How did you find the cause? What changed, how did you measure it, and what did you choose to leave alone? A code sample, technical write-up, open-source contribution, or homelab project is welcome. A description of a private production project works too. Please keep confidential details out of your application. We'd like to understand your judgment, your persistence, and the care you put into the result. Explore our open-source repositories, particularly Hollywood and Bessemer (BSMR) , and star them if you'd like to follow their development. We especially welcome substantive contributions: a useful bug fix, a measured performance improvement, a thoughtful API change, or tests that establish an important guarantee. Start by discussing the problem and proposed scope with the maintainers. Working through an issue, implementation, and code review gives us a direct sense of how you investigate, explain your decisions, and respond to feedback. It also gives you a chance to experience how we work and whether you'd enjoy building with us. Contributing is an optional way to show your work, and we're equally interested in the work you've already done.