Staff Software Engineer - Observability
Remote·Posted today
aiinfrastructuregorustawsgcpterraformkafka
At Snowflake, we are powering the era of the agentic enterprise. To usher in this new era, we seek AI-native thinkers across every function who are energized by the opportunity to reinvent how they work. You don’t just use tools; you possess an innate curiosity, treating AI as a high-trust collaborator that is core to how you solve problems and accelerate your impact. We look for low-ego individuals who thrive in dynamic and fast-moving environments and move with an experimental mindset — who rapidly test emerging capabilities to discover simpler, more powerful ways to deliver results. At Snowflake, your role isn't just to execute a function, but to help redefine the future of how work gets done. Staff Software Engineer - External Observability Platform Location: Bellevue, WA (Hybrid: 3 days/week in-office) Team: Infrastructure & Observability Platform Engineering About the Role Snowflake’s Data Cloud processes exabytes of data across multi-cloud global environments every day. Delivering seamless reliability and real-time visibility to thousands of global enterprise customers requires an Observability Platform built on hyper-scalable backend distributed systems. We are seeking a Staff / Lead Software Engineer to architect, design, and scale our External Observability Platform . In this role, you will lead the technical strategy for customer-facing telemetry, system metrics, audit logs, distributed tracing, and actionable operational insights. You will build high-throughput, low-latency infrastructure capable of ingesting, processing, and serving petabytes of telemetry data with strict SLA guarantees. You will join a team of world-class engineers in our Bellevue, WA office. To be successful, you must be deeply technical, capable of leading complex cross-functional architecture initiatives, and skilled at mentoring senior engineers while holding your own with the brightest technical minds in the industry. Key Responsibilities Architect & Scale Distributed Infrastructure: Lead the technical vision, design, and execution of Snowflake’s External Observability platform capable of handling trillions of events per day across multi-cloud deployments (AWS, Azure, GCP). Define Observability Standards: Standardize telemetry generation (metrics, logs, traces, and events) across all Snowflake application engineering teams, ensuring consistent schema, zero-data-loss ingestion, and optimized storage access. Build High-Throughput Engines: Write scalable, reliable, and testable backend services to process time-series data, high-cardinality metrics, and distributed traces at massive scale. Automate Infrastructure Lifecycle: Practice infrastructure-as-code (IaC) using tools like Terraform to deliver self-healing, automated telemetry pipelines and dynamic monitoring topology. Drive Platform Adoption & Diplomacy: Collaborate closely with Application Engineering, Security, and Customer Support teams to make systems measurable, translate telemetry into actionable customer-facing insights, and resolve cross-organizational technical dependencies. Technical Leadership & Mentorship: Drive engineering excellence through rigorous design reviews, technical roadmapping, performance tuning, and mentoring engineers across the broader organization. Incident Escalation & Root Cause Analysis: Serve as an expert troubleshooter for critical, complex system failure modes across large distributed clusters, conducting deep-dive post-mortems and building automation to eliminate repeat incidents. Minimum Qualifications 10+ years of professional experience building infrastructure and backend distributed systems at scale using languages such as Go, C++, Java, or Rust. Proven track record of architecting, deploying, and maintaining hyper-scale distributed platforms in public cloud environments (AWS, Azure, or GCP). Deep theoretical and practical CS fundamentals (data structures, algorithms, concurrency patterns, storage engines, distributed consensus, networking protocols). Strong experience with infrastructure-as-code (IaC) tools such as Terraform , or Pulumi. Hands-on expertise with time-series databases, high-cardinality metric stores, distributed tracing frameworks (OpenTelemetry, Jaeger), or log streaming systems (Kafka, Flink, ClickHouse, Prometheus/Thanos). Superior communication, collaboration, and diplomatic skills—ability to align cross-functional teams around architectural standards and lead technical decisions with empathy and clarity. Preferred Qualifications Massive Scale Experience: Prior experience in high-performance computing (HPC) or handling global installations processing petabytes of telemetry per day. Customer-Facing Observability: Experience building external or customer-exposed observability tools, APIs, and analytics dashboards with strict latency and access-control constraints. Network & Systems Deep Dive: Deep operational understanding of high-performance Linux kernel tuning, ebpf, eBPF-based profiling, or low-level network optimization. Multi-Cloud Expertise: Prior experience running cloud-agnostic platform infrastructure across AWS, Azure, and GCP simultaneously. Every Snowflake employee is expected to follow the company’s confidentiality and security standards for handling sensitive data. Snowflake employees must abide by the company’s data security plan as an essential part of their duties. It is every employee's duty to keep customer information secure and confidential. Snowflake is growing fast, and we’re scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake. How do you want to make your impact? For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com