Senior Data Platform Engineer

India, Remote·Posted today
mldata-platformspythonawsgcpterraformkafkasparkragllm
<div class="content-intro"><p style="line-height: 1.4;"><span style="color: rgb(0, 0, 0); font-family: arial, helvetica, sans-serif; font-size: 12pt;"><strong>Who we are</strong></span></p> <p style="line-height: 1.4;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">DigiCert is a global leader in intelligent trust. We protect the digital world by ensuring the security, privacy, and authenticity of every interaction. Our AI-powered DigiCert ONE platform unifies PKI, DNS, and certificate lifecycle management, to secure infrastructure, software, devices, messages, AI content and agents. Learn why more than 100,000 organizations, including 90% of the Fortune 500, choose DigiCert to stop today’s threats and prepare for a quantum-safe future at&nbsp;<a href="http://www.digicert.com/">www.digicert.com</a></span></p></div><p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><strong>Job summary</strong></span></p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">We are looking for a Senior Data Platform Engineer to design, build, and operate the foundational data and ML infrastructure that powers analytics, reporting, and machine learning across DigiCert. This role sits at the intersection of data engineering and ML platform work — you will own the systems, pipelines, and tooling that data scientists, analysts, and engineers rely on every day. You bring deep expertise in Databricks, a strong engineering mindset, and hands-on experience building and maintaining ML pipelines in production. You will partner closely with data science, analytics, product, and business teams to deliver a platform that is reliable, governed, and built to scale.</span></p> <p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><strong>What you will do</strong></span></p> <ul> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Design, build, and maintain scalable data and ML pipelines using Python and SQL, processing large-scale datasets across batch and streaming workloads</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Own and evolve core platform infrastructure on Databricks — including Delta Lake table architecture, Unity Catalog governance, Databricks Workflows orchestration, and compute optimization</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Build and maintain end-to-end ML pipelines: feature engineering, model training pipelines, experiment tracking (MLflow), and model deployment/serving infrastructure</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Collaborate with data scientists to operationalize models — bridging the gap between experimentation and production-grade ML systems</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Define and enforce data platform standards: ingestion patterns, data modeling conventions, medallion architecture (Bronze/Silver/Gold), and pipeline reliability practices</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Implement data quality, observability, and monitoring frameworks to ensure platform health and data trustworthiness</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Optimize pipelines for performance, cost, and reliability at scale using Spark and PySpark</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Evaluate, integrate, and govern new platform tooling and data sources within the Databricks ecosystem</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Contribute to architectural decisions and help drive the long-term data platform roadmap</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Participate in code reviews, technical design discussions, and engineering standards</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Mentor junior engineers and elevate overall platform and data engineering practices</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Document platform architecture, pipeline design, and operational runbooks</span></li> </ul> <p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><strong>What you will have</strong></span></p> <ul> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">6+ years of experience in data engineering, data platform, or ML engineering roles</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Strong proficiency in Python and SQL, with a track record of building production-grade data pipelines using both</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Hands-on Databricks expertise: Delta Lake, Unity Catalog, Databricks Workflows, PySpark, and the Databricks ecosystem broadly</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience building and maintaining ML pipelines in production — feature engineering, training pipelines, experiment tracking, and model deployment</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Familiarity with MLflow or comparable experiment tracking and model registry tools</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience working on cloud data platforms (AWS, Azure, or GCP)</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Strong understanding of data modeling, dimensional design, and analytics-friendly data architecture</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience with batch and incremental/CDC pipeline patterns</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Proficiency with Git, version control, and CI/CD practices for data and ML workflows</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Strong engineering judgment — you think about reliability, maintainability, and cost, not just correctness</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Clear communication and comfort working with both technical and non-technical stakeholders</span></li> </ul> <p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><strong>Nice to have</strong></span></p> <ul> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience with streaming or near real-time pipelines (Kafka, Kinesis, Spark Structured Streaming)</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Familiarity with feature store platforms (Databricks Feature Store, Feast, or Tecton)</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience with LLM pipelines, RAG architectures, or AI/BI tooling (Genie, AI Functions)</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Knowledge of data quality and observability tooling (Great Expectations, Monte Carlo, etc.)</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Exposure to dbt or similar SQL-based transformation frameworks</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Infrastructure-as-code experience (Terraform, Databricks Asset Bundles)</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Experience working in Agile or Scrum environments</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Prior experience mentoring engineers or shaping platform standards</span></li> </ul> <p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><strong>Benefits</strong></span></p> <ul> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Generous time off policies</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Top shelf benefits</span></li> <li style="font-family: arial, helvetica, sans-serif; font-size: 12pt;"><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">Education, wellness and lifestyle support</span></li> </ul> <p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">To protect candidate information and maintain a secure hiring process, all applications must be submitted through our careers portal. Resumes or CVs sent directly via email will not be reviewed or considered.</span></p> <p>&nbsp;</p> <p><span style="font-family: arial, helvetica, sans-serif; font-size: 12pt;">#LI-SD1</span></p> <p>&nbsp;</p>