Mid Data Engineer (Barcelona hybrid)

Barcelona·Posted 9mo ago
pythonawsspark
<p><strong>We are:</strong><br>Wizeline, a global AI-native technology solutions provider, develops cutting-edge, <strong>AI-powered</strong> digital products and platforms. We partner with clients to leverage data and AI, accelerating market entry and driving business transformation. As a global community of innovators, we foster a culture of <strong>growth, collaboration, </strong>and <strong>impact.</strong><strong><br></strong><strong><br></strong><strong>With the right people and the right ideas, there’s no limit to what we can achieve</strong></p> <p><strong>Are you a fit?</strong></p> <p>Sounds awesome, right? Now, let’s make sure you’re a good fit for the role:</p> <p><strong>Responsibilities:</strong></p> <p><strong>Existing platform (Databricks)</strong></p> <ul> <li>Keep production pipelines running: ingestion, transformation, and delivery to downstream consumers.</li> <li>Diagnose and resolve pipeline failures and data quality issues, often without documentation to fall back on.</li> <li>Reverse-engineer and document existing transformation logic and business rules — this is the input the migration depends on.</li> <li>Migrate legacy tables from Hive Metastore to Unity Catalog.</li> <li>Maintain Iceberg-enabled table sharing between Databricks and Snowflake.</li> </ul> <p><strong>New development (Snowflake, dbt, Airflow)</strong></p> <ul> <li>Build and test dbt models, including incremental materializations and data tests.</li> <li>Develop and maintain Airflow DAGs for orchestration.</li> <li>Validate that migrated pipelines produce output equivalent to the Databricks versions.</li> <li>Contribute to Snowflake modeling, performance, and cost decisions.</li> </ul> <p><strong>Across both</strong></p> <ul> <li>Work directly with client stakeholders on technical topics, alongside the team lead.</li> </ul> <p><strong>Technical Requirements</strong></p> <p><strong>Databricks</strong></p> <ul> <li><strong>PySpark and SQL</strong> — able to read, debug, and modify existing pipelines. Deep Spark tuning is not required.</li> <li><strong>Delta Lake</strong>: MERGE/upsert patterns, table properties, OPTIMIZE, partitioning.</li> <li><strong>Databricks Workflows</strong>, cluster configuration, job troubleshooting.</li> <li><strong>Unity Catalog</strong>: catalogs, schemas, grants, lineage, and the metastore model.</li> </ul> <p><strong>Snowflake</strong></p> <ul> <li>Warehouses, roles and grants, and the general operating model.</li> <li>Query performance and an awareness of how compute cost behaves.</li> </ul> <p><strong>Dbt</strong></p> <ul> <li>Models, sources, tests, and incremental materializations.</li> <li>Project structure and how dbt fits into a deployment workflow.</li> </ul> <p><strong>Airflow</strong></p> <ul> <li>Writing and maintaining DAGs, operators, scheduling, and dependency management.</li> <li>Understanding retries, backfills, and idempotent task design.</li> </ul> <p><strong>Fundamentals</strong></p> <ul> <li>3+ years operating production data pipelines.</li> <li>Strong <strong>SQL</strong> — window functions, complex joins, reading transformation logic written by someone else.</li> <li><strong>Python</strong> for scripting, automation, and API integration.</li> <li>Incremental loading patterns, idempotency, late-arriving data, reprocessing.</li> <li><strong>AWS</strong>: S3, IAM basics. Basic working knowledge of <strong>Redshift</strong> and its role in the wider architecture.</li> </ul> <p><strong>Ways of working</strong></p> <ul> <li><strong>Fluent English</strong> — client-facing role with stakeholders based abroad.</li> <li>Self-directed. Able to make progress on an unfamiliar codebase without a structured onboarding path, and comfortable asking good questions when context is missing.</li> <li>Clear communicator: can explain a production incident to a non-technical stakeholder and give a realistic ETA.</li> </ul> <p><strong><em>Nice-to-have:</em></strong></p> <ul> <li>Experience with an actual platform migration, not only greenfield work.</li> <li>Open table formats, particularly <strong>Iceberg</strong> and cross-platform sharing.</li> <li>Clickstream or web analytics data (Adobe Analytics, Google Analytics, Segment).</li> <li>Experience taking over an undocumented system and stabilizing it.</li> <li><strong>AI Tooling Proficiency</strong>: Leverage one or more AI tools to optimize and augment day-to-day work, including drafting, analysis, research, or process automation. Provide recommendations on effective AI use and identify opportunities to streamline workflows.&nbsp;</li> </ul> <p><strong>What we offer:</strong></p> <ul> <li>A High-Impact Environment</li> <li>Commitment to Professional Development</li> <li>Flexible and Collaborative Culture</li> <li>Global Opportunities</li> <li>Vibrant Community</li> <li>Total Rewards</li> </ul> <p><em>*Specific benefits are determined by the employment type and location.</em></p> <p>Find out more about our culture&nbsp;<a href="https://www.instagram.com/wizelineglobal/">here</a>.</p>