Job Summary
A Senior Data Engineer (ETL) designs, builds, and optimizes scalable data pipelines to transform raw data into analytics-ready datasets. They own the architecture of the extraction, transformation, and loading processes, ensuring data quality, reliability, and optimal performance across enterprise data lakes and warehouses.
Key Responsibilities
- Pipeline Development: Design and build complex, high-throughput ETL/ELT pipelines to ingest structured and unstructured data from diverse sources.
- Data Architecture: Structure scalable data models, data warehouses (e.g., Snowflake, BigQuery, Redshift), and data lakes to support advanced analytics and reporting.
- Performance Tuning: Optimize SQL queries, data indexing, partitioning, and cluster configurations to reduce compute costs and pipeline latency.
- Data Quality & Governance: Implement automated testing, data validation frameworks, and data lineage tracking to guarantee data integrity and compliance.
- CI/CD & Automation: Automate deployment workflows using CI/CD tools (e.g., GitHub Actions, Jenkins) and orchestrate pipelines using scheduling tools (e.g., Airflow, Prefect).
Skill Requirements
- Experience: 7+ years of dedicated data engineering experience, with a proven track record of architecting production-grade ETL pipelines.
- Programming Languages: Advanced proficiency in Python, Scala, or Java, alongside expert-level SQL optimization skills.
- Big Data Frameworks: Hands-on experience with distributed computing tools such as Apache Spark, PySpark, Databricks, or Hadoop.
- Cloud Infrastructure: Deep familiarity with cloud data ecosystems on AWS, Azure, or GCP.
- Orchestration Tools: Strong experience managing complex workflows with Apache Airflow, Dagster, or similar orchestration engines