Job Summary
We are seeking an experienced ETL Data Engineer to design, develop, and optimize scalable data pipelines that support analytics, reporting, and data-driven decision-making. The ideal candidate will have strong expertise in Python, PySpark, Apache Spark, Apache Airflow, and data serialization/storage formats such as Avro and Parquet.
You will work closely with data architects, data analysts, and business stakeholders to build robust, high-performance ETL solutions capable of processing large-scale datasets in distributed environments.
Key Responsibilities
Key Responsibilities
Data Engineering & ETL Development
• Design, develop, and maintain scalable ETL/ELT pipelines using Python and PySpark.
• Build and optimize distributed data processing applications using Apache Spark.
• Develop data ingestion frameworks for structured and semi-structured data sources.
• Transform, cleanse, validate, and enrich data to meet business requirements.
• Implement data quality checks and monitoring mechanisms across data pipelines.
Workflow Orchestration
• Design and manage workflow orchestration using Apache Airflow.
• Create, schedule, monitor, and troubleshoot DAGs for reliable data processing.
• Ensure pipeline resiliency through automated recovery and alerting mechanisms.
Data Storage & Optimization
• Work with columnar storage formats such as Parquet and Avro.
• Optimize data partitioning, compression, and storage strategies.
• Improve query performance and processing efficiency for large datasets.
Performance & Reliability
• Tune Spark jobs for performance and resource utilization.
• Analyze bottlenecks and optimize distributed processing workloads.
• Ensure scalability, availability, and reliability of data platforms.
Collaboration & Governance
• Collaborate with Data Architects and Business Analysts to translate requirements into technical solutions.
• Follow data governance, security, and compliance standards.
• Document technical designs, workflows, and operational procedures.
Skill Requirements
Technical Skills
• Strong programming experience in Python.
• Hands-on experience with PySpark and Apache Spark.
• Expertise in ETL/ELT design and implementation.
• Experience with Apache Airflow for workflow orchestration.
• Knowledge of Avro and Parquet file formats.
• Strong understanding of distributed data processing concepts.
• Experience working with SQL and relational databases.
• Familiarity with data modeling concepts and best practices.
• Experience with Git and CI/CD practices.
Preferred Skills
• Experience with cloud platforms such as Azure, AWS, or GCP.
• Knowledge of Delta Lake, Iceberg, or Hudi.
• Experience with Kafka or other streaming technologies.
• Familiarity with containerization technologies such as Docker and Kubernetes.
• Understanding of DataOps and MLOps practices.
Other Requirements
Experience
• Bachelor’s or master’s degree in computer science, Information Technology, Engineering, or a related field.
• 4+ years of experience in Data Engineering or ETL development.
• Proven experience developing enterprise-scale data pipelines and data platforms.
Nice-to-Have Certifications
• Databricks Certified Data Engineer
• Azure Data Engineer Associate (DP-203)
• AWS Certified Data Analytics
• Google Professional Data Engineer
Success Criteria
The successful candidate will:
• Deliver reliable and scalable ETL pipelines.
• Improve data processing efficiency and performance.
• Ensure high data quality and operational excellence.
• Enable business teams with timely and trusted data.