Job Summary
- US Citizen, Senior Data Engineer with 11+ years of experience designing, building, and optimizing enterprise-scale data platforms, Lakehouse architecture, and analytics solutions in Azure cloud ecosystems. Proven expertise in developing scalable data pipelines, ETL/ELT frameworks, and cloud-native data solutions using Python, Spark, Azure Databricks, Azure Data Factory, Iceberg, Redshift, Synapse, S3, and ADLS Gen2 to support high-volume data processing, advanced analytics, and business intelligence.
Key Responsibilities
- Design and implement scalable Lakehouse and data platform architectures leveraging Apache Iceberg, Delta Lake, and cloud-native storage solutions to enable ACID transactions, schema evolution, data versioning, and multi-engine analytics.
- Develop and maintain high-performance ETL/ELT pipelines using Python, PySpark, Azure Databricks, Azure Data Factory (ADF), and Spark for batch and real-time data processing.
- Design and optimize data models, data warehouses, and semantic layers to support enterprise reporting, self-service analytics, and AI/ML workloads.
- Build cloud-native solutions using Azure (Databricks, Synapse Analytics, ADLS Gen2, Event Hubs, Azure Functions) to deliver scalable and cost-effective data solutions.
- Implement robust data governance, lineage, quality, and compliance frameworks through automated validation, data contracts, metadata management, and monitoring solutions.
- Optimize large-scale data processing workloads through partitioning, clustering, compaction, caching, and query tuning techniques to improve performance and reduce infrastructure costs.
- Establish and maintain DataOps and DevOps best practices, including CI/CD pipelines, Infrastructure-as-Code (Terraform), automated testing, monitoring, and release automation.
- Develop secure data platforms using IAM, RBAC, KMS/CMK encryption, Private Endpoints, Lake Formation, and Microsoft Purview, ensuring compliance with enterprise security and regulatory requirements.
- Build and support streaming and real-time data solutions using Kafka, Kinesis, Event Hubs, Flink, and Spark Streaming.
- Enable AI/ML and advanced analytics initiatives by integrating data platforms with SageMaker, MLflow, Azure Machine Learning, and MLOps frameworks.
- Collaborate closely with business stakeholders, architects, data scientists, analysts, and engineering teams to translate complex business requirements into scalable, reliable, and high-performing data solutions.
Provide technical leadership, mentoring engineering teams, and establish best practices for architecture, coding standards, performance optimization, and cloud-native data engineering.
Skill Requirements
Programming & Data Engineering:
Python, SQL, PySpark, Scala, ETL/ELT, Data Modeling, Data Warehousing
Azure Technologies:
Azure Databricks, Azure Synapse Analytics, Azure Data Factory (ADF), ADLS Gen2, Event Hubs, Azure Functions, Azure Machine Learning, Microsoft Purview
Lakehouse Technologies:
Apache Iceberg, Delta Lake, Hudi, Data Lake Architecture
DataOps & Automation:
Terraform, GitHub Actions, Azure DevOps, Jenkins, Airflow, dbt, Great Expectations
Analytics & BI:
Power BI, Semantic Models, Data Visualization, Executive Reporting
Governance & Security:
RBAC, IAM, KMS/CMK, Data Lineage, Data Governance, Compliance Frameworks
Other Requirements
- Preferred Qualifications
- Experience with Power BI or other data visualization tools.
- Exposure to cloud certifications (e.g., Azure Cloud Practitioner).
- Knowledge of REST APIs and web frameworks (e.g., Flask).
- Experience with stakeholder engagement and requirements gathering.
- Additional experience in project management or leadership roles is a plus.
- Additional Attributes
- Strong organizational and multitasking abilities.
- Willingness to learn new technologies and adapt to changing requirements.
- Ability to analyze business requirements and design robust data pipelines, data models, and ETL/ELT processes to enable reliable data-driven decision-making.