Job Summary
Key Responsibilities
Key Responsibilities:
- Design and implement scalable data architectures using Azure services
- Build and optimize ETL/ELT pipelines using Azure Databricks and Apache Spark
- Architect data lake and data warehouse solutions using Azure Data Lake and Azure Synapse Analytics
- Define data modeling strategies (batch & real-time processing)
- Ensure data security, governance, and compliance standards
- Collaborate with stakeholders to understand business requirements and translate them into technical solutions
- Optimize performance and cost efficiency of data workloads
- Implement CI/CD pipelines for data engineering workflows
- Mentor and guide data engineers and development teams
- Integrate advanced analytics, AI, and ML solutions when required
Skill Requirements
Required Skills & Qualifications:
- Strong experience with Azure Databricks and Spark (PySpark/Scala)
- Hands-on experience with Azure services (ADF, ADLS, Synapse, Event Hub)
- Expertise in big data architecture and distributed systems
- Strong knowledge of SQL, Python, and data engineering concepts
- Experience with data modeling techniques (star schema, dimensional modeling)
- Understanding of real-time streaming (Kafka/Event Hub)
- Knowledge of DevOps and CI/CD practices
Strong Development Area:
Data Engineering Foundations
Batch vs streaming; lakehouse concepts; medallion (Bronze/Silver/Gold); file formats (Parquet/Delta/CSV/JSON/Avro); partitioning & clustering; schema evolution; data governance basics; DevOps/CI-CD for data.
SQL
Joins, subqueries, CTEs, window functions, set operations; aggregation & rollups; MERGE/UPSERT; analytic functions; performance (indexes, partition pruning, statistics); data validation scenarios (dedupe, top-N, SCD keys).
Apache Spark (Core)
Spark architecture (driver/executors), DAG, stages/tasks; RDD vs DataFrame/Dataset; wide vs narrow transformations; shuffle mechanics; caching/persistence; partitioning; broadcast joins; skew handling; checkpointing; job tuning.
PySpark
DataFrame API, Spark SQL; UDF vs pandas UDF; windowing; incremental loads; structured streaming (triggers, watermarks); handling semi-structured data; optimizing with predicates, pushdown, join strategies; error handling; unit testing (pytest + chispa).
Databricks Platform
Workspace basics; clusters (Single Node/All-Purpose/Jobs), cluster policies; DBR/LTS; notebooks & Repos; Jobs & Workflows; Delta Lake & Delta Live Tables; Unity Catalog (catalog/schema/table, permissions, lineage); MLflow basics; secret scopes; DBFS; REST APIs; Databricks Connect.
Delta Lake / Lakehouse Patterns
ACID transactions; schema enforcement/evolution; time travel; OPTIMIZE/ZORDER; VACUUM; CDC patterns (MERGE INTO, change data feed); streaming vs batch Delta; expectations/constraints; table maintenance strategies.
Orchestration & Scheduling
Databricks Workflows; Azure Data Factory/Synapse pipelines; triggers; parameter passing; fail/retry; alerts; integration with AutoSys/Control-M/Jenkins/GitHub Actions; event-driven patterns.
Python (Core for Data)
Core syntax; typing & data structures; file I/O; logging; virtual environments; packaging; testing (pytest); common data libs (pandas, pyarrow); error handling; performance considerations (vectorization, generators).
Data Quality & Testing
Great Expectations/dbx expectations or custom checks; unit/integration tests; reconciliation (row/amount); anomaly detection; contract testing for schemas; data observability (metrics, SLAs, freshness).
Security & Governance
Unity Catalog permissions, row/column-level security; secrets management (Key Vault/Secret scopes); PII handling; audit logs; token management; compliance basics.
Cost & Performance Optimization
Cluster sizing, autoscaling; DBU awareness; spot/preemptible instances; storage formats; caching; efficient joins; job scheduling; monitoring with metrics & Ganglia/Spark UI; cost tagging and chargeback.
Analytical Skills & Problem Solving
Break down data problems; root-cause incidents; propose alternatives; estimate complexity; communicate clearly with stakeholders.
Stake Holder Management
Interaction with Client Stakeholders, Communication.
Other Requirements
2. Microsoft Certified: Azure Data Engineer Associate (Optional But Valuable