Job Summary
We are seeking a highly skilled SQL-focused Data Engineer / Developer with strong expertise in Databricks. The ideal candidate will be responsible for designing, optimizing, and managing data pipelines, with a strong focus on query performance tuning, scalability, and efficient data processing.
Key Responsibilities
- Key Responsibilities
- Develop, optimize, and maintain SQL queries, stored procedures, and data pipelines
- Design scalable and efficient data processing solutions using Databricks
- Work with large-scale structured and semi-structured datasets
- Ensure data quality, integrity, and reliability across systems
- Collaborate with data analysts, engineers, and business stakeholders
- Support data migration, transformation, and integration initiatives
- Performance Tuning & Optimization (Mandatory Focus)
- Optimize SQL queries for performance using:
- Indexing strategies (clustered/non-clustered)
- Query execution plan analysis
- Partitioning and data distribution techniques
- Improve data processing efficiency in Databricks:
- Spark job optimization (partitioning, caching, broadcast joins)
- Delta Lake optimizations (Z-ordering, vacuum, optimize commands)
- Identify and resolve performance bottlenecks in pipelines and queries
- Reduce query runtime and resource utilization
- Additional Skills (Preferred)
- Experience with cloud platforms (Azure preferred)
- Knowledge of ETL tools and pipelines
- Familiarity with Python / PySpark
- Exposure to CI/CD pipelines and version control (Git)
- Understanding of big data frameworks and distributed systems
Skill Requirements
- Secondary Skills (Good to Have)
- 2. Databricks
- Hands-on experience with Azure Databricks / Databricks Lakehouse platform
- Strong knowledge of:
- Apache Spark (SQL & PySpark)
- Delta Lake architecture
- Experience building:
- ETL/ELT data pipelines
- Data transformations using Spark
- Familiarity with:
- Notebooks, clusters, jobs, and workflows
- Understanding of data lake and medallion architecture (Bronze, Silver, Gold layers)
Other Requirements
- Secondary Skills (Good to Have)
- 2. Databricks
- Hands-on experience with Azure Databricks / Databricks Lakehouse platform
- Strong knowledge of:
- Apache Spark (SQL & PySpark)
- Delta Lake architecture
- Experience building:
- ETL/ELT data pipelines
- Data transformations using Spark
- Familiarity with:
- Notebooks, clusters, jobs, and workflows
- Understanding of data lake and medallion architecture (Bronze, Silver, Gold layers)