Job Summary
We are seeking an experienced Databricks Architect to design, build, and optimize scalable cloud-based data platforms and Lakehouse solutions. The ideal candidate will have strong expertise in Databricks, Apache Spark, cloud platforms (AWS/Azure/GCP), data engineering, and modern analytics architecture.
The role involves leading end-to-end architecture, migration, modernization, governance, performance optimization, and implementation of enterprise data solutions supporting analytics, AI/ML, and business intelligence initiatives.
Key Responsibilities
- Design enterprise-scale Lakehouse architecture using Databricks
- Architect scalable ETL/ELT pipelines using PySpark, Spark SQL, and Databricks Workflows
- Implement Delta Lake, Medallion Architecture (Bronze/Silver/Gold), and Unity Catalog
- Lead migration of legacy data warehouses and ETL platforms to Databricks
- Design batch and real-time streaming data pipelines
- Optimize Spark jobs for performance, scalability, and cost efficiency
- Implement CI/CD pipelines and DevOps automation
- Define data governance, security, lineage, and access control standards
- Collaborate with business stakeholders, data engineers, analysts, and leadership teams
- Provide technical leadership, code reviews, architecture reviews, and best practices
- Support AI/ML enablement using MLflow and Databricks capabilities
- Establish monitoring, observability, and operational support frameworks
Skill Requirements
Technical Skills
- Strong hands-on experience with:
- Databricks
- Apache Spark
- PySpark
- SQL
- Python
- Experience with Delta Lake and Unity Catalog
- Expertise in data modeling and Lakehouse architecture
- Experience with cloud platforms:
- AWS
- Azure
- GCP
- Knowledge of:
- Data Lakes
- Data Warehousing
- Streaming frameworks
- CI/CD pipelines
- Git/DevOps
- Familiarity with orchestration tools:
- Airflow
- Azure Data Factory
- Databricks Workflows
Preferred Skills
- MLflow / AI & ML integration
- Kafka / Structured Streaming
- Snowflake integration
- Infrastructure as Code (Terraform)
- Kubernetes exposure
Data governance and security frameworks
Other Requirements
- 10+ years in Data Engineering / Data Architecture
- 5+ years working with Databricks
- Experience leading enterprise cloud modernization projects
- Strong stakeholder management and client-facing experience
- Experience designing scalable distributed data systems