Job Summary
We are seeking a Senior Databricks Data Engineer with 6+ years of data engineering experience to design and scale our next-generation data platform.
In this role, you will be the technical lead for our Databricks on AWS ecosystem.
You will architect robust, cost-effective pipelines that ingest, process, and serve massive datasets to power our analytical and AI initiatives.
You will work natively with AWS services, enforce strict governance, and ensure our infrastructure is optimized for maximum performance.
Key Responsibilities
Key Responsibilities
Databricks & AWS Data Pipeline Engineering
Architect end-to-end data products using the Medallion Architecture (Bronze, Silver, Gold layers) natively built on Amazon S3 and Delta Lake.
Design and scale production pipelines using PySpark, SQL, and Delta Live Tables (DLT) for both high-throughput batch and real-time streaming data.
Ingest diverse data streams by integrating Databricks with AWS messaging systems like Amazon Kinesis or Managed Streaming for Apache Kafka (MSK).
Optimize Lakehouse storage layouts by leveraging Delta features such as Liquid Clustering, Z-Ordering, and data compaction.
Solid understanding of data testing methodologies, including unit, integration, and negative testing.
Experience with model monitoring, alerting, and pipeline observability.
Strong documentation and communication skills.
AWS Infrastructure, Security & Governance
Implement data governance and strict role-based access control across all AWS-hosted data assets using Unity Catalog.
Manage infrastructure security by configuring security roles and personas
Optimize cloud spend by managing Databricks clusters, selecting appropriate Amazon EC2 instance types (compute vs. memory-optimized), and tracking DBUs.
Integrate security frameworks utilizing AWS Secrets Manager to securely handle pipeline credentials and API keys.
DevOps & Engineering Excellence
Azure DevOps via Git integration, automated unit testing, and custom CI/CD pipelines (e.g., GitHub Actions, Databricks DAB).
Orchestrate data workflows seamlessly using Databricks Workflows or Apache Airflow
Deploy Infrastructure as Code (IaC) using Terraform to provision and scale Databricks workspaces and AWS data resources.
Mentor and guide mid-level engineers, establishing coding standards and running comprehensive peer code reviews.
Required Skills & QualificationsTechnical Expertise
Experience: 6+ years in Data Engineering, with 4+ years of deep focus on Databricks within an AWS environment.
Languages: Expert-level Python (PySpark) and complex, analytical SQL.
Storage & Format: Deep technical mastery of Delta Lake, Amazon S3, and advanced Spark optimization techniques.
AWS Ecosystem: Strong hands-on proficiency with EC2, IAM, S3, Secrets Manager, and VPC networking fundamentals.[basic level]
CI/CD & IaC: Proven experience with Terraform for data platform provisioning and automated deployment tools.
Soft SkillsArchitectural Ownership:
Ability to lead technical design discussions and choose cost-efficient cloud patterns.
Technical Communication: Clear documentation skills to map complex data lineage and infrastructure dependencies.
Preferred Certifications[optional]
Databricks Certified Data Engineer Associate/Professional
Skill Requirements
Key Responsibilities
Databricks & AWS Data Pipeline Engineering
Architect end-to-end data products using the Medallion Architecture (Bronze, Silver, Gold layers) natively built on Amazon S3 and Delta Lake.
Design and scale production pipelines using PySpark, SQL, and Delta Live Tables (DLT) for both high-throughput batch and real-time streaming data.
Ingest diverse data streams by integrating Databricks with AWS messaging systems like Amazon Kinesis or Managed Streaming for Apache Kafka (MSK).
Optimize Lakehouse storage layouts by leveraging Delta features such as Liquid Clustering, Z-Ordering, and data compaction.
Solid understanding of data testing methodologies, including unit, integration, and negative testing.
Experience with model monitoring, alerting, and pipeline observability.
Strong documentation and communication skills.
AWS Infrastructure, Security & Governance
Implement data governance and strict role-based access control across all AWS-hosted data assets using Unity Catalog.
Manage infrastructure security by configuring security roles and personas
Optimize cloud spend by managing Databricks clusters, selecting appropriate Amazon EC2 instance types (compute vs. memory-optimized), and tracking DBUs.
Integrate security frameworks utilizing AWS Secrets Manager to securely handle pipeline credentials and API keys.
DevOps & Engineering Excellence
Azure DevOps via Git integration, automated unit testing, and custom CI/CD pipelines (e.g., GitHub Actions, Databricks DAB).
Orchestrate data workflows seamlessly using Databricks Workflows or Apache Airflow
Deploy Infrastructure as Code (IaC) using Terraform to provision and scale Databricks workspaces and AWS data resources.
Mentor and guide mid-level engineers, establishing coding standards and running comprehensive peer code reviews.
Required Skills & QualificationsTechnical Expertise
Experience: 6+ years in Data Engineering, with 4+ years of deep focus on Databricks within an AWS environment.
Languages: Expert-level Python (PySpark) and complex, analytical SQL.
Storage & Format: Deep technical mastery of Delta Lake, Amazon S3, and advanced Spark optimization techniques.
AWS Ecosystem: Strong hands-on proficiency with EC2, IAM, S3, Secrets Manager, and VPC networking fundamentals.[basic level]
CI/CD & IaC: Proven experience with Terraform for data platform provisioning and automated deployment tools.
Soft SkillsArchitectural Ownership:
Ability to lead technical design discussions and choose cost-efficient cloud patterns.
Technical Communication: Clear documentation skills to map complex data lineage and infrastructure dependencies.
Preferred Certifications[optional]
Databricks Certified Data Engineer Associate/Professional
Other Requirements
Key Responsibilities
Databricks & AWS Data Pipeline Engineering
Architect end-to-end data products using the Medallion Architecture (Bronze, Silver, Gold layers) natively built on Amazon S3 and Delta Lake.
Design and scale production pipelines using PySpark, SQL, and Delta Live Tables (DLT) for both high-throughput batch and real-time streaming data.
Ingest diverse data streams by integrating Databricks with AWS messaging systems like Amazon Kinesis or Managed Streaming for Apache Kafka (MSK).
Optimize Lakehouse storage layouts by leveraging Delta features such as Liquid Clustering, Z-Ordering, and data compaction.
Solid understanding of data testing methodologies, including unit, integration, and negative testing.
Experience with model monitoring, alerting, and pipeline observability.
Strong documentation and communication skills.
AWS Infrastructure, Security & Governance
Implement data governance and strict role-based access control across all AWS-hosted data assets using Unity Catalog.
Manage infrastructure security by configuring security roles and personas
Optimize cloud spend by managing Databricks clusters, selecting appropriate Amazon EC2 instance types (compute vs. memory-optimized), and tracking DBUs.
Integrate security frameworks utilizing AWS Secrets Manager to securely handle pipeline credentials and API keys.
DevOps & Engineering Excellence
Azure DevOps via Git integration, automated unit testing, and custom CI/CD pipelines (e.g., GitHub Actions, Databricks DAB).
Orchestrate data workflows seamlessly using Databricks Workflows or Apache Airflow
Deploy Infrastructure as Code (IaC) using Terraform to provision and scale Databricks workspaces and AWS data resources.
Mentor and guide mid-level engineers, establishing coding standards and running comprehensive peer code reviews.
Required Skills & QualificationsTechnical Expertise
Experience: 6+ years in Data Engineering, with 4+ years of deep focus on Databricks within an AWS environment.
Languages: Expert-level Python (PySpark) and complex, analytical SQL.
Storage & Format: Deep technical mastery of Delta Lake, Amazon S3, and advanced Spark optimization techniques.
AWS Ecosystem: Strong hands-on proficiency with EC2, IAM, S3, Secrets Manager, and VPC networking fundamentals.[basic level]
CI/CD & IaC: Proven experience with Terraform for data platform provisioning and automated deployment tools.
Soft SkillsArchitectural Ownership:
Ability to lead technical design discussions and choose cost-efficient cloud patterns.
Technical Communication: Clear documentation skills to map complex data lineage and infrastructure dependencies.
Preferred Certifications[optional]
Databricks Certified Data Engineer Associate/Professional