Job Summary
Key Responsibilities
- Design, develop, test, deploy, and maintain scalable ETL/ELT pipelines for structured and unstructured data.
- Build batch and real-time processing solutions using BigQuery, Dataflow, Cloud Storage, Pub/Sub, and related GCP services.
- Develop data-processing components and integrations using Java or Python, SQL, REST APIs, Kafka, and containerized services on GKE.
- Implement reusable ingestion and transformation frameworks supporting full, incremental, and event-driven loads.
- Design data models and optimize BigQuery tables through partitioning, clustering, efficient SQL, and cost-aware processing patterns.
- Produce high-level and low-level technical designs, document assumptions, and contribute to architecture reviews.
- Implement data-quality checks, schema validation, reconciliation, audit logging, lineage, monitoring, alerting, and error-handling mechanisms.
- Troubleshoot complex production issues, perform root-cause analysis, and deliver sustainable corrective actions.
- Apply security, privacy, access-control, reliability, and operational-support standards across data solutions.
- Use source control, automated testing, CI/CD, and infrastructure-as-code practices to promote repeatable deployments.
- Participate in code reviews, improve engineering standards, and mentor less-experienced engineers.
- Collaborate with architects, analysts, application teams, platform teams, and business stakeholders to deliver reliable data products.
Required Technical Skills
- Strong hands-on experience with GCP data services, particularly BigQuery, Dataflow, Cloud Storage, and Pub/Sub.
- Proficiency in SQL, including complex transformations, query tuning, data validation, and analytical processing.
- Proficiency in Java or Python for data pipelines, automation, API integration, and production support.
- Good understanding of ETL/ELT, data warehousing, data lakes, dimensional modelling, batch processing, and stream processing.
- Experience with Apache Beam, Kafka, REST APIs, Docker, Kubernetes, or GKE.
- Experience with data formats such as JSON, CSV, Avro, and Parquet.
- Working knowledge of Git, automated testing, CI/CD pipelines, observability, and release management.
- Ability to design solutions for scalability, resilience, security, performance, operability, and cost efficiency.
Preferred Skills
- Experience with Cloud Composer or Apache Airflow, Dataproc or Spark, Dataform or dbt, and metadata-driven pipeline frameworks.
- Exposure to Terraform or another infrastructure-as-code tool.
- Knowledge of Dataplex, data governance, metadata management, lineage, and access-control practices.
- Experience modernizing legacy or on-premises data workloads to GCP.
- Google Cloud Professional Data Engineer certification or an equivalent cloud data certification.
Typical Qualifications
Degree in Computer Science, Software Engineering or a related discipline, or equivalent practical experience.
Experience delivering software solutions in agile teams.
Knowledge of modern development frameworks, tools and engineering practices.
Key Responsibilities
Key Responsibilities
- Design, develop, test, deploy, and maintain scalable ETL/ELT pipelines for structured and unstructured data.
- Build batch and real-time processing solutions using BigQuery, Dataflow, Cloud Storage, Pub/Sub, and related GCP services.
- Develop data-processing components and integrations using Java or Python, SQL, REST APIs, Kafka, and containerized services on GKE.
- Implement reusable ingestion and transformation frameworks supporting full, incremental, and event-driven loads.
- Design data models and optimize BigQuery tables through partitioning, clustering, efficient SQL, and cost-aware processing patterns.
- Produce high-level and low-level technical designs, document assumptions, and contribute to architecture reviews.
- Implement data-quality checks, schema validation, reconciliation, audit logging, lineage, monitoring, alerting, and error-handling mechanisms.
- Troubleshoot complex production issues, perform root-cause analysis, and deliver sustainable corrective actions.
- Apply security, privacy, access-control, reliability, and operational-support standards across data solutions.
- Use source control, automated testing, CI/CD, and infrastructure-as-code practices to promote repeatable deployments.
- Participate in code reviews, improve engineering standards, and mentor less-experienced engineers.
- Collaborate with architects, analysts, application teams, platform teams, and business stakeholders to deliver reliable data products.
Required Technical Skills
- Strong hands-on experience with GCP data services, particularly BigQuery, Dataflow, Cloud Storage, and Pub/Sub.
- Proficiency in SQL, including complex transformations, query tuning, data validation, and analytical processing.
- Proficiency in Java or Python for data pipelines, automation, API integration, and production support.
- Good understanding of ETL/ELT, data warehousing, data lakes, dimensional modelling, batch processing, and stream processing.
- Experience with Apache Beam, Kafka, REST APIs, Docker, Kubernetes, or GKE.
- Experience with data formats such as JSON, CSV, Avro, and Parquet.
- Working knowledge of Git, automated testing, CI/CD pipelines, observability, and release management.
- Ability to design solutions for scalability, resilience, security, performance, operability, and cost efficiency.
Preferred Skills
- Experience with Cloud Composer or Apache Airflow, Dataproc or Spark, Dataform or dbt, and metadata-driven pipeline frameworks.
- Exposure to Terraform or another infrastructure-as-code tool.
- Knowledge of Dataplex, data governance, metadata management, lineage, and access-control practices.
- Experience modernizing legacy or on-premises data workloads to GCP.
- Google Cloud Professional Data Engineer certification or an equivalent cloud data certification.
Typical Qualifications
Degree in Computer Science, Software Engineering or a related discipline, or equivalent practical experience.
Experience delivering software solutions in agile teams.
Knowledge of modern development frameworks, tools and engineering practices.