Job Summary
e are looking for a detail-oriented ETL Tester to validate data pipelines built on a Medallion Architecture (Bronze, Silver, and Gold layers). You will ensure data accuracy, integrity, and consistency as data moves through ingestion, transformation, and curation stages, working closely with data engineers, analysts, and business stakeholders to deliver trustworthy, production-ready data.
Key Responsibilities
Key Responsibilities
-
Design, develop, and execute test cases for ETL/ELT pipelines across Bronze (raw), Silver (cleansed/conformed), and Gold (curated/business-ready) layers.
-
Validate data ingestion from source systems into the Bronze layer for completeness and schema conformity.
-
Test data transformation logic, including cleansing, deduplication, standardization, and enrichment, applied in the Silver layer.
-
Verify business logic, aggregations, and KPIs computed in the Gold layer against source-of-truth data and business requirements.
-
Perform source-to-target data validation and reconciliation across all Medallion layers.
-
Conduct row-count checks, checksum/hash validation, referential integrity checks, and duplicate/null value analysis.
-
Validate incremental loads, CDC (Change Data Capture), and SCD (Slowly Changing Dimension) logic.
-
Test data pipeline jobs orchestrated through tools such as Databricks, Azure Data Factory, Apache Airflow, or Informatica.
-
Write and execute SQL queries to validate large datasets across data lakes and warehouses, including Delta Lake, Snowflake, Synapse, and BigQuery.
-
Identify, document, and track defects using tools such as JIRA, and work with engineering teams to resolve data quality issues.
-
Develop reusable test scripts and frameworks to automate regression testing of ETL pipelines.
-
Validate data lineage, metadata, and audit/logging mechanisms across pipeline stages.
-
Participate in Agile ceremonies, including sprint planning, stand-ups, and retrospectives, and contribute to test strategy documentation.
-
Perform performance and load testing on large-volume data pipelines where required.
Skill Requirements
-
Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent experience.
-
7+ years of experience in ETL/data testing, with hands-on exposure to Medallion Architecture (Bronze/Silver/Gold layers).
-
Strong proficiency in SQL for complex data validation and reconciliation queries.
-
-
Understanding of data warehousing concepts, including fact/dimension tables, star/snowflake schemas, and SCD types.
-
Experience with test management and defect tracking tools such as JIRA, TestRail, or qTest.
-
Knowledge of data quality frameworks and tools, such as Great Expectations, dbt tests, or Deequ, is a plus.
-
Basic scripting skills in Python or PySpark for test automation.
-
Strong analytical skills and attention to detail when working with large, complex datasets.
-
Excellent communication skills to document defects and collaborate with cross-functional teams
-
-
Experience testing pipelines built on Databricks, Apache Spark, and Delta Lake.
-
Familiarity with cloud data platforms: Azure (ADF, Synapse), AWS (Glue, Redshift), or GCP (BigQuery, Dataflow)
Other Requirements
-
Experience with CI/CD pipelines for data testing, such as Azure DevOps, Jenkins, or GitHub Actions.
-
Exposure to data cataloging and lineage tools, including Unity Catalog, Purview, or Collibra.
-
Familiarity with BI tools such as Power BI or Tableau to validate Gold-layer reporting outputs.
-
Prior experience in a regulated industry, including finance, healthcare, or insurance, with data governance requirements.
-
Certification in Databricks, Azure Data Engineering, or a similar discipline