Job Summary
Data Validation & Verification
- Validate transformation logic against business rules and documented specifications
- Perform source-to-target data reconciliation — verifying completeness, accuracy, and consistency
- Identify data anomalies, silent failures, and drift in pipeline outputs
- Build and maintain automated data validation suites that execute as part of pipeline runs
- Conduct periodic data audits beyond automated checks
QA Process Definition & Governance
- Define acceptance criteria for each ETL pipeline and transformation step
- Define Definition of Done (DoD)
- Create and maintain data quality test plans covering functional correctness, edge cases, regression, and performance
- Design test cases for new transformations
- Establish data quality SLAs
- Define entry and exit criteria for pipeline releases
- Maintain a defect taxonomy — categorizing data issues (schema drift, logic errors, source issues, timing issues) for root cause tracking and trend analysis
- Define sign-off workflows
Test Strategy & Frameworks, Documentation & Traceability, Monitoring, Observability & Reporting, Collaboration
Key Responsibilities
Data Validation & Verification
- Validate transformation logic against business rules and documented specifications
- Perform source-to-target data reconciliation — verifying completeness, accuracy, and consistency
- Identify data anomalies, silent failures, and drift in pipeline outputs
- Build and maintain automated data validation suites that execute as part of pipeline runs
- Conduct periodic data audits beyond automated checks
QA Process Definition & Governance
- Define acceptance criteria for each ETL pipeline and transformation step
- Define Definition of Done (DoD)
- Create and maintain data quality test plans covering functional correctness, edge cases, regression, and performance
- Design test cases for new transformations
- Establish data quality SLAs
- Define entry and exit criteria for pipeline releases
- Maintain a defect taxonomy — categorizing data issues (schema drift, logic errors, source issues, timing issues) for root cause tracking and trend analysis
- Define sign-off workflows
Test Strategy & Frameworks, Documentation & Traceability, Monitoring, Observability & Reporting, Collaboration
Skill Requirements
SQL-Advanced — window functions, CTEs, set comparisons, complex joins, data profiling queries
AWS Data Services-Hands-on experience querying and validating data in Amazon Redshift, AWS Lake Formation, Athena, and S3-based data lakes
Python (or equivalent scripting) - Validation scripts, data comparison tools, automation frameworks
ETL/ELT Concepts-Deep understanding of extraction, transformation, and loading patterns, including common failure modes
QA Methodology-Test planning, test case design, acceptance criteria definition, defect lifecycle management
Data Profiling-Statistical profiling, distribution analysis, completeness and uniqueness checks
Validation Frameworks-Hands-on experience with at least one: Great Expectations, dbt tests, Soda Core, or equivalent custom frameworks
Version Control-Git — managing test suites alongside pipeline code
Experience
- 6-10 years of combined experience in data engineering, data QA, or analytics engineering
- Has owned data quality for at least one production system end-to-end (not just contributed)
- Has defined acceptance criteria and quality gates that blocked defective releases
- Has built automated validation suites that caught real production issues
- Comfortable reading and reasoning about pipeline code (transformation logic, orchestration DAGs)
- Experience working with curated/aggregated datasets that serve application UIs
- Familiarity with AWS Glue, Redshift Spectrum, and AWS data pipeline services
Preferred Experience
- Experience with BDD-style data testing (Given/When/Then for data transformations)
- CI/CD integration for data quality — automated gates in deployment pipelines
- Experience defining and tracking data SLAs/SLOs
- Knowledge of regulatory or compliance data requirements
- Performance testing for pipelines — verifying latency and throughput
- Exposure to chaos engineering for data — intentionally injecting bad data to test resilience
- Experience with pipeline orchestration tools (Glue Orchestrator, Step Functions, Airflow)
- Experience with IAM permissions and Lake Formation access controls for data governance