Job Summary
Analytics Engineer
|
Role |
Analytics Engineer |
|
Experience |
7 - 10 years |
|
Primary Stack |
AWS Lakehouse & Analytics (Iceberg, Athena, Lake Formation, Glue, SageMaker) |
Key Responsibilities
Key Responsibilities
- Design and build scalable data lakehouse architectures on AWS using Apache Iceberg as the open table format.
- Develop and optimize analytics query layers using Amazon Athena, including performance tuning, partitioning, and cost optimization.
- Implement fine-grained data access, governance, and permissions using AWS Lake Formation.
- Build and maintain ETL/ELT pipelines using AWS Glue (Jobs, Crawlers, Catalog) for ingestion, transformation, and curation of data.
- Support ML feature engineering and model-ready data pipelines in collaboration with data science teams using Amazon SageMaker.
- Own analytics/dimensional data modelling (star/snowflake schemas, semantic layers, fact & dimension design) to support BI and reporting needs.
- Apply DBT (where applicable) for transformation workflows, testing, and documentation of analytics models.
- Partner with data engineers, data scientists, BI teams, and client stakeholders to translate business requirements into robust analytics solutions.
- Ensure data quality, lineage, and governance standards are met across the lakehouse environment.
- Contribute to architecture decisions, code reviews, and best practices for the analytics engineering team.
Skill Requirements
Must-Have Technical Skills
- Strong hands-on experience with AWS analytics services: Glue, Athena, Lake Formation, S3.
- Practical experience with Apache Iceberg (or similar open table formats such as Delta Lake/Hudi) in a production lakehouse setup.
- Experience with Amazon SageMaker for supporting ML/analytics workflows and feature pipelines.
- Solid grounding in analytics/dimensional data modelling — star schema, snowflake schema, slowly changing dimensions, semantic modelling.
- Strong SQL skills and experience optimizing queries on large-scale distributed data.
- Experience with Python and/or PySpark for data transformation and pipeline development.
- Understanding of data governance, cataloging, and access control concepts on AWS.
Good to Have
- Hands-on experience with DBT (data build tool) for transformation, testing, and documentation.
- Exposure to CI/CD for data pipelines and infrastructure-as-code (Terraform/CloudFormation).
- Familiarity with BI tools such as QuickSight, Tableau, or Power BI.
- AWS certifications (Data Analytics Specialty, Machine Learning Specialty).