Job Summary
Role: Data Engineer
Primary: AWS Glue, Kafka, or SNS/SQS, Python / PySpark, Data lake, Cloudwatch, Cloudtrail, SNS/SQS, DB design and SQL
Secondary: AWS IAM, EKS
Location: US - Remote (Seattle, WA preferred)
Citizenship: US Citizen or GC holder
Job Description:
• Develop Services to enable data ingestion from and synchronization with system which exposes required data access mechanisms ensuring near-real-time updates
• Ingest data from multiple sources using the python and any other ETL tools
• Design and implement an event-driven architecture using AWS EventBridge, Kafka, or SNS/SQS for real-time data streaming
• Design, implement, and maintain scalable data pipelines that integrate both on-prem and AWS cloud environments.
• Develop efficient Python scripts and applications using libraries like pandas, NumPy, etc., to handle and process large datasets.
• Work with various NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB) to support high-performance data storage and retrieval.
• Develop and deploy applications in a cloud-native architecture, leveraging modern cloud technologies for scalability and resilience.
• Continuously monitor data workflows and systems, troubleshoot issues, and optimize performance for reliability and scalability
Transition existing pipeline to MSSQL server
• Collaborate with the business application owner on the existing data architecture, including data ingestion, data pipelines, business logic, data consumption patterns, and analytics requirements
• Design and document the target data architecture, pipelines, processing and analytics architecture
• Identify opportunities for optimization and consolidation
• Collaboration with data team on decomposition of business logic and data transformation patterns
Key Responsibilities
Develop Services to enable data ingestion from and synchronization with system which exposes required data access mechanisms ensuring near-real-time updates
• Ingest data from multiple sources using the python and any other ETL tools
• Design and implement an event-driven architecture using AWS EventBridge, Kafka, or SNS/SQS for real-time data streaming
• Design, implement, and maintain scalable data pipelines that integrate both on-prem and AWS cloud environments.
• Develop efficient Python scripts and applications using libraries like pandas, NumPy, etc., to handle and process large datasets.
• Work with various NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB) to support high-performance data storage and retrieval.
• Develop and deploy applications in a cloud-native architecture, leveraging modern cloud technologies for scalability and resilience.
• Continuously monitor data workflows and systems, troubleshoot issues, and optimize performance for reliability and scalability
Transition existing pipeline to MSSQL server
• Collaborate with the business application owner on the existing data architecture, including data ingestion, data pipelines, business logic, data consumption patterns, and analytics requirements
• Design and document the target data architecture, pipelines, processing and analytics architecture
• Identify opportunities for optimization and consolidation
• Collaboration with data team on decomposition of business logic and data transformation patterns
Skill Requirements
Role: Data Engineer
Primary: AWS Glue, Kafka, or SNS/SQS, Python / PySpark, Data lake, Cloudwatch, Cloudtrail, SNS/SQS, DB design and SQL
Secondary: AWS IAM, EKS