Job Summary
Job Description – Senior GCP Data Engineer (Healthcare / RCM Analytics)
Job Title
Senior GCP Data Engineer – Data Pipelines & Semantic Layer Enablement
Job Summary
We are seeking a highly skilled Senior GCP Data Engineer to design, develop, and support scalable healthcare analytics solutions on Google Cloud Platform (GCP). The candidate will work closely with Data Analysts, Data Architects, and Business SMEs to support the implementation of the IBM Unified Data Model for Healthcare (UDMH), create robust data pipelines, and enable semantic-layer-driven reporting and analytics for Revenue Cycle Management (RCM) data.
The ideal candidate should have extensive experience with BigQuery, Dataform, Dataflow, Dataproc, Airflow, Python, Java, and SQL, and possess a strong understanding of data warehousing, semantic layer concepts, and healthcare data processing.
Key Responsibilities
Data Engineering & Pipeline Development
- Design, develop, and maintain scalable and high-performance data pipelines on GCP.
- Build ingestion, transformation, and orchestration frameworks using Dataflow, Dataform, Dataproc, Airflow, Python, Java, and SQL.
- Develop ETL/ELT pipelines to load and transform healthcare and RCM data into analytical data models.
- Implement automated and metadata-driven data processing frameworks.
- Optimize data pipelines for performance, scalability, reliability, and cost efficiency.
UDMH & Data Model Enablement
- Support implementation of the IBM Unified Data Model for Healthcare (UDMH).
- Develop transformation logic to map source data into UDMH-compliant structures.
- Collaborate with Data Analysts and Architects to implement source-to-target mappings and business rules.
- Validate transformed datasets and ensure compliance with enterprise data standards.
Semantic Layer Support
- Support creation and maintenance of semantic layer datasets used by reporting and analytics platforms.
- Develop curated data models, business views, dimensions, facts, and derived metrics.
- Collaborate with BI teams to ensure consistent KPI calculations and business definitions.
- Enable performant and reusable data assets for reporting and self-service analytics.
BigQuery & Data Warehouse Engineering
- Design and optimize BigQuery data models and datasets.
- Implement partitioning, clustering, and query optimization strategies.
- Develop reusable SQL frameworks and transformation pipelines.
- Support structured, semi-structured, and unstructured data processing requirements.
Workflow Orchestration & Automation
- Develop and manage workflow orchestration using Apache Airflow (Cloud Composer).
- Implement dependency management, scheduling, monitoring, alerting, and recovery mechanisms.
- Build CI/CD pipelines and automated deployment processes for data engineering assets.
- Ensure operational excellence through monitoring, logging, and troubleshooting practices.
Data Quality & Governance
- Implement data quality validations and reconciliation checks within pipelines.
- Support metadata management, lineage tracking, and data governance initiatives.
- Build automated controls to detect data anomalies and processing failures.
- Maintain documentation for data flows, transformations, and operational procedures.
Stakeholder Collaboration
- Work closely with Data Analysts, Business SMEs, Data Architects, and Product Owners.
- Participate in technical design reviews and architecture discussions.
- Support testing, deployment, and production stabilization activities.
Provide technical mentorship and guidance to junior engineers
Skill Requirements
Required Qualifications
- Bachelor's degree in Computer Science, Information Systems, Engineering, or related field.
- 7+ years of experience in Data Engineering and Data Warehousing.
- 4+ years of hands-on experience with Google Cloud Platform (GCP).
- Strong experience with:
- BigQuery
- Dataform
- Dataflow
- Dataproc
- Cloud Storage
- Cloud Composer (Airflow)
- Advanced SQL development skills.
- Strong programming experience in Python and/or Java.
- Experience building enterprise-scale ETL/ELT data pipelines.
- Strong understanding of dimensional modeling and data warehousing concepts.
- Experience with CI/CD and DevOps practices for data platforms.
Other Requirements
Preferred Qualifications
- Experience working with Healthcare Provider data platforms.
- Understanding of Revenue Cycle Management (RCM) business processes.
- Experience implementing or supporting IBM UDMH.
- Experience supporting semantic layers using Power BI, Looker, Cognos, Tableau, or similar BI platforms.
- Familiarity with HL7, FHIR, EMPI, Claims, Billing, and Patient Financial data domains.
- Experience with data governance, metadata management, and data lineage frameworks.
Technical Skills
GCP Technologies
- BigQuery
- Dataform
- Dataflow
- Dataproc
- Cloud Composer (Airflow)
- Cloud Storage
- IAM
- Monitoring & Logging
Programming
- Python
- Java
- SQL
- Shell Scripting
Data Engineering
- ETL/ELT Development
- Data Warehousing
- Data Modeling
- Data Quality Frameworks
- Metadata-Driven Pipelines
- Performance Optimization
Analytics & Semantic Layer
- Semantic Data Modeling
- Metrics & KPI Enablement
- Dimensional Modeling
- Fact & Dimension Design
- Business View Development
Domain Knowledge
- Healthcare Analytics
- Revenue Cycle Management (RCM)
- IBM UDMH
- Clinical and Financial Data Structures
-
Key Deliverables
- GCP Data Pipelines
- Dataform Transformation Frameworks
- Dataflow and Dataproc Jobs
- Airflow DAGs and Orchestration Workflows
- UDMH Data Integration Components
- Semantic Layer Supporting Datasets
- Data Quality Validation Frameworks
- Technical Design Documentation
- Production Support and Optimization Artifacts
Experience: 7–12+ Years
Role Level: Senior Data Engineer
Industry: Healthcare / Revenue Cycle Management (RCM)
Cloud Platform: Google Cloud Platform (GCP)
Location: Remote / HybridSuccess Criteria
- Successfully delivers scalable and reliable GCP data pipelines.
- Enables accurate and efficient source-to-UDMH transformations.
- Supports semantic layer development for business reporting and analytics.
- Ensures high data quality, performance, and operational reliability.