Job Summary
Job Summary : The Machine Learning Engineer designs, builds, and maintains machine learning and statistical models that support the organization\'s clinical and operational needs. The role will be working across the entire modeling lifecycle -- framing requirements, preparing model-ready data, engineering features, developing and validating models in Python and/or R, and keeping those models accurate in production over time. Job Description : Machine Learning Engineer\\\\r\\\\nRole Overview\\\\r\\\\nThe Machine Learning Engineer designs, builds, and maintains machine learning and statistical models that support the organization\\\\\\\'s clinical and operational needs. The role will be working across the entire modeling lifecycle -- framing requirements, preparing model-ready data, engineering features, developing and validating models in Python and/or R, and keeping those models accurate in production over time.\\\\r\\\\nKey Responsibilities\\\\r\\\\n1. Requirement Framing and Approach Selection\\\\r\\\\n• Collaborate with the Data Engineering & Analytics team to translate operational and clinical needs into well-defined machine learning problems (classification, regression, clustering, forecasting, anomaly detection, etc.), as well as the success criteria and ongoing evaluation metrics with stakeholders. \\\\r\\\\n• Select the right model for the problem at hand -- linear / logistic regression, regularized models, tree-based or gradient-boosted models, clustering models, etc. including recognizing if a simpler approach will serve the need better.\\\\r\\\\n2. Data Engineering & Preparation\\\\r\\\\n• Write the SQL needed to source, join, and transform data from enterprise source systems (EHR, Financial, and People Operations platforms) into analysis-ready datasets.\\\\r\\\\n• Work on upstream data engineering tasks (building ETL/ELT pipelines, integrating with semi structured data such as REST APIs, building stored procedures, etc.) to ensure the data feeding the machine learning datasets are reliable and well structured. \\\\r\\\\n• Design well validated, clean training and testing datasets with correct granularity, point-in-time correctness, and prevention of data leakage.\\\\r\\\\n• Conduct exploratory analysis to identify patterns, anomalies, and data quality issues before modeling begins.\\\\r\\\\n3. Feature Engineering\\\\r\\\\n• Design, build and document features from raw and transformed data -- aggregations, ratios, categorical coding, business-driven calculated variables, etc.\\\\r\\\\n• Handle real-world data problems systematically, based on the specific context and data problem at hand -- missing data, outliers, class imbalance, etc. \\\\r\\\\n• Work with subject matter experts to understand and incorporate domain knowledge into features and keep feature logic consistent throughout the process.\\\\r\\\\n\\\\r\\\\n4. Model Development and Training\\\\r\\\\n• Develop models in Python (pandas, NumPy, scikit-learn) and/or R, writing reproducible, well-documented code versioned in Git.\\\\r\\\\n• Apply rigorous protocols -- train/validation/test splits, cross-validation, and time-based splits for temporal or seasonal data.\\\\r\\\\n• Tune parameters appropriately to ensure against overfitting.\\\\r\\\\n5. Statistical Modeling and Analysis\\\\r\\\\n• Apply core statistical techniques -- regression, correlation, hypothesis testing, and significance testing to explain relationships and validate findings.\\\\r\\\\n• Understand and be able to distinguish explanatory features from predictive features, and clearly communicate assumptions, limitations, and instances w
Key Responsibilities
Job Responsibilities : Requirement Framing and Approach Selection • Collaborate with the Data Engineering & Analytics team to translate operational and clinical needs into well-defined machine learning problems (classification, regression, clustering, forecasting, anomaly detection, etc.), as well as the success criteria and ongoing evaluation metrics with stakeholders. • Select the right model for the problem at hand -- linear / logistic regression, regularized models, tree-based or gradient-boosted models, clustering models, etc. including recognizing if a simpler approach will serve the need better. 2. Data Engineering & Preparation • Write the SQL needed to source, join, and transform data from enterprise source systems (EHR, Financial, and People Operations platforms) into analysis-ready datasets. • Work on upstream data engineering tasks (building ETL/ELT pipelines, integrating with semi structured data such as REST APIs, building stored procedures, etc.) to ensure the data feeding the machine learning datasets are reliable and well structured. • Design well validated, clean training and testing datasets with correct granularity, point-in-time correctness, and prevention of data leakage. • Conduct exploratory analysis to identify patterns, anomalies, and data quality issues before modeling begins. 3. Feature Engineering • Design, build and document features from raw and transformed data -- aggregations, ratios, categorical coding, business-driven calculated variables, etc. • Handle real-world data problems systematically, based on the specific context and data problem at hand -- missing data, outliers, class imbalance, etc. • Work with subject matter experts to understand and incorporate domain knowledge into features and keep feature logic consistent throughout the process. 4. Model Development and Training • Develop models in Python (pandas, NumPy, scikit-learn) and/or R, writing reproducible, well-documented code versioned in Git. • Apply rigorous protocols -- train/validation/test splits, cross-validation, and time-based splits for temporal or seasonal data. • Tune parameters appropriately to ensure against overfitting. 5. Statistical Modeling and Analysis • Apply core statistical techniques -- regression, correlation, hypothesis testing, and significance testing to explain relationships and validate findings. • Understand and be able to distinguish explanatory features from predictive features, and clearly communicate assumptions, limitations, and instances where the data cannot support accurate conclusions / predictions. 6. Model Evaluation, Validation, and Testing • Evaluate models with metrics appropriate to the problem -- precision/recall, ROC-AUC, and calibration for classification; RMSE/MAE for regression; custom business metrics where needed. • Test model behavior across meaningful subgroups to identify any gaps or bias. • Validate against testing / holdout data and real-world outcomes before results inform decisions, and document evaluation methodology so results are reviewable. 7. Model Lifecycle and Operationalization • Build models that can be deployed across a large organization and expanded on an ongoing basis when needed. The models need to be well parameterized, documented, and structured. • Maintain reproducible machine learning workflows using Git-based version control, model versioning, and documented training processes to ensure models can be audited, retrained, and enhanced over time. • Develop and support automated machine learning pipelines for model training, validation, deployment, and periodic retraining, reducing manual effort and improving reliability. • Establish monitoring processes for ongoing models to detect and respond to data drift and perform