Job Summary
Job Description: AI/ML Technical Lead Predictive Observability & Reliability Engineering Role Title AI/ML Technical Lead Experience 5–8 years in AI/ML engineering, with 2–3 years in a technical lead/mentoring capacity Location Bangalore Employment Type Full-time Job Description We are looking for a hands-on AI/ML Technical Lead to drive model development, mentor engineers, and deliver production-grade AI/ML solutions for predictive observability and reliability use cases. The role bridges platform architecture with day-to-day technical execution and team delivery. Key Responsibilities • Lead the design, development, and deployment of ML/DL models for anomaly detection, predictive failure analysis, and Root Cause Analysis (RCA). • Own the end-to-end ML pipeline: data ingestion, feature engineering (rolling windows, lag features, aggregations), model training, scoring, and RCA generation. • Build and tune unsupervised models (Isolation Forest, Autoencoders) and supervised/time-series models (LSTM) for degradation prediction and risk scoring. • Translate business/operational requirements into technical model specifications and success metrics (MTTD, MTTR, MTBF improvements). • Guide and mentor engineers, including upskilling them in Python, data engineering, and model development practices. • Collaborate with architecture stakeholders to ensure models integrate cleanly into the broader observability/data platform. • Drive technical delivery across phased rollouts — from POC/MVP to validation to full production experience — including dashboards, alerting, and explanation summaries. • Establish coding standards, model validation practices, and documentation for model architecture and deployment strategy. Required Skill Set • 5–8 years in AI/ML engineering, with 2–3 years in a technical lead or mentoring capacity. • Strong hands-on Python skills; proven experience with scikit-learn, TensorFlow/PyTorch. • Practical expertise in unsupervised learning (clustering, Isolation Forest, Autoencoders) and deep learning for time-series (LSTM/GRU). • Experience building feature engineering pipelines from raw telemetry/log data (Azure Log Analytics, Databricks, or similar). • Solid understanding of observability concepts — anomaly detection, event correlation, RCA, SLIs/SLOs. • Experience leading or mentoring small technical teams, including reviewing code and guiding model development decisions. • Strong communication skills to bridge between operations stakeholders and technical implementation. Preferred Qualifications • Experience with real-time streaming pipelines and automated alerting/remediation systems. • Familiarity with Microsoft Fabric, Azure Databricks, or unified data platforms. • Exposure to Copilot Studio or agentic AI frameworks for conversational operational insights. • Prior experience upskilling engineers or ops teams into AI/ML-capable roles.
Key Responsibilities
Same as above
Skill Requirements
Same as above
Other Requirements
Same as above