Job Summary
We are looking for an AI Observability Architect to design and lead the transformation of monitoring and reliability operations from reactive to predictive and self-healing. The role involves architecting AI-driven observability platforms, defining ML/DL model strategy, and enabling automated detection, correlation, and remediation across IT systems.
Key Responsibilities
- Architect an end-to-end AI observability platform unifying logs, metrics, traces, and events into a single observability data store.
- Design the strategy for predictive/unsupervised ML & DL models (Isolation Forest, Autoencoders, LSTM) for anomaly detection, degradation prediction, and Root Cause Analysis (RCA).
- Define architecture for risk scoring, feature engineering pipelines, and time-window based telemetry aggregation across infra, network, and process signals.
- Build the framework for automated alerting and proactive remediation, progressing toward AI-driven and AI-led resolution with reduced human intervention.
- Define KPIs/SLOs for critical systems and translate them into monitored signals within the observability platform.
- Integrate observability insights into conversational/agentic interfaces and dashboards for natural-language querying of health, risk, and RCA.
- Establish governance, security (RBAC, Entra ID), and scalability patterns so the platform extends across multiple domains.
- Mentor and guide engineers transitioning into AI/ML-driven observability roles.
Skill Requirements
- 6+ years in observability/SRE/monitoring, with 2+ years architecting AI/ML-driven solutions.
- Strong Python skills; hands-on experience with scikit-learn, TensorFlow/PyTorch (LSTM, Autoencoders, Isolation Forest).
- Deep experience with Azure Log Analytics, KQL, and Databricks/PySpark for large-scale telemetry processing.
- Proven ability to design anomaly detection, event correlation, and automated RCA systems.
- Experience with monitoring tool ecosystems (AppDynamics, SolarWinds, Azure Monitor) and unifying them into a common data model.
- Familiarity with agentic AI concepts, Copilot Studio, or conversational interfaces for operational insights.
- Strong architectural and stakeholder communication skills, with the ability to align technical design with business outcomes (MTTD, MTTR, MTBF).
Other Requirements
- Experience building self-healing/automated remediation systems.
- Exposure to Microsoft Fabric or unified data platforms for observability data consolidation.
- Background mentoring SRE teams transitioning into AI/ML-capable roles.