Job Summary
Senior Observability Engineer:
We are building a next-generation Operational Data Lake (ODL) and Customer Anomaly Detection capability focused on proactively identifying customer-impacting issues before they result in incidents or business disruption. The initiative combines observability, analytics, anomaly detection, alerting, and executive-level dashboarding to provide visibility into customer, regional, and global health indicators across the payment ecosystem. The goal is to move from traditional threshold-based monitoring to baseline-driven, predictive detection and actionable insights. This includes monitoring approval rates, transaction volumes, stand-in activity, response code behaviors, customer experience metrics, and other leading indicators of service degradation.
Role Overview
We are looking for a Senior Engineer (5–7 years experience) with strong hands-on expertise in observability platforms and production monitoring environments. The individual will help design, build, and operationalize customer anomaly detection solutions, dashboards, automated alerting, and observability capabilities across multiple data sources.
Key Responsibilities
· Develop and enhance anomaly detection dashboards using Splunk, Grafana, and related observability platforms.
- Build operational dashboards for executive, engineering, and support audiences.
- Create and optimize Splunk searches, correlation logic, alerts, and reports.
- Identify customer, issuer, acquirer, merchant, and regional anomalies through data analysis and baselining techniques.
- Design meaningful KPIs, SLOs, SLIs, and health indicators.
- Integrate data from multiple operational and observability sources.
- Develop automated alerting and notification mechanisms.
- Partner with SRE, Engineering, Operations, and Business teams to investigate trends and improve detection capabilities.
- Support observability strategy, monitoring standards, and dashboard governance.
- Help evolve the platform toward predictive analytics and AI/ML-driven anomaly detection.
Skill Requirements
· Strong experience in Site Reliability Engineering, Production Engineering, Observability, or Monitoring domains.
- Strong hands-on experience with:
o Splunk Enterprise
- Grafana
- Dashboard creation and visualization
- Alerting and monitoring frameworks
- Strong knowledge of observability concepts:
o Metrics
- Logs
- Traces
- SLI/SLO implementation
- Incident detection and response
- Experience building production-grade dashboards and monitoring solutions.
- Strong data analysis and troubleshooting skills.
- Experience writing complex SPL queries.
- Knowledge of anomaly detection, baselining, and trend analysis.
- Experience working with distributed systems and high-volume transaction environments.
- Strong communication and stakeholder management skills.
Preferred Skills
· Experience with OpenTelemetry and modern observability ecosystems.
- Knowledge of AI/ML-based anomaly detection concepts.
- Experience with data lakes, analytics platforms, or streaming data architectures.
- Experience in financial services, payments, or mission-critical production environments.
- Familiarity with Kubernetes, cloud platforms, and automation frameworks.