Job Summary
Key Responsibilities
2. To create the root cause analysis for critical issues/faults (hands on working).
3. To implement any necessary preventive measures to reduce future defects.
4. To provide technical assistance to the team members in resolving customer issues.
5. To execute continuous improvement activities and to improve the teamâs performance.
Skill Requirements
- Monitoring: AppDynamics, Splunk, Big Panda, Grafana (added)
- Schedulers: Autosys (advanced), Control‑M (optional)
- OS: Good Unix/Linux skills; shell scripting a plus
- Databases: Oracle, MSSQL, MongoDB - Ability to write queries in anyone of the DB
- Cloud / DevOps (Good to Have): AWS basics, CI/CD, Docker, Kubernetes
- Languages: Basic understanding of Java/.NET logs; scripting (Python/Shell) helpful
Other Requirements
Advanced Troubleshooting & RCA
- Handle escalations from L1 and resolve complex application issues.
- Perform detailed Root Cause Analysis (RCA) and drive permanent fixes.
- Work closely with Development, DevOps, Infrastructure, and Release teams.
Batch Management & Monitoring Optimization
- Review and optimize Autosys batch jobs, including failure pattern analysis.
- Configure/tune alerts and thresholds in AppDynamics, Splunk, Big Panda, Grafana, Kibana.
- Conduct log analysis, query analysis, and application health checks.
Incident, Problem & Change Management
- Take ownership of P1/P2 high‑severity incidents.
- Engage in Problem Management to eliminate repeat incidents.
- Validate and implement releases, patches, configuration changes.
Environment & Application Support
- Support DEV / SIT / UAT / PROD environments.
- Deploy patches, troubleshoot integration/API issues, and validate releases.
- Perform post‑deployment checks and environment stability reviews.
Collaboration & Process Improvement
- Work with cross‑functional teams (Development, QA, SMEs).
- Create and maintain SOPs, knowledge articles, and troubleshooting guides.
- Mentor L1 analysts and drive process optimization/automation.