Job Summary
Key responsibilities
- Provide technical support by handling and consulting on BAU, Incidents/emails/alerts for the respective applications.
- Perform post-mortem, root cause analysis using ITIL standards of Incident Management, Service Request fulfilment, Change Management, Knowledge Management, and Problem Management.
- Manage regional L2 team and vendor teams supporting the application. Ensure the team is up to speed and picks up the support duties.
- Build up technical subject matter expertise on the applications being supported including business flows, application architecture, and hardware configuration.
- Define and track KPIs, SLAs and operational metrics to measure and improve application stability and performance.
- Conduct real time monitoring to ensure application SLAs are achieved and maximum application availability (up time) using an array of monitoring tools.
- Build and maintain effective and productive relationships with the stakeholders in business, development, infrastructure, and third-party systems / data providers & vendors.
- Assist in the process to approve application code releases as well as tasks assigned to support to perform. Keep key stakeholders informed using communication templates.
- Approach support with a proactive attitude, desire to seek root cause, in-depth analysis, and strive to reduce inefficiencies and manual efforts.
- Mentor and guide junior team members, fostering technical upskill and knowledge sharing.
- Provide strategic input into disaster recovery planning, failover strategies and business continuity procedures
- Collaborate and deliver on initiatives and install these initiatives to drive stability in the environment.
- Perform reviews of all open production items with the development team and push for updates and resolutions to outstanding tasks and reoccurring issues.
- Drive service resilience by implementing SRE(site reliability engineering) principles, ensuring proactive monitoring, automation and operational efficiency.
- Ensure regulatory and compliance adherence, managing audits, access reviews, and security controls in line with organizational policies.
- The candidate will have to work in shifts as part of a Rota covering APAC and EMEA & USA hours providing 24x7 support. In the event of major outages or issues we may ask for flexibility to help provide appropriate cover.
Your skills and experience
- 7-12 years of experience in providing hands on IT application support.
- Experience in managing vendor team’s providing 24x7 support.
- Preferred: Team lead role experience, Experience in an investment bank, financial institution.
- Bachelor’s degree from an accredited college or university with a concentration in Computer Science or IT-related discipline (or equivalent work experience/diploma/certification).
- Preferred: ITIL v3 foundation certification or higher.
- Knowledgeable in cloud products like Google Cloud Platform (GCP) and Kubernetes and hybrid applications.
- Strong understanding of ITIL /SRE/ DEVOPS best practices for supporting a production environment.
- Understanding of production KPIs, SLO, SLA and SLI.
- Monitoring Tools: Knowledge of Elastic Search, Control M, Grafana, Geneos, OpenShift, Prometheus, Google Cloud Monitoring, Airflow, Splunk.
- Working Knowledge of creation of Reporting Dashboards and reports for senior management
Key Responsibilities
Key responsibilities
- Provide technical support by handling and consulting on BAU, Incidents/emails/alerts for the respective applications.
- Perform post-mortem, root cause analysis using ITIL standards of Incident Management, Service Request fulfilment, Change Management, Knowledge Management, and Problem Management.
- Manage regional L2 team and vendor teams supporting the application. Ensure the team is up to speed and picks up the support duties.
- Build up technical subject matter expertise on the applications being supported including business flows, application architecture, and hardware configuration.
- Define and track KPIs, SLAs and operational metrics to measure and improve application stability and performance.
- Conduct real time monitoring to ensure application SLAs are achieved and maximum application availability (up time) using an array of monitoring tools.
- Build and maintain effective and productive relationships with the stakeholders in business, development, infrastructure, and third-party systems / data providers & vendors.
- Assist in the process to approve application code releases as well as tasks assigned to support to perform. Keep key stakeholders informed using communication templates.
- Approach support with a proactive attitude, desire to seek root cause, in-depth analysis, and strive to reduce inefficiencies and manual efforts.
- Mentor and guide junior team members, fostering technical upskill and knowledge sharing.
- Provide strategic input into disaster recovery planning, failover strategies and business continuity procedures
- Collaborate and deliver on initiatives and install these initiatives to drive stability in the environment.
- Perform reviews of all open production items with the development team and push for updates and resolutions to outstanding tasks and reoccurring issues.
- Drive service resilience by implementing SRE(site reliability engineering) principles, ensuring proactive monitoring, automation and operational efficiency.
- Ensure regulatory and compliance adherence, managing audits, access reviews, and security controls in line with organizational policies.
- The candidate will have to work in shifts as part of a Rota covering APAC and EMEA & USA hours providing 24x7 support. In the event of major outages or issues we may ask for flexibility to help provide appropriate cover.