Sr Administrator (Tools & Automation)
India
Job Description
Sr Administrator (Tools & Automation)
Hyderabad, Telangana

Job Summary

Job Summary : Problem Manager is responsible for identifying, reviewing, and analysing IT-related issues to prevent future occurrences and reduce their impact on business operations. This role involves detailed analysis, strategic planning, and effective communication to ensure that problems are addressed promptly and thoroughly. Problem Manager is responsible for conducting post-incident reviews & retrospectives, analysing trends from multiple incidents and identifying common themes to prioritize post-incident review action items across the organization. This function is essential within the 3P Service Management team, ensuring the overall health of production sites and the performance of core customer-facing journeys. The ideal candidate will exhibit exceptional analytical skills, with the ability to analyse complex issues and data to identify root causes and develop effective solutions. A methodical thinker, understanding the broader organisational impact of problems and aligning resolution efforts with business objectives. The candidate should possess direct experience in IT problem management, incident management, or a related field, with a proven track record of successfully resolving complex IT issues and implementing long-term solutions.

Key Responsibilities

Job Responsibilities : Primary Role & Responsibilities: Conduct post incident reviews and retrospectives: After every major incident, ensure that all incident data and descriptions are accurately recorded. This involves reviewing detailed information about the incident, including the timeline of events, actions taken, and outcomes. Analyze this data to identify any gaps in the response process and areas for improvement. Facilitate discussions with relevant teams to understand the root causes and contributing factors. Compile a comprehensive report with findings and recommendations to prevent future occurrences and enhance the overall incident response strategy. Repair Items: Collaborate with incident managers, development teams, professional services, service delivery managers and product managers to identify service improvement items and track them to closure. This involves detailed tracking of each item from identification through to implementation, ensuring that all necessary steps are completed. Regular updates and communication with stakeholders are crucial to ensure transparency and alignment. Once implemented, validate the effectiveness of the improvements and document the outcomes for future reference and learning. Executive Reports: Perform comprehensive trend analysis on recurring problems and root causes to identify patterns and themes within the incident data. Create detailed reports outlining the analysis, highlighting significant findings and providing actionable recommendations to address these problem themes. Present these reports to management, offering insights and strategic advice to prevent future incidents and improve overall system reliability and efficiency. Ensure that the recommendations are practical, measurable, and aligned with the organization\'s goals and capabilities. Incident Analysis: Perform a thorough analysis of incident, event, and change data to identify and prioritize problem trends. This involves reviewing patterns in the data to understand recurring issues and their root causes. Collaborate with relevant service owners to drive the resolution of problem tickets, ensuring that each problem is effectively addressed and closed. The process includes gathering detailed information, coordinating with various teams, and implementing solutions that enhance the overall stability and reliability of the system. Monitor aging incident tickets and Repair items: Regular review & monitoring of outstanding issues and tracking their progress. Engage with various teams to ensure updates are provided and actions are taken promptly. Weekly follow-up meetings with responsible teams to discuss the status of these items, address any obstacles, and escalate issues as needed. Ensuring these tasks are completed within the agreed timelines. Once the service improvement items are implemented, validate their effectiveness by conducting tests and gathering feedback from stakeholders. This process ensures that the initial problem has been resolved and that similar issues will not recur. Continuous Improvement: Take an active role in continuously refining the Problem Management and Incident Response processes. This involves gathering feedback from team members and stakeholders on the effectiveness of current procedures and identifying areas where enhancements can be made. Develop and implement new strategies and best practices to streamline workflows, reduce the time to resolution, and improve communication and coordination among teams. Regularly review and update documentation to reflect changes in processes and ensure all team members are aware of and trained on the latest procedures. Facilitate workshops and training sessions to share knowledge and promote a culture of continuous improvement within the organization.

Skill Requirements

Skill Requirement : Experience working with ticketing tools: DfM, Jira, ServiceNow, Zendesk, Remedyforce, Salesforce Familiarity with monitoring tools: Azure Monitoring, Zenoss, Pingdom, Zabbix, Grafana, ALA Understanding of application, platform, OS, and infrastructure layers Basic understanding of Azure cloud technology Understanding of Linux systems in support of a SaaS product Understanding of virtualization & containerized platforms: Kubernetes, VMware Understanding of networking technologies: TCP/IP, DNS, Routing, HTTP

Other Requirements

Other Requirement : Basic understanding of any cloud technology (Azure, AWS, GCP) and related concepts (SaaS, PaaS, IaaS, Firewalls, Load Balancing, Network Security, TCP/IP, VPN, BGP, MPLS, Routing, databases, storage technologies) Experience with automation tools or scripting languages (Bash, PowerShell, Python, Power platform) Preferred certifications: ITIL Foundation V3 or ITIL-4, Azure Cloud (AZ-900, AZ-104, PL-900)

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.