SeniorAdministrator - Social Network Development, Netwok Automation
India
Job Description
SeniorAdministrator - Social Network Development, Netwok Automation
Hyderabad, Telangana

Job Summary

Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor\'s or Master\'s degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.

Key Responsibilities

Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor\'s or Master\'s degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.

Skill Requirements

Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor\'s or Master\'s degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.

Other Requirements

Key Responsibilities Agentic AI Development • Design and develop autonomous and assistive AI agents for NOC and fleet operations. • Build multi-agent workflows leveraging LLMs, RAG, planning, orchestration, and tool-calling frameworks. • Develop AI-driven operational assistants capable of diagnostics, root cause analysis, remediation recommendations, and autonomous execution. • Implement human-in-the-loop governance controls for production AI systems. Platform Engineering and Automation • Develop integrations with ServiceNow, Prometheus/Grafana, SolarWinds, PagerDuty, Ansible Automation Platform, Terraform, and cloud and virtualization platforms. • Build APIs, connectors, and automation workflows enabling agent actions across operational systems. • Create self-healing and auto-remediation capabilities for infrastructure and application incidents. Reliability Engineering • Partner with SRE teams to improve SLI/SLO-based monitoring and operational excellence. • Design event correlation and incident intelligence solutions. • Automate incident triage, escalation, and root cause analysis workflows. • Reduce Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR). AI and Data Engineering • Implement Retrieval Augmented Generation (RAG) solutions using operational knowledge bases, runbooks, and historical incident data. • Build evaluation frameworks for AI agent accuracy, safety, and effectiveness. • Develop telemetry, observability, and performance analytics for agent operations. • Ensure responsible and secure AI deployment practices. Collaboration and Leadership • Work closely with AMD Fleet Operations architects, infrastructure teams, SREs, and platform engineering teams. • Mentor junior developers and guide architecture decisions. • Participate in design reviews, technical roadmaps, and innovation initiatives. • Drive adoption of modern AI engineering and automation best practices. Required Qualifications • Bachelor\'s or Master\'s degree in Computer Science, Engineering, Artificial Intelligence, or a related field. • 8+ years of software engineering experience. • Strong hands-on expertise in Python; Java or Go; REST APIs and microservices; Git, CI/CD, and DevOps practices. • Experience building enterprise-scale automation solutions. • Strong understanding of networking, infrastructure, cloud, and operational support environments. • Experience with LangGraph, LangChain, Semantic Kernel, AutoGen, CrewAI, or equivalent multi-agent frameworks. • Experience integrating AI with enterprise tools and operational platforms. Preferred Qualifications • Experience in NOC, SRE, NetOps, CloudOps, or Infrastructure Operations. • Knowledge of VMware, Kubernetes, Azure, AWS, and Linux systems. • Experience with observability platforms such as Prometheus, Grafana, SolarWinds, Datadog, or Splunk. • Experience with ServiceNow development and integration. • Knowledge of routing, switching, DNS, firewalls, and load balancing. Desired Skills Agentic AI Architecture Generative AI and LLMs Multi-Agent Systems RAG and Knowledge Engineering Observability Engineering Infrastructure Automation Terraform and Ansible SRE Practices Event Correlation and AIOps Incident Management Automation Success Metrics • Reduction in manual operational effort. • Increase in automated incident resolution rates. • Improved NOC productivity and operational efficiency. • Reduction in MTTR and operational escalations. • Increased adoption of agent-driven operational workflows. • Delivery of production-ready autonomous operational capabilities. Program Alignment This position supports the AMD Fleet Operations direction of automation-first operations, SRE-led reliability, observability-first architecture, self-healing remediation, and autonomous NetOps, CloudOps, and virtualization operations agents.

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.