SME - JBoss Application Server, Apache Tomcat
India
Job Description
SME - JBoss Application Server, Apache Tomcat
Bengaluru, Karnataka

Job Summary

The SRE L3 Engineer is responsible for ensuring the stability, availability, performance, scalability, and reliability of Java-based enterprise applications running on WebSphere, JBoss/WildFly, Linux, and Oracle databases. The role extends beyond advanced operational support and focuses on deep technical troubleshooting, root cause elimination, reliability engineering, performance optimization, automation, and technical leadership. The engineer serves as the highest level of operational escalation for complex production issues and acts as the primary technical liaison between application development, middleware, infrastructure, database, and support teams, driving continuous service improvement and operational excellence.

• Strong hands-on experience with Java/J2EE enterprise applications. \\\\r\\\\n• Advanced expertise in: \\\\r\\\\no IBM WebSphere Application Server (WAS)\\\\r\\\\no JBoss/WildFly\\\\r\\\\no Oracle Database\\\\r\\\\no Linux/Unix environments \\\\r\\\\n• Strong knowledge of: \\\\r\\\\no JVM internals and tuning\\\\r\\\\no Thread dumps and heap dump analysis\\\\r\\\\no Application deployment and configuration\\\\r\\\\no Middleware clustering and high availability\\\\r\\\\no Performance troubleshooting

Key Responsibilities

Application Operations & Technical Ownership • Own and manage L3 support and technical operations for Java/J2EE applications hosted on WebSphere and JBoss/WildFly platforms. • Provide expert-level troubleshooting for application, middleware, database, and infrastructure-related issues impacting production services. • Ensure application availability, resilience, performance, and operational stability against agreed service objectives. • Review deployments, configurations, and architectural dependencies to improve service reliability and maintainability. , Incident & Major Incident Management • Act as the L3 technical escalation point for incidents escalated from L1 and L2 support teams. • Lead technical resolution of critical P1/P2 incidents and ensure rapid service restoration. • Perform deep technical diagnosis using application logs, thread dumps, heap dumps, JVM metrics, middleware logs, Oracle performance data, and infrastructure telemetry. • Participate in major incident management and provide technical guidance during production outages. Problem Management & Root Cause Elimination • Lead detailed Root Cause Analysis (RCA) for recurring and business-critical incidents. • Identify and eliminate reliability risks through permanent fixes, automation, configuration improvements, and architecture recommendations. • Drive reduction of recurring incidents, technical debt, and manual operational effort. • Collaborate with development and architecture teams to address systemic application and platform issues. Middleware Operations – WebSphere / JBoss • Provide advanced administration and troubleshooting for IBM WebSphere Application Server (WAS) and JBoss/WildFly environments. , • Perform and govern: o Application deployments (EAR/WAR) o JVM tuning o Thread pool management o Datasource configuration o Cluster administration o Load balancing optimization o Middleware performance tuning • Analyze thread dumps, heap dumps, garbage collection logs, and server diagnostics to identify performance bottlenecks. • Govern middleware operational standards, health checks, and lifecycle management. Oracle Database Operations • Perform advanced troubleshooting of Oracle database-related issues impacting application performance and availability. , • Analyze SQL performance, locking issues, waits, execution plans, and database resource utilization. • Support release activities involving database changes and schema updates. • Collaborate with DBA teams to implement performance improvement and capacity management initiatives. Monitoring, Observability & Reliability Engineering • Design, implement, and optimize monitoring, alerting, dashboards, and observability solutions. • Analyze logs, metrics, traces, JVM behavior, middleware health, and database performance indicators. • Drive adherence to SLI, SLO, SLA, and Error Budget objectives and continuously improve system reliability. , • Identify trends and reliability risks through proactive monitoring and analytics. Automation & Continuous Improvement • Develop and maintain automation solutions using Shell, Python, Ansible, Groovy, or equivalent technologies. • Automate deployments, health checks, diagnostics, recovery procedures, and operational workflows. • Drive reduction of operational toil through automation and self-healing mechanisms. • Support CI/CD pipelines, release validations, and production readiness activities. Technical Leadership & Governance • Mentor and coach L1/L2 engineers on technical troubleshooting, reliability practices, and operational excellence. • Review operational procedures, SOPs, runbooks, and knowledge articles. • Provide technical leadership during service reviews, problem reviews, and reliability improvement initiatives. 

Skill Requirements

• Participate in architecture discussions and recommend platform reliability improvements.

Core Technical Skills • Strong hands-on experience with Java/J2EE enterprise applications. • Advanced expertise in: o IBM WebSphere Application Server (WAS) o JBoss/WildFly o Oracle Database o Linux/Unix environments • Strong knowledge of: o JVM internals and tuning o Thread dumps and heap dump analysis o Application deployment and configuration o Middleware clustering and high availability o Performance troubleshooting SRE & Operations Skills • Strong experience in Production Support, Application Operations, or Site Reliability Engineering (SRE). • Extensive experience with: o Incident Management o Problem Management o Change Management o Major Incident Handling o RCA methodologies • Expertise in monitoring and observability tools such as Splunk, AppDynamics, Dynatrace, Grafana, or similar. • Strong understanding of: o SLI/SLO/SLA concepts o Error Budgets o Capacity Management o Reliability Engineering practices Soft Skills • Strong analytical and problem-solving skills. • Ability to lead technical troubleshooting during major incidents. • Excellent written and verbal communication skills. • Strong stakeholder management and cross-functional collaboration capabilities. • Ability to mentor and guide junior engineers.

Other Requirements

 • 6–10 years of experience in Java Application Support, Middleware Operations, Application Operations, Production Support, or SRE roles. Adapted from the L2 requirement of 3–6 years. • Bachelor\'s degree in Computer Science, Information Technology, Engineering, or related discipline. • Experience supporting large-scale enterprise and distributed environments. • Experience as a senior support engineer, technical lead, middleware specialist, or SRE is preferred. Preferred / Good-to-Have Skills • OpenShift, Kubernetes, Docker, or container platform experience. • CI/CD tools such as Jenkins, GitHub Actions, GitLab, or Azure DevOps. • Cloud platforms (Azure, AWS, GCP). • AppDynamics, Dynatrace, Splunk Observability, Prometheus, Grafana. • Infrastructure as Code and configuration management automation. • Experience in enterprise transformation, reliability engineering, and platform modernization initiatives.

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.