SME - AzureMSSQL DBA, Microsoft Azure
India
Job Description
SME - AzureMSSQL DBA, Microsoft Azure
Bengaluru, Karnataka

Job Summary

The SRE L3 Engineer – Windows Based Applications is responsible for ensuring the stability, availability, performance, scalability, and reliability of enterprise production applications built on the Microsoft technology stack, including.NET / ASP.NET, IIS, Windows Server, and Microsoft SQL Server. The role extends beyond advanced production support and focuses on deep technical troubleshooting, permanent problem resolution, reliability engineering, automation, performance optimization, and technical governance. The L3 Engineer acts as a technical escalation point and subject matter expert for complex incidents, recurring problems, performance issues, and platform-level application stability concerns. The role collaborates closely with L1/L2 support teams, development teams, infrastructure teams, database teams, and business stakeholders to drive service restoration, eliminate repeat incidents, improve SLO adherence, and strengthen the overall reliability posture of business-critical applications.

• Strong hands-on experience supporting and troubleshooting.NET / ASP.NET applications in enterprise production environments. ,

• Advanced knowledge of IIS, including application pools, bindings, SSL certificates, configuration, deployment, performance tuning, and troubleshooting.

• Strong working knowledge of Microsoft SQL Server, including querying, performance tuning, troubleshooting, database health monitoring, and capacity considerations. ,

• Strong understanding of Windows Server environments, application deployment, configuration management, log analysis, and debugging.

• Experience with observability and monitoring tools such as Splunk or equivalent platforms. 

Key Responsibilities

Application Operations & Technical Ownership • Own and manage L3-level application operations for Windows-based enterprise applications built on.NET / ASP.NET, hosted on IIS, and integrated with Microsoft SQL Server. • Provide expert-level support for complex production issues impacting availability, latency, performance, errors, capacity, and reliability. • Act as the technical owner for application stability, ensuring production environments remain reliable, resilient, and aligned with agreed service levels. • Review application configuration, deployment patterns, environment dependencies, and operational risks to improve service reliability. 

 Incident & Major Incident Management • Serve as the L3 technical escalation point for incidents escalated from L1/L2 support teams. • Lead technical investigation and resolution of complex P1/P2 incidents, ensuring timely service restoration and effective stakeholder communication. • Perform deep-dive analysis across application logs, IIS logs, Windows event logs, SQL performance indicators, metrics, traces, and monitoring dashboards. • Support on-call and 24x7 major incident response as required for business-critical applications.

Problem Management & Root Cause Elimination • Lead root cause analysis for recurring, complex, or business-critical incidents and define corrective and preventive actions. • Identify systemic reliability gaps and drive permanent fixes, workarounds, automation, configuration changes, or architectural improvement recommendations. • Work with development, infrastructure, and database teams to resolve underlying application, platform, configuration, and performance bottlenecks. • Ensure known errors, recurring issues, and technical risks are captured, tracked, and addressed through structured problem management.

Monitoring, Observability & Reliability Engineering • Design, configure, and enhance monitoring, alerting, dashboards, and observability solutions for application health, user impact, performance, and reliability indicators. • Analyze metrics, logs, traces, and monitoring trends to proactively detect performance degradation, stability risks, capacity constraints, and emerging failure patterns. • Drive adherence to SLOs, SLIs, SLAs, and error budget principles, ensuring reliability objectives are measurable and actionable. , • Recommend improvements to monitoring coverage, alert quality, runbooks, dashboards, and operational readiness. ,

Technical Operations – IIS, Windows Server & SQL Server • Provide expert-level troubleshooting and administration support for IIS-hosted applications, including website/application configuration, application pools, deployments, SSL certificates, performance tuning, and issue resolution. • Diagnose and resolve complex Windows Server and application runtime issues impacting application availability or performance. • Analyze and troubleshoot Microsoft SQL Server issues related to query performance, connectivity, capacity, database health, and application-level database dependencies. • Partner with database teams to drive SQL tuning, database health monitoring, and long-term performance improvement actions.

Automation & Continuous Improvement • Develop and maintain automation scripts and tools using PowerShell, C#, or related technologies to reduce manual effort, operational toil, and repeat incidents. • Identify opportunities to automate health checks, log analysis, deployment validation, incident diagnostics, and routine operational activities. • Support CI/CD, release validation, deployment readiness, rollback planning, and production acceptance activities. • Drive continuous improvement actions to improve scalability, reliability, operational efficiency, and supportability. 

Skill Requirements

Technical Leadership, Governance & Documentation • Provide technical guidance and mentoring to L1/L2 support engineers for troubleshooting, incident handling, monitoring, and operational best practices. , • Maintain and improve runbooks, SOPs, troubleshooting guides, knowledge articles, and technical documentation. • Contribute to service improvement plans, architectural fix recommendations, reliability reviews, and operational governance forums. • Communicate effectively with technical teams, business stakeholders, and leadership during incidents, problem reviews, and service improvement discussions.

Core Technical Skills • Strong hands-on experience supporting and troubleshooting.NET / ASP.NET applications in enterprise production environments. , • Advanced knowledge of IIS, including application pools, bindings, SSL certificates, configuration, deployment, performance tuning, and troubleshooting. • Strong working knowledge of Microsoft SQL Server, including querying, performance tuning, troubleshooting, database health monitoring, and capacity considerations. , • Strong understanding of Windows Server environments, application deployment, configuration management, log analysis, and debugging. • Experience with observability and monitoring tools such as Splunk or equivalent platforms. ________________________________________ SRE / Operations Skills • Strong experience in production support, application operations, and SRE practices, preferably at L3 or senior technical escalation level. • Strong understanding of Incident, Problem, and Change Management processes aligned with ITIL practices. • Practical knowledge of monitoring, observability, performance tuning, capacity management, SLOs, SLAs, and error budgets. • Ability to analyze complex application behavior using logs, metrics, traces, SQL data, IIS configuration, and infrastructure signals. • Strong automation mindset with experience in scripting and operational tooling. ________________________________________ Soft Skills • Strong analytical, diagnostic, and problem-solving capabilities for complex production issues. • Ability to work effectively under pressure during major incidents and high-priority escalations. • Strong communication skills with the ability to explain technical issues clearly to technical and business stakeholders. • Collaborative mindset with the ability to work across application, development, infrastructure, database, and business teams. , • Ability to mentor junior engineers and contribute to knowledge sharing and operational maturity.

 

Other Requirements

• 6–10 years of experience in Application Support, Production Support, Site Reliability Engineering, or Application Operations, with strong exposure to Microsoft technology-based enterprise applications. • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related discipline. • Experience working in enterprise-scale production environments supporting business-critical applications. • Prior experience in an L3, technical lead, senior support engineer, or SRE role is preferred. • Experience supporting applications in regulated, large-scale, or complex multi-team environments will be an added advantage. • Experience with advanced observability platforms such as Splunk, AppDynamics, Azure Monitor, or equivalent tools. • Experience with CI/CD pipelines, deployment automation, and release validation. • Exposure to cloud-hosted Windows/.NET workloads, hybrid environments, or containerized application platforms. • Knowledge of secure application operations, certificate management, access controls, and compliance-driven production support. • Experience contributing to service improvement plans, reliability reviews, and operational excellence initiatives.

Information at a Glance

Why HCLTech?

At HCLTech, you'll supercharge your potential. You'll find your career. And you'll find your spark. All at a place that knows that helping its customers stay on top starts by putting its people first.

HCLTech is a global technology company, home to more than 223,000 people across 60 countries, delivering industry-leading capabilities centered around digital, engineering, cloud and AI, powered by a broad portfolio of technology services and products. We work with clients across all major verticals, providing industry solutions for Financial Services, Manufacturing, Life Sciences and Healthcare, Technology and Services, Telecom and Media, Retail and CPG, and Public Services. Consolidated revenues as of 12 months ending June 2026 totaled $14.8 billion.