Job Summary
Provide operational ownership for business-critical SaaS products, ensuring tenant health, secure access, integration reliability, release readiness and effective vendor engagement. This role operates as highest technical escalation, engineering authority and service reliability leadership for the domain.
Key Responsibilities
Key outcomes
- Stable, secure and recoverable business services with minimized user and operational impact
- Predictable SLA performance, transparent communication and high-quality ticket ownership
- Reduced recurrence through measurable problem actions, automation and knowledge reuse
- Safe releases and changes supported by evidence, risk assessment and rollback readiness
- Strong partnership with business teams, global resolver groups and technology vendors
Core responsibilities
- Lead recovery for critical and ambiguous incidents and make risk-based technical decisions
- Own problem-management strategy, complex RCA and permanent engineering remediation
- Define reference architecture, support standards, non-functional requirements and technical guardrails
- Review high-risk changes, designs, data fixes, integrations, upgrades and resilience plans
- Drive observability, automation, capacity, performance, security and disaster-recovery improvements
- Partner with product owners, enterprise architects, cybersecurity, vendors and global operations leaders
- Build capability through mentoring, technical reviews, communities of practice and succession planning
- Tenant configuration, user provisioning, role and entitlement support
- SSO, SCIM, API, webhook and integration troubleshooting
- Vendor-release impact assessment and regression coordination
- Audit-log review, data integrity checks and service-status correlation
- Licence consumption, feature enablement and environment governance
- Escalate product defects and manage vendor communications through closure
Skill Requirements
Technology and technical skill set
Enterprise SaaS platforms, SaaS administration consoles, SAML/OIDC SSO, SCIM provisioning, REST APIs and webhooks, iPaaS / middleware, Browser and endpoint diagnostics, Audit logs, Data export/import tools, Monitoring and synthetic checks, ITSM and vendor support portals.
Primary Skills
- SaaS Application Support Operations, Microsoft Azure Integration Services (Logic Apps, Service Bus, Azure Functions), REST/SOAP API monitoring, enterprise cloud workflow troubleshooting, SaaS vendor incident handling, production environment problem-solving, SLA/OLA target governance
Secondary Skills
- Microsoft Entra ID (Azure Active Directory), Single Sign-On (SSO) protocols (OAuth 2.0, SAML 2.0, OIDC), Azure Monitor, Application Insights, SQL data querying, Git repositories, Postman API testing tools, cloud platform security posture compliance, multi-vendor environment coordination
Other Requirements
Required knowledge and capabilities
- 8-12+ years of relevant experience, with extensive platform engineering, architecture and production-support leadership
- Strong working knowledge of SaaS, tenant, SSO, SCIM, SAML, OIDC, API, webhook, release management.
- Practical understanding of ITIL-aligned incident, problem, change, request, knowledge and major-incident processes.
- Ability to read logs, correlate events, form testable hypotheses and document a defensible technical conclusion.
- Working knowledge of security controls including least privilege, privileged access, secrets, certificates, vulnerability remediation and audit evidence.
- Clear written and verbal communication with the ability to translate technical findings into business impact and recovery options.
Education and preferred certifications
- Bachelor’s degree in computer science, Information Technology, Engineering or a related discipline, or equivalent demonstrable professional experience.
- ITIL 4 Foundation
- Vendor administrator certification
- Cloud security or identity fundamentals
Behavioural competencies
- Customer and business-impact orientation
- Ownership and calm execution under pressure
- Analytical troubleshooting and evidence-based reasoning
- Collaboration across cultures, time zones and suppliers
- Control discipline, attention to detail and documentation quality
- Continuous learning, automation mindset and constructive challenge
- Technical leadership without dependency on formal authority
- Ability to make and communicate risk-based decisions during critical events
Performance measures
- Critical-incident restoration and business-impact reduction
- Availability, performance and SLO attainment
- Permanent fix delivery and technical-debt reduction
- Architecture/change risk reduction
- Observability and automation maturity
- Team capability and vendor effectiveness
Working model
- May support global services across time zones through rotational shifts and on-call coverage.
- Expected to participate in major-incident bridges, planned releases, disaster-recovery exercises and critical business windows.
- Must follow approved change, security, data-handling, validation and segregation-of-duties controls.