Job Summary
The Bento Platform Operations Engineer is responsible for providing Tier-2 operational support, administration, monitoring, troubleshooting, security, and lifecycle management for GIC's Bento Platform built on AWS and Red Hat OpenShift Dedicated (OSD). The role ensures seamless operation of containerized application platforms, supporting approximately 40 business applications running across production and non-production OpenShift clusters.
The engineer will manage platform reliability, performance, security compliance, patch management, incident resolution, and operational governance across AWS cloud infrastructure, OpenShift clusters, CI/CD toolchains, security platforms, databases, and supporting services.
Key Responsibilities
Key Responsibilities Platform Operations & Administration Provide Tier-2 operational support for Bento Platform services. Manage and support: OpenShift Dedicated (OSD) clusters AWS cloud services and infrastructure Platform-as-a-Service (PaaS) components Database services (RDS) Supporting platform services and middleware Ensure platform stability, availability, and operational excellence. OpenShift & Container Platform Management Monitor and administer OpenShift Dedicated clusters across production and non-production environments. Support container orchestration, deployment, scaling, and resource optimization. Analyze cluster health, node capacity, pod utilization, and application performance. Troubleshoot container, networking, storage, and platform-related issues. Support approximately: 1,600 pods in Non-Production 800 pods in Production AWS Cloud Operations Manage operational support across GIC AWS cloud accounts. Monitor AWS resources, utilization, cost optimization, and compliance. Support services including: EC2 VPC Load Balancers IAM Storage Services Security Services Monitoring and Logging Components Incident & Problem Management Perform incident diagnosis, troubleshooting, root cause analysis, and resolution. Meet agreed production SLAs and operational KPIs. Participate in major incident management and service restoration activities. Drive proactive problem management and continuous service improvements. Patch, Upgrade & Lifecycle Management Conduct patch assessments and risk evaluations. Manage upgrades and patching activities across: OpenShift clusters Supporting infrastructure RHEL virtual machines Platform services Coordinate maintenance activities while minimizing business i
Skill Requirements
Security & Compliance Perform security vulnerability assessments and remediation. Support platform security compliance aligned with GIC policies. Conduct security scanning, auditing, and reporting. Support centralized security policy enforcement and access controls. Manage and support: CyberArk Conjur AWS Secrets Manager Platform authentication and authorization controls CI/CD & DevOps Support Support platform DevOps toolchain including: GitHub JFrog Artifactory/Repositories CI/CD Pipelines Control-M Agents Troubleshoot deployment pipeline issues. Support container image lifecycle management and repository administration. Monitoring, Logging & Observability Monitor platform health, availability, capacity, and performance. Support logging and observability solutions including: Fluentd Logging Sidecar Monitoring and Alerting Services Generate operational and capacity reports. Track resource consumption and cost optimization opportunities. Platform Services Administration Support and maintain: Red Hat SSO Squid Proxy Fluentd Logger Sidecar Conjur Sidecar AWS Secrets Manager integrations Control-M Agents Supporting RHEL VMs Documentation & Governance Maintain operational runbooks, SOPs, and knowledge repositories. Update technical documentation based on platform changes. Participate in governance reviews, audits, operational reporting, and service reviews. Required Technical Skills OpenShift & Kubernetes Red Hat OpenShift Dedicated (OSD) Kubernetes Administration Container Platforms Pod, Deployment, Service, Route, Ingress Management OpenShift Networking and Security Cloud Technologies AWS Cloud Services EC2 IAM VPC CloudWatch AWS Security Services Cost Optimization and Governance Linux Administration Red Hat Enterprise Linux (RHEL) System Administration Shell Scripting Performance Analysis and Troubleshooting DevOps & Automation GitHub JFrog Artifactory CI/CD Pipelines Jenkins/GitHub Actions (preferred) Infrastructure Automation Security CyberArk Conjur AWS Secrets Manager Vulnerability Management Security Compliance and Auditing Identity and Access Management Monitoring & Logging Fluentd Prometheus Grafana ELK/OpenSearch (preferred) Log Analysis and Alert Management Preferred Qualifications Bachelor's Degree in Computer Science, Information Technology, Engineering, or related discipl
Other Requirements
- 5–10 years of IT Infrastructure and Platform Operations experience.
- Minimum 3+ years supporting Kubernetes/OpenShift environments.
- Experience managing enterprise AWS environments.
- Experience supporting production-critical container platforms.
Preferred Certifications
- Red Hat OpenShift Administration Certification
- Red Hat Certified Engineer (RHCE)
- AWS Certified Solutions Architect Associate/Professional
- AWS SysOps Administrator
- Certified Kubernetes Administrator (CKA)
- ITIL Foundation Certification
Key Competencies
- Strong troubleshooting and analytical skills
- Incident and problem management expertise
- Automation mindset
- Operational excellence and service reliability focus
- Stakeholder communication and coordination
- Ability to work in 24x7 support and critical incident environments
- Continuous improvement and platform optimization mindset