Job Summary
- Work Overview
Responsible for build and 24x7 L2 operation of BMM application(Rakuten’s Internal) and container platform (k8s-based)
-
- Operations & Management
- Monitor, Manage and operate the BMM(Containerized app owned by Rakuten) and take necessary action for seamless running of application
- Drive the issue/cases raised by user to closure utilizing the knowledge base existing
- Technical Skills
- Ability to work at both at infrastructure and application layer of Kubernetes in Prod and Lab environments
- Development & Automation
- Code modifications in python and Ansible environment to deploy in production.
- Discipline on Training and Education
- Creation of documentation and playbooks necessary for knowledge hub pages and trainings.
- Issues / Incident Support in production environment
- Recovery, Root Cause Analysis (RCA)
- Communication with platform and BMM vendors
- Creation of presentation materials
- Operations & Management
- Working Conditions :(Location: India - 3 Resources)
- Location:Bangalore/India、Rakuten Bangalore Office (RBD olympus)
- Time:Shift Work for 24x7 Coverage
*Provide How to Calculate Costs for Shift Work
Key Responsibilities
Experience & Track Record
- Over 5 or 8 years of experience designing, building, and operating on telecom clouds (on-premises private clouds).
- Proven experience in incident management and root cause analysis in production environments.
- Extensive experience with Kubernetes, including deploying, managing, and troubleshooting clusters in production.
- Development expertise in operational efficiency and automation tools (e.g., Ansible, Python).
- Experience in deploying and modifying Helm charts to adjust configurations on demand, complying with Change Management Policies.
Skill Requirements
2. Proficiency In Aws Cloud Formation, Eks, And Linux Environments
3. Solid Knowledge Of Support Processes And Service Level Agreements (Slas)
4. Excellent Communication And Presentation Skills For Stakeholder Engagement
5. Ability To Perform Root Cause And Trend Analysis Effectively
Other Requirements
Technical Skills & Certifications
- Certified Kubernetes Administrator (CKA) or equivalent certification.
- Strong Linux expertise, including experience with performance tuning, troubleshooting, and scripting.
- Strong troubleshooting skills across hardware (servers), OS (Linux), networks, Kubernetes, and cloud platform software.
- Networking fundamentals for cloud and telecom environments.
Documentation & Leadership
- Proven ability to create detailed technical documentation, including requirements, HLD (High-Level Design), and LLD (Low-Level Design) documents.
- Proven management and communication skills to lead operations as a technical leader, including progress and risk management.
Language Proficiency
- English: Business Level
Cloud-Native & DevOps Expertise
- Experience with cloud-native technologies, including Istio, Helm, and Prometheus.
- Understanding of observability practices, including experience with Grafana for visualization and Prometheus for monitoring metrics.
- Familiarity with Continuous Integration/Continuous Deployment (CI/CD) pipelines and tools like Jenkins, GitLab.
- Expertise in hybrid cloud environments, integrating on-premises private clouds with public cloud services.
Operations & Database Knowledge
- Experience with working in an operations environment and following SLA/SLOs.
- Knowledge of working in database technologies and building applications connecting to them.
Development & Scripting
- Strong skills in debugging code in Python and Ansible.
Certifications
- Relevant Linux certifications (e.g., RHCSA, RHCE).