Job Summary
Responsible for leading technical teams in the implementation, automation, and maintenance of DevOps practices using Python and Kubernetes. The role involves ensuring efficient delivery of software projects, optimizing system performance, and providing technical guidance to team members.
Key Responsibilities
- Infrastructure Management & Deployment: Design, provision, and maintain the underlying infrastructure for LLM-based debugging agents and backend services on highly scalable environments (GCP, AWS). Manage CI/CD pipelines to ensure safe, iterative rollouts to Staging and Production.
- Networking & Authentication: Troubleshoot and configure complex networking pathways and security perimeters. Manage configurations for API Gateways, handle CORS policies, and implement secure authentication protocols (OAuth, SSO, End User Credentials, IAM).
- Observability & Log Analytics: Configure and optimize Cloud Logging, monitoring metrics, and dashboards to ensure platform stability. Assist in optimizing system quotas, API rate limits, and token utilization for generative models.
- Automation & Provisioning: Collaborate with software engineering counterparts to automate environment provisioning and testing (using Go-based frameworks and Python scripting).
- Reliability & Troubleshooting: Act as the infrastructure SME to triage backend connection drops, resolve errors proactively, investigate memory pressure issues, and establish safety guardrails for automated data parsing.
Skill Requirements
- Cloud & Orchestration: Deep expertise in Cloud computing platforms (GCP preferred) and container orchestration (Kubernetes, Helm, etc).
- Networking/Security Expertise: Strong understanding of web security, API management, reverse proxies, and identity/access management (IAM, OAuth2.0).
- Scripting & Automation: Proficiency in scripting and automation using Python, Bash, or Go to maintain infrastructure as code and CI/CD pipelines.
- Infrastructure: Familiarity with different infrastructure deployment patterns and automation.
- Observability: Extensive experience with centralized logging, telemetry, and monitoring tools.
- Problem Solving: Proven ability to dive deep into distributed system logs, debug microservice interactions, and resolve infrastructural bottlenecks under pressure.
- Ability to work independently