Job Summary
Owns the CI/CD, infrastructure automation, and release engineering for the Intel AI Agent Factory. Establishes and operates the automated build-test-deploy pipeline for AI agents on Google Cloud, ensuring secure, repeatable, and observable deployments across development, staging, and production environments in line with the factory delivery model.
Key Responsibilities
- Design, build, and maintain CI/CD pipelines for agent code, prompts, configs, and RAG assets using Google Cloud Build, Cloud Deploy, and GitHub/Cloud Source Repositories.
- Automate infrastructure provisioning using Terraform / Infrastructure-as-Code across GKE, Cloud Run, and Agent Engine runtimes.
- Implement containerization, artifact registry management, versioning, and rollback strategies for agent deployments.
- Configure and manage IAM, VPC-SC, secrets management, and security controls within the deployment pipeline.
- Set up observability, logging, tracing, and monitoring (Cloud Monitoring, Cloud Logging, Cloud Trace) for pipeline and runtime health.
- Enforce quality and security gates in the pipeline (unit/eval tests, vulnerability scans) before promotion to production.
- Support environment readiness, access provisioning, and deployment cadence aligned to the 2-week sprint model and August MVP timeline.
- Collaborate with FDE, AI Agent Developers, and AgenticOps SME to ensure smooth handoffs from build to deploy to operate.
Skill Requirements
- Strong hands-on DevOps/CI-CD engineering on Google Cloud (Cloud Build, Cloud Deploy, Artifact Registry).
- Infrastructure-as-Code with Terraform; containerization with Docker + Kubernetes (GKE) and Cloud Run.
- Scripting proficiency in Python and/or Bash/Go.
- Experience with Git-based workflows, branch strategies, and release automation.
- Working knowledge of cloud security: IAM, VPC-SC, secrets management.
Other Requirements
- Exposure to AI/ML or agent deployment pipelines (Vertex AI, Agent Engine, ADK).
- Observability tooling (Cloud Monitoring/Logging/Trace, Prometheus, Grafana).
- FinOps / cost-management practices on GCP.
- Multi-cloud (AWS/Azure) DevOps familiarity.
- 4–9 years in DevOps / SRE / Cloud/Platform engineering (Tier 3–4).
- Google Cloud DevOps/Professional certification preferred.
- Offshore (India) with overlap hours for US coordination, or Onsite (USA) as required.