Job Summary
Role Overview
We are looking for a highly skilled DevOps/MLOps/LLMOps Engineer to build, automate, and manage scalable AI/ML and Generative AI platforms. The ideal candidate will have experience operationalizing ML models and LLM-based applications, implementing CI/CD pipelines, monitoring AI systems, and enabling production-grade AI deployments.
Key Responsibilities
- Design and maintain CI/CD pipelines for AI/ML and Generative AI applications.
- Build and manage MLOps and LLMOps platforms for model training, deployment, monitoring, and governance.
- Automate provisioning and deployment using Infrastructure as Code (IaC).
- Implement observability, performance monitoring, and cost optimization for AI workloads.
- Manage containerized workloads using Docker and Kubernetes.
- Support deployment and lifecycle management of LLMs, RAG solutions, and Agentic AI applications.
- Ensure security, compliance, and reliability of AI platforms.
Skill Requirements
- 5-8 years of experience in DevOps, MLOps, or Platform Engineering.
- Strong expertise in Azure, AWS, or GCP.
- Hands-on experience with Docker, Kubernetes, Terraform, GitHub Actions, Azure DevOps, or Jenkins.
- Experience with ML lifecycle tools such as MLflow, Kubeflow, Azure ML, or SageMaker.
- Knowledge of LLMOps, prompt management, model evaluation, vector databases, and RAG architectures.
- Experience monitoring production AI systems using tools such as Prometheus, Grafana, Azure Monitor, or dynatrace
- Strong scripting/programming skills in Python, Bash, or PowerShell.
Preferred Skills
- Experience with Azure AI Foundry, Azure OpenAI, Azure AI Search, and Agentic AI solutions.
- Familiarity with LangChain, LangGraph, Semantic Kernel, or AutoGen.
- Knowledge of Responsible AI, model governance, and AI security practices.
- Azure, AWS, Kubernetes, or Terraform certifications.
Ideal Candidate
A hands-on engineer with proven experience building and operating enterprise-scale DevOps, MLOps, and LLMOps platforms, enabling reliable deployment, monitoring, governance, and optimization of AI and Generative AI solutions in production environments.