Job Summary
Platform Engineer III
Job Summary
Lead the deployment and solutioning of Generative AI and agentic systems in customer environments by owning end-to-end implementation across integration, testing, and stabilization phases. Collaborate with architects, product, and platform teams to design and deliver scalable, reliable, and production-grade AI solutions, ensuring adherence to standards, governance, and performance expectations.
Key Responsibilities
- Lead end-to-end customer deployments across pilot, stabilization, and production phases
- Design and implement AI solution architectures, including APIs, workflows, integrations, and orchestration layers
- Build and optimize RAG pipelines, agent workflows, and orchestration frameworks for real-world use cases
- Define and execute testing, validation, and evaluation frameworks to ensure solution quality and performance
- Monitor system performance and proactively identify, diagnose, and resolve issues across system layers
- Work across data, APIs, and platform components to ensure scalable and resilient system integration
- Drive best practices for safety, governance, and responsible AI implementation across deployments
- Create reusable solution patterns, accelerators, and documentation to improve delivery efficiency
- Collaborate with cross-functional teams (engineering, platform, support, product) to align solution design and execution
- Lead debugging and troubleshooting efforts across production systems, ensuring timely resolution
- Mentor and guide FDE I, II, and III engineers, contributing to team capability building
Skill Requirements
- Strong expertise in Python and software engineering best practices
- Deep understanding of APIs, microservices, and distributed system integration
- Hands-on experience with cloud platforms (Azure / AWS / GCP) and deployment architectures
- Proven experience working with Generative AI / LLM-based systems in production environments
- Strong problem-solving and system-level debugging skills
- Effective communication skills, including ability to engage with customer and internal stakeholders
Other Requirements
- Experience with OpenAI / Azure OpenAI or similar APIs
- Strong knowledge of RAG architectures, prompt engineering, and evaluation frameworks
- Hands-on experience with agent frameworks (LangChain, LangGraph, AutoGen, etc.)
- Familiarity with vector databases, data pipelines, and large-scale data handling
- Experience with observability, monitoring, and performance optimization tools
- Exposure to cost, performance, and reliability optimization for AI systems