Job Summary
Description
In this role, you will:
• Build and operate AI evaluation workflows that measure the quality of LLM outputs across chat,
summarization, recommendations, and agent-based features.
• Implement LLM-as-a-Judge and rubric-based evals to score outputs for correctness, relevance, grounding,
and consistency.
• Instrument LLM and agent workflows to capture traces, prompts and responses, metadata, and user
feedback.
• Support release readiness by running evals before launches and highlighting regressions or quality risks.
• Help define agent-specific evals (task completion, tool correctness, error recovery).
• Partner with AI engineers and AI platform teams to translate product requirements into evals criteria.
• Review eval results and recommend improvements.
• Contribute to system design for observability, retries, and logging.
Key Responsibilities
Minimum Qualifications
• 4+ years of experience in data and AI-related fields such AI engineering, software development, ML
engineering, data science, or QA roles.
• We’re looking for someone with an eagerness and ability to learn new skills and solve dynamic problems
in an encouraging and expansive environment.
• Working across global teams to ensure alignment of product development.
• Strong Python skills.
• Applied knowledge of GenAI and RAG strategies, micro-services, recommendation systems, and context
engineering.
• Familiarity of AI evaluation techniques, such as Golden datasets, LLM-as-judge, or rubric-based scoring.
• Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector
databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
• Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop,
Spark, or Snowflake.
• Experience with CI/CD or release validation workflows.
• Familiarity with telemetry and evaluation frameworks for AI agents.
• Experience working with data science teams on insights generation leveraging LLMs.
• Knowledge of project management, and productivity tools such as Wrike and Miro.
• Strong time management skills with the ability to collaborate across multiple teams.
• Able to balance competing priorities, long-term projects, and ad hoc requirements.
• Ability to work in a fast-paced, dynamic, constantly evolving business environment.
Preferred Qualifications
• Hands-on experience with Langfuse or similar tools for LLMs observability.
• Sound communication skills - expert at messaging domain and technical content, at a level appropriate
for the audience. Strong ability to gain trust with stakeholders and senior leadership.
• Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector development
graphs.
• Other complementary technologies for distributed systems architecture and asynchronous messaging,
agent communication, and catching like RabbitMQ, Redis, and Valkey are preferred.
Skill Requirements
Minimum Qualifications
• 4+ years of experience in data and AI-related fields such AI engineering, software development, ML
engineering, data science, or QA roles.
• We’re looking for someone with an eagerness and ability to learn new skills and solve dynamic problems
in an encouraging and expansive environment.
• Working across global teams to ensure alignment of product development.
• Strong Python skills.
• Applied knowledge of GenAI and RAG strategies, micro-services, recommendation systems, and context
engineering.
• Familiarity of AI evaluation techniques, such as Golden datasets, LLM-as-judge, or rubric-based scoring.
• Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector
databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
• Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop,
Spark, or Snowflake.
• Experience with CI/CD or release validation workflows.
• Familiarity with telemetry and evaluation frameworks for AI agents.
• Experience working with data science teams on insights generation leveraging LLMs.
• Knowledge of project management, and productivity tools such as Wrike and Miro.
• Strong time management skills with the ability to collaborate across multiple teams.
• Able to balance competing priorities, long-term projects, and ad hoc requirements.
• Ability to work in a fast-paced, dynamic, constantly evolving business environment.
Preferred Qualifications
• Hands-on experience with Langfuse or similar tools for LLMs observability.
• Sound communication skills - expert at messaging domain and technical content, at a level appropriate
for the audience. Strong ability to gain trust with stakeholders and senior leadership.
• Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector development
graphs.
• Other complementary technologies for distributed systems architecture and asynchronous messaging,
agent communication, and catching like RabbitMQ, Redis, and Valkey are preferred.