Job Summary
Key Responsibilities
2. To conduct comprehensive code reviews, establish and oversee quality assurance processes, performance optimization , implementation of best practices and coding standards to ensure successful delivery of complex projects.
3. To ensure process compliance in the assigned module| and participate in technical discussions/review as a technical consultant for feasibility study (technical alternatives, best packages, supporting architecture best practices, technical risks, breakdown into components, estimations).
4. To collaborate with stakeholders to define project scope, objectives, deliverables and accordingly prepare and submit status reports for minimizing exposure & closure of escalations.
Skill Requirements
5–8+ years data engineering; 2+ years production AI/ML or LLM-era data infrastructure. ▪ Proven experience building production pipelines at scale — batch and streaming, Snowflake,AWS/Azure. ▪ Deep expertise: Python, PySpark, Snowflake, Delta Lake, Kafka, Spark Structured Streaming. ▪ Hands-on with vector stores, embedding pipelines, and retrieval infrastructure in production RAG environments. ▪ Working knowledge of MLOps: MLflow, CI/CD for AI, automated evaluation, and production monitoring. ▪ Strong grounding in data governance, quality frameworks, and compliance aligned engineering.
Technical Skills: Expert-Python, SQL, PySpark, Kafka, Delta Lake, AWS (S3, Glue, Kinesis, EKS, Redshift), Docker, Kubernetes, GitHub Actions, Snowflake Strong- LangChain, LlamaIndex, LLM APIs (OpenAI, Bedrock, Claude, HuggingFace), Pinecone, FAISS, ChromaDB, OpenSearch, MLflow, FastAPI, Neo4j Solid- CI/CD pipelines, CloudWatch, Grafana, data lineage platforms, MCP Familiar- LangGraph, prompt engineering, RLHF dataset prep, LLM fine-tuning workflows