Job Summary
• 15+ years of hands-on data engineering and architecture experience, with 3–5+ years building production AI/ML and LLM-era data infrastructure. • Proven experience designing enterprise-scale AI data platforms that serve multiple AI consumers —not just one application or pipeline. • Deep expertise in lakehouse and data mesh architectures: Databricks, Delta Lake, PySpark, Kafka, Spark Structured Streaming, cloud-native data services (AWS, Azure). • Hands-on experience with vector stores, semantic models, knowledge graphs, and retrieval infrastructure in production environments. • Working knowledge of LLMOps: model serving pipelines, MLflow, CI/CD for AI, automated evaluation, and production monitoring. • Strong background in data governance, security, and compliance in regulated industries (financial services, payments, cybersecurity, healthcare). • Experience defining data access controls for AI agents and automated systems — not just human users.
Key Responsibilities
1. To architect| design and develop [through team] solution for product / sustenance delivery.
2. To train and develop team so as to ensure that there is an adequate supply of trained manpower in the said technology and delivery risks are mitigated.
3. To ensure knowledge up-gradation and work with new technologies so that the solution is current and meets quality standards and the client requirements.
4. To gather specifications and deliver solutions to the client organization based on understanding of a domain or technology.
Skill Requirements
Technical Skills • Expert: Python, SQL, PySpark, Kafka, Databricks, Delta Lake, Snowflake,AWS (S3, Glue, EKS, Bedrock, Kinesis, Redshift), Docker, Kubernetes, Terraform, GitHub Actions. • Strong: LangChain, LlamaIndex, LLM APIs (OpenAI, AWS Bedrock, Claude, HuggingFace), vector databases (Pinecone, FAISS, ChromaDB, OpenSearch), knowledge graphs (Neo4j). Cisco Confidential • Solid: MLflow, FastAPI, CI/CD pipelines, observability tooling (CloudWatch, Grafana, or equivalent), data lineage and metadata management platforms. TECH STACK YOU'LL WORK WITH Databricks · Delta Lake · PySpark · Kafka · Spark Structured Streaming · Apache NiFi · Snowflake · AWS (S3, Glue, EKS, Bedrock, Kinesis, Redshift, Lambda) · Azure · Kubernetes · Docker · Terraform · GitHub Actions · Jenkins · MLflow · LangChain · LlamaIndex · HuggingFace · OpenAI · AWS Bedrock · Claude ·Pinecone · FAISS · ChromaDB · OpenSearch · Neo4j · FastAPI · Python · SQL · MCP · LangGraph · Prompt Engineering · MLOps · CI/CD · Grafana / CloudWatch