Job Summary
Key Responsibilities
1. Architect end-to-end AI/ML solutions using Python, TensorFlow, PyTorch, and scikit-learn, ensuring scalable model deployment and integration with enterprise systems.
2. Design and implement distributed data processing workflows with Apache Spark and Kafka to support real-time and batch ML model operations.
3. Develop robust data pipelines and feature engineering processes using pandas, NumPy, and Apache Airflow to optimize model performance and data quality.
4. Oversee the development and validation of machine learning models for NLP, deep learning, and time series forecasting, applying advanced techniques and frameworks such as XGBoost and LightGBM.
5. Define and enforce architectural standards for model storage, versioning, and reproducibility using MySQL, PostgreSQL, and DataBricks.
6. Mentor team members on AI/ML best practices and emerging technologies, ensuring continuous skill enhancement and technical excellence.
7. Collaborate with internal stakeholders to gather requirements and translate business needs into technical specifications for AI/ML solutions.
8. Evaluate and integrate new tools and technologies to maintain solution relevance and meet evolving client requirements.
9. Architect and implement RESTful API integrations to enable seamless communication between AI/ML components and external systems, ensuring scalable, secure, and efficient data exchange across diverse enterprise environments.
MUST HAVE:
• 15+ years of hands-on data engineering and architecture experience, with 3–5+ years building
production AI/ML and LLM-era data infrastructure.
• Proven experience designing enterprise-scale AI data platforms that serve multiple AI
consumers —not just one application or pipeline.
• Deep expertise in lakehouse and data mesh architectures: Databricks, Delta Lake, PySpark,
Kafka, Spark Structured Streaming, cloud-native data services (AWS, Azure).
• Hands-on experience with vector stores, semantic models, knowledge graphs, and retrieval
infrastructure in production environments.
• Working knowledge of LLMOps: model serving pipelines, MLflow, CI/CD for AI, automated
evaluation, and production monitoring.
• Strong background in data governance, security, and compliance in regulated industries
(financial services, payments, cybersecurity, healthcare).
• Experience defining data access controls for AI agents and automated systems — not just
human users
Skill Requirements
1. Expert Proficiency In Ai/Ml Model Development Using Python, Tensorflow, Pytorch, And Scikitlearn.
2. Excellent Knowledge Of Distributed Data Processing With Apache Spark And Kafka.
3. Advanced Skills In Data Engineering, Feature Extraction, And Pipeline Automation Using Pandas, Numpy, And Apache Airflow.
4. Solid Understanding Of Classical Machine Learning, Deep Learning, Nlp, And Time Series Forecasting Techniques.
5. Indepth Experience With Relational Databases Such As Mysql And Postgresql For Data Management And Model Storage.
6. Strong Ability To Architect Scalable Solutions Integrating Multiple Data Sources And Ml Frameworks.
7. Excellent Communication And Mentoring Skills To Guide Technical Teams.
Skills:
• Expert: Python, SQL, PySpark, Kafka, Databricks, Delta Lake, Snowflake,AWS (S3, Glue, EKS,
Bedrock, Kinesis, Redshift), Docker, Kubernetes, Terraform, GitHub Actions.
• Strong: LangChain, LlamaIndex, LLM APIs (OpenAI, AWS Bedrock, Claude, HuggingFace), vector
databases (Pinecone, FAISS, ChromaDB, OpenSearch), knowledge graphs (Neo4j).
Cisco Confidential
• Solid: MLflow, FastAPI, CI/CD pipelines, observability tooling (CloudWatch, Grafana, or
equivalent), data lineage and metadata management platforms.
Other Requirements
1. Recommended: TensorFlow Developer Certificate
2. AWS Certified Machine Learning � Specialty
3. Databricks Certified Data Engineer Professional (optional but valuable)
TECH STACK the candidate would work with:
Databricks · Delta Lake · PySpark · Kafka · Spark Structured Streaming · Apache NiFi · Snowflake ·
AWS (S3, Glue, EKS, Bedrock, Kinesis, Redshift, Lambda) · Azure · Kubernetes · Docker ·
Terraform · GitHub Actions · Jenkins · MLflow · LangChain · LlamaIndex · HuggingFace · OpenAI ·
AWS Bedrock · Claude ·Pinecone · FAISS · ChromaDB · OpenSearch · Neo4j · FastAPI · Python ·
SQL · MCP · LangGraph · Prompt Engineering · MLOps · CI/CD · Grafana / CloudWatch