Job Summary
- Databricks (Delta Lake, Unity Catalog)
- Snowflake (Snowpark, Cortex AI)
- Agentic AI & LLMs (LangChain, AutoGen)
- Modern ETL/ELT (dbt, Airflow)
- Azure / AWS / GCP
- Vector DBs (Pinecone, Qdrant)
- PySpark & SQL
- Legacy Architecture Modernizatio
Key Responsibilities
2. Develop and refine complex algorithms for large-scale data processing, feature engineering, and pattern recognition.
3. Conduct deep exploratory data analysis, visualization, and statistical modeling to extract actionable insights.
4. Partner with business and technology leaders to integrate data-driven solutions into enterprise strategies.
5. Implement deep learning, natural language processing, and big data technologies to enhance analytics capabilities.
6. Drive data governance, model interpretability, and ethical AI practices for responsible data science implementation.
7. Optimize, automate, and scale data science pipelines for improved operational efficiency and impact.
8. Mentor junior data scientists, foster a data-driven culture, and stay ahead of emerging trends in AI and analytics.
Skill Requirements
- Modern Data Platforms: Deep hands-on expertise with Databricks (Delta Lake, Delta Live Tables, Unity Catalog) and/or Snowflake (Snowpark, Dynamic Tables, Cortex AI, Stream & Tasks).
- Generative AI & Agentic Workflows: Demonstrated experience building GenAI/LLM solutions, agentic frameworks (e.g., LangChain, LlamaIndex, AutoGen, CrewAI), semantic search, and vector database integrations (e.g., Pinecone, Qdrant, Chroma).
- Large-Scale Data Handling: Proven track record of designing and operating large-scale distributed data processing systems with PySpark, Spark SQL, and parallel computing architectures handling multi-terabyte to petabyte datasets.
- Multi-Cloud Infrastructure: Proficient in cloud-native data architecture across at least two major cloud providers (Azure, AWS, GCP), including cloud storage, IAM, serverless compute, and security patterns. Software & Data Engineering Practices: Advanced proficiency in Python and SQL; solid experience with CI/ CD, dbt, Airflow, Docker, Git, unit testing, and Infrastructure as Code (Terraform).
- Legacy Modernization Experience: Tangible experience refactoring and migrating legacy ETL workflows (e.g., SSIS, Informatica, Teradata, Netezza, legacy Hadoop/HDFS) to modern cloud stack architectures.
Preferred Qualifications
- Experience with MLOps and LLMOps frameworks (e.g., MLflow, LangSmith, Weights & Biases) for model deployment and tracking.
- Familiarity with Data Mesh concepts, data clean rooms, and semantic layer management (e.g., Cube, AtScale).
- Relevant certifications: Databricks Certified Data Engineer Professional, Snowflake SnowPro Core / Advanced, or AWS/Azure Data Engineer certifications.