Job Summary
Builds the data and knowledge foundation for grounded, high-quality AI agents. Owns data pipelines, embeddings, vector stores, and RAG components on Google Cloud, extracting and normalizing data from Intel's enterprise sources for AI consumption.
Key Responsibilities
- Design and implement data ingestion, transformation, and embedding pipelines.
- Extract and normalize data from Databricks, ServiceNow, Snowflake, OneDrive, Adobe, and PDFs.
- Build RAG pipelines using Vertex AI Search, BigQuery, AlloyDB, and Vector Search.
- Optimize retrieval quality, chunking strategies, and prompt grounding.
- Ensure data privacy, PII controls, and access governance; monitor pipeline health and cost.
Skill Requirements
- Strong data engineering with GenAI/RAG experience.
- Hands-on BigQuery, Vertex AI Search, AlloyDB/Cloud SQL, Dataproc/Dataflow.
- Proficient in Python and SQL; embeddings and semantic search.
- Experience integrating enterprise data sources (Databricks, ServiceNow, Snowflake).
Other Requirements
- Knowledge graphs, data privacy/DLP, IAM.
- Adobe systems and unstructured/PDF data extraction.
- Pipeline orchestration and observability.
- 6+ years data engineering with 2+ years GenAI/RAG.
- Google Cloud Data Engineer certification preferred.
- Offshore (India), aligned to overlap with US stakeholders as needed.