Job Summary
Job Summary : Job Description • Build and maintain data pipelines that ingest infrastructure metrics, logs, and health check outputs • Design ETL processes to feed the RAG pipeline with structured runbook and operational data • Implement vector embedding pipelines for knowledge base ingestion into pgvector/ChromaDB • Create data quality checks and validation layers before AI reasoning layer ingestion • Integrate data from ServiceNow, DB Run scripts, and observability platforms • Optimize data throughput and storage for real-time weather map updates
Job Responsibilities : • Build and maintain data pipelines that ingest infrastructure metrics, logs, and health check outputs • Design ETL processes to feed the RAG pipeline with structured runbook and operational data • Implement vector embedding pipelines for knowledge base ingestion into pgvector/ChromaDB • Create data quality checks and validation layers before AI reasoning layer ingestion • Integrate data from ServiceNow, DB Run scripts, and observability platforms • Optimize data throughput and storage for real-time weather map update
Key Responsibilities
NA
Skill Requirements
Skill Requirement : Skills Required • Python (intermediate-advanced), SQL, Bash scripting • ETL/ELT frameworks: Apache Airflow, Pandas, Polars • Databases: PostgreSQL, MongoDB, Redis, SQLite • Vector databases: pgvector, ChromaDB, FAISS • Cloud data services: GCP BigQuery, Cloud Storage • Data modeling, schema design, and data quality • JSON/YAML parsing and data transformation
Other Requirements
Support with multi stake holders