Gaurav Gurjar

Senior data engineer & AI data platform architect · Dubai

Reliable data
and governed AI
systems.

From architecture through production. I help teams modernize data platforms, automate high-value workflows, and build policy-aware AI systems that remain reliable in production.

Discuss your data challenge → Download CV ↓

300+production pipelines built and operated

2M+people reached by a public-health application

7+ yearsacross data engineering and AI

300+

production pipelines built and operated across regulated, geospatial, API and PDF sources

2M+

people reached by a public-health risk application with serverless data services

7+ years

across data engineering, data science, and AI, in regulated eDiscovery, insurance, public health, veterinary healthcare, and AI governance

Make complex systems dependable.

Three capabilities, one standard: what gets built stays governable, testable, and operable after handoff.

01 Data platforms

Ingestion to warehouse, with discipline.

Cloud ingestion, dimensional modeling, data quality, incremental processing, and production operations. Tested fact and dimension models, not just pipelines that run.

Python · SQL · AWS · Redshift · dbt · PySpark
02 Governed AI systems

Policy before orchestration.

Knowledge ingestion, policy processing, grounding, routing, provenance, PII controls, and runtime telemetry—including on-device execution, where answers ground to a verified knowledge base and the data never leaves the device. The system's decisions live in one place and its actions in another.

JavaScript · TypeScript · Python · FastAPI · Grounding · Provenance
03 Delivery leadership

Architecture through implementation.

Stakeholder communication, testing, release discipline, and end-to-end delivery ownership—the reusable packages and modules underneath, not only the pipeline on top.

Architecture · Reviews · Release discipline · Delivery ownership

Regulated market-data ingestion platform

Regulated market data300+ production pipelines

Cloud ingestion and transformation behind 300+ production pipelines, spanning regulated-business records, geospatial sources, APIs, and PDF documents.

Runs as the governed ingestion standard behind 300+ production pipelines — auditable, reproducible, and no longer dependent on one-off scripts.

View case study →

Public-health risk data services

Public health2M+ people reached

Statistical components, data pipelines, and serverless backend services behind a public COVID-19 mortality-risk application used by 2M+ people.

Backs a public risk application serving 2M+ people, with serverless services that hold up under real public-load spikes.

View case study →

Insurance analytics data platform

InsuranceDaily and monthly incremental snapshots

Tested warehouse models, compliance datasets, and analytical marts on a fixed daily and monthly cadence, replacing manual extracts behind regulatory filings.

Gives compliance and actuarial teams reconciled, trusted marts on a fixed daily and monthly cadence instead of manual extracts.

View case study →

Multi-location veterinary inventory ETL

Veterinary healthcare46 clinic locations

A modular inventory-adjustment package with typed validation at the API boundary and a weekly schedule, proving location coverage before anything is published.

Same standard as the platforms above: coverage checked at publish time, exceptions handled rather than dropped, releases reproducible run to run.

View all case studies →

Maintained projects

Independent work, shipped and maintained.

Independent work

Maintained projects live outside client engagements. Same standard, public by default.

Setu

A Rust data-activation engine that reacts to PostgreSQL change streams and delivers matching events to webhooks, Slack, or Telegram without polling or middleware.

Rust · PostgreSQL · Webhooks

Rental Market Dynamics Dubai

An automated pipeline for extracting Dubai rent-contract data, transforming it to Parquet, and publishing analysis-ready releases and property-usage reports.

Python · Parquet · Automated releases

Hermes Google Sheets

A maintained Hermes Agent plugin for reading, searching, and updating Google Sheets through focused spreadsheet tools.

Hermes · Google Sheets API

Open-source contributions

PUDL and sportsdataverse-py: test modernization, AssetSpec migration, cache handling, and analysis examples. Reviewed upstream.

PUDL · sportsdataverse-py

Trusted in the work.

Three voices, not mine. Each one names a different thing: technical range, collegiality, delivery under difficulty.

“I have worked with Gaurav for close to a year and he is a very well rounded, skilled, and innovative data scientist. He has helped me in the development of statistical methods, backend server infrastructure, and data science tasks. Gaurav is not only a great developer but also a great communicator. He has always been very prompt, responsive, and completes tasks on time. He goes above and beyond to ensure that the customer requirements and needs are met. I recommend him to anyone seeking expert level data science services.”
Benjamin Harvey, Ph.D. · Founder of AI Squared
“Gaurav is a thoughtful person with a very creative mind. He is intellectually curious and looks for efficient solutions to any problems. I enjoyed my time working with him and appreciate the collegial relationship we developed.”
Ivette Basterrechea · Department of Justice
“I worked with Gaurav on several projects, he has strong technical skills and is also a good team player. He was able to deliver high-quality work and found solutions to difficult problems.”
Le Zhang · Google

Field notes

View all insights →

Technical depth, delivery focus.

7+ years, in three acts. Regulated eDiscovery ingestion at a FedRAMP-authorized legal platform, where every dataset had to be defensible downstream. Public-health data science used by 2M+ people. Then five years of governed cloud data platforms in insurance and regulated markets: 300+ production pipelines, dimensional models, and compliance datasets.

Most recently, on-device AI—policy-aware routing, grounded retrieval, and provenance controls for a browser agent that keeps data on the device.

Dubai-based. UAE Golden Visa holder. Available for remote global consulting engagements.

Core technology: Python · SQL · TypeScript · AWS · Redshift · dbt · Airflow · PySpark · FastAPI · Pydantic

Certified: IBM Data Engineer · IBM Data Warehouse Engineer · dbt Fundamentals · Dagster & dbt · Coursera Data Engineering Professional · UAE AI Camp · Full credentials →

Start a conversation

Have a data challenge that has outgrown quick fixes?

Discuss your data challenge →

15 minutes, no pitch deck. Bring the messy source.

Gaurav Gurjar · Senior data engineering and governed AI consulting, from architecture through production.

© 2026 Gaurav Gurjar · Built on paper, shipped on the web.