sagar das✳ENGINEER BY TRADE.
CURIOUS BY DEFAULT.

Résumé

Data & AI Engineer

I build scalable data platforms and AI infrastructure for trusted enterprise workflows.

Python, SQL, Spark, Beam, Kafka, Airflow, Iceberg, BigQuery, Rust, TypeScript, RAG, vector search, and agentic AI systems.

Experience

  1. 00

    Senior Data Engineer

    Quantiphi

    Aug 2026 – Present · Marlborough, MA

    Building scalable batch and streaming data platforms for analytics and AI, with a focus on healthcare data interoperability.

    • Designing and building scalable ETL/ELT pipelines and batch plus real-time processing systems, tuned for performance, scalability, and cost efficiency across large structured and semi-structured datasets.
    • Building and maintaining data integration workflows across heterogeneous source systems, with quality, integrity, and reliability controls spanning the full data lifecycle.
    • Designing data models that support downstream analytics and AI use cases, translating stakeholder requirements into technical design.
    • Supporting healthcare data ingestion and transformation against interoperability standards, with CI/CD, testing, and pipeline monitoring built into delivery.
    • 01

      Data Engineer

      Wells Fargo · via Capgemini America Inc.

      Oct 2025 – Aug 2026 · Charlotte, NC

      Leading GenAI-assisted modernization, lakehouse architecture, and event-driven ingestion for regulated financial data.

      • 900+
        Java modules migrated
      • 35+
        financial datasets
      • 20+ yrs
        reporting context
      • Led a GenAI-assisted modernization initiative using GitHub Copilot to migrate 900+ legacy Java modules into Spark and Apache Beam pipelines in roughly 7 months versus a 2+ year manual rewrite estimate.
      • Architected an Apache Iceberg + BigQuery lakehouse for an on-prem-to-GCP migration POC, onboarding 35+ financial datasets with SAR/CSAR compliance considerations.
      • Built event-driven ingestion on GCS, Airflow, and Dataflow with schema enforcement, DQ checks, curated lakehouse outputs, and a BigQuery semantic layer over 20+ years of financial data.
      • GCP
      • GitHub Copilot
      • Apache Beam
      • Spark
      • Airflow
      • Dataflow
      • BigQuery
      • Apache Iceberg
    • 02

      Data Specialist

      University of Maryland

      Sep 2023 – May 2025 · College Park, MD

      Built high-throughput analytics and AI systems for academic operations and research.

      • 20M/day
        ELMS events → BQ
      • 6h→45m
        119-table CDC
      • 5d→2d
        survey RAG review
      • Engineered a Pub/Sub + Dataflow streaming pipeline ingesting 20M+ daily ELMS events into BigQuery for real-time Superset dashboards tracking 15 KPIs across 230 academic programs.
      • Overhauled 119 Redshift ingestion workflows in Python, SQL, and AWS with CDC and validation, cutting runtime from 6 hours to 45 minutes.
      • Prepared an LLM-powered RAG POC over 100K+ open-ended survey responses using Python, LangChain, and Elasticsearch, reducing qualitative review effort by 60%.
      • BigQuery
      • Pub/Sub
      • Dataflow
      • LangChain
      • Elasticsearch
      • Python
    • 03

      Senior Software Engineer. Data Platform

      Tiger Analytics

      Jul 2021 – Jul 2023 · Chennai, India

      Led delivery of a self-serve data fabric used by Fortune 500 clients.

      • 6
        enterprise clients
      • 75%
        onboarding time cut
      • 85+
        analysts + ML engineers
      • Led a 7-engineer team building a self-serve AWS data platform that enabled data mesh adoption across 6 enterprise clients.
      • Designed an Apache Iceberg lakehouse on S3 with Glue Catalog, schema evolution, ACID transactions, and time-travel queries for reproducible ML training datasets and historical analytics.
      • Developed Airbyte/Airflow batch and streaming ingestion with CDC, encryption, Macie PII detection, and a Spark + Deequ DQ framework with 30+ rules, reducing onboarding time by 75% and blocking 85% of bad data before downstream ML.
      • Shipped FastAPI microservices on Kubernetes exposing ELT orchestration as self-serve REST APIs for 85+ analysts and ML engineers.
      • AWS
      • Spark
      • Airflow
      • Airbyte
      • FastAPI
      • Kubernetes
      • Apache Iceberg
      • Deequ

    Education

    1. Master of Information Management

      University of Maryland. College Park

      Aug 2023 – May 2025 · College Park, MD

      • Graduate assistantship as Data Specialist. Built analytics + AI systems across 230 academic programs.
      • Coursework: distributed systems, retrieval, applied ML, data governance.
    2. B.E. Information Technology

      Panjab University, Chandigarh

      2015 – 2019 · India