Deepeka Gurunathan

Senior Data Scientist & AI/ML Engineer

9+ years building production ML systems and agentic AI across banking, healthcare, and pharma — from fraud detection to multi-agent oncology pipelines.

View My Work → Get in Touch
Deepeka Gurunathan
ABOUT ME

End-to-end AI builder with a research publication and a production track record.

I build AI systems that ship — from LangGraph multi-agent pipelines and RAG architectures to production fraud detection models scoring 500K+ customers monthly. My work spans healthcare, banking, pharma, and telecom across the clients from US, UK, and Indonesia.

Currently a Data Analyst at the Texas Education Agency and a volunteer data scientist at Living Stones Foundation, working on the LumenIndex rural development index for Latin America.

LangGraph GPT-4o RAG Neo4j XGBoost GCP Python SQL Snowflake Tableau Power BI AWS
Let's Connect →
EXPERTISE

What I bring to the table.

🤖

Agentic AI & LLMs

LangGraph · LangChain · GPT-4o · Claude API · OpenAI Realtime API · RAG · Vector DB · ChromaDB · FAISS · Pinecone

📊

Machine Learning

XGBoost · LightGBM · scikit-learn · TensorFlow · PyTorch · Time Series · NLP · Deep Learning · Reinforcement Learning

🔧

Data Engineering

ETL Pipelines · Apache Kafka · Databricks · Snowflake · AWS S3/SageMaker · GCP · Docker · Kubernetes · CI/CD

🗄️

Databases & Analytics

Python · SQL · Neo4j · PostgreSQL · MongoDB · Redis · Tableau · Power BI · Excel

🚀

MLOps & Deployment

MLflow · FastAPI · Streamlit · Render · GitHub Actions · LangSmith · Prometheus · Docker · Kubernetes

☁️

Cloud Platforms

GCP Compute Engine · AWS (S3, EC2, SageMaker, Glue) · Azure ML Studio · Blob Storage

FEATURED PROJECTS

Things I have built and shipped.

View All on GitHub →
🧬

AI · ONCOLOGY · MULTI-AGENT

Molecular Intel Agent

5-agent precision oncology pipeline using LangGraph, GPT-4o, Neo4j knowledge graph, ChromaDB RAG, and XGBoost risk scoring — deployed end-to-end on GCP with LangSmith observability.

LangGraph GPT-4o Neo4j ChromaDB GCP
Live Demo →
🔍

BANKING · AI COPILOT · RAG

Fraud Investigation Copilot

AI-powered copilot for fraud analysts and risk teams — dual RAG pipelines over policy documents and historical fraud cases, returning structured JSON decisions with rule IDs via FastAPI.

LangChain FastAPI ChromaDB GPT-4o RAG
View on GitHub →
🔄

E-COMMERCE · AGENTIC · LANGGRAPH

Refund Easy

Autonomous AI refund agent using LangGraph and GPT-4o that resolves e-commerce refund requests end-to-end with policy validation, duplicate detection, and full LangSmith trace observability.

LangGraph GPT-4o FastAPI LangSmith Streamlit
Live Demo →
✈️

AVIATION · ML · TIME SERIES

Sequence Spoilage Risk Analysis

Sequence spoilage prediction model for pilot training operations using time series analysis, engineering 55+ features from Snowflake SQL pipelines, achieving 73% AUC with temporal train/test split.

XGBoost Snowflake Time Series SQL Python
View on GitHub →
🧪

HEALTHCARE · COMPUTER VISION · CNN

Kidney Disease Classification — EfficientNet

Transfer learning model for kidney CT scan classification across 4 classes (Normal, Cyst, Stone, Tumor) using EfficientNet-B0. Achieved 97% accuracy and 96% macro F1-score with class-imbalance-aware training.

PyTorch EfficientNet Transfer Learning CNN Medical Imaging
View on GitHub →
📚

EDUCATION · RAG · NLP

RAG FAQ Assistant

End-to-end RAG pipeline using ChromaDB for semantic search, LangChain for prompt orchestration, and OpenAI embeddings for university department FAQ response generation.

ChromaDB LangChain OpenAI RAG Streamlit
Live Demo →
EXPERIENCE

Where I have worked.

Jun 2026 – PresentTEA logo

Data Analyst Intern

Texas Education Agency · Texas, USA

  • Analyzing complex education datasets and developing dashboards to support data-driven decision-making for special populations programs serving students across Texas public schools.
Apr 2025 – May 2026UNT logo

Graduate Research & Teaching Assistant

University of North Texas · Denton, TX

  • Designed AI agents using Monte Carlo Tree Search for automated game testing — improved win rate from ~48% to ~77% through policy and rollout tuning.
  • Built simulation pipelines generating 10K+ game scenarios for deck balance and strategy evaluation; research published in JBAM and adopted by Guildhouse Games.
  • Mentored 100+ students on statistical and ML concepts; contributed to course redesign reducing withdrawal rates by ~35%.
Dec 2020 – May 2022Lloyds logo

Data Scientist

Lloyds Banking Group · India

  • Led GBM-based fraud model under extreme class imbalance (~0.2% fraud rate) — achieved OOT AUC ~0.88, capturing ~40% of confirmed fraud in top 5% risk scores with 3+ vintage validation.
  • Deployed REST API churn scoring on Azure ML, integrating with CRM to trigger retention workflows for high-risk customers identified above optimized probability threshold.
  • Led end-to-end A/B test across 3.7M visits (iOS & Android) — +3% CTR lift and +4% product page visits at 95% confidence.
  • Built NLP similarity and clustering system for transaction data, enabling efficient fraud pattern detection and improved reviewer experience.
Aug 2018 – Dec 2020GSK logo

Analyst, Clinical Data Systems

GlaxoSmithKline · India

  • Built patient stratification models using Random Forest and XGBoost — 87% classification accuracy on clinical trial data with automated preprocessing reducing training latency by 47%.
  • Designed XGBoost regression to quantify biomarker reduction in confirmed responders (MAE ~10%, R² ~0.78), supporting dose-response analysis for clinical team.
  • Engineered 50+ clinical features from disparate sources; built automated monitoring framework tracking PSI, feature drift, and recall degradation monthly.
Aug 2016 – Jul 2018UHG logo

Data Engineer

UnitedHealth Group · India

  • Designed scalable ETL pipelines processing 10M+ healthcare claim records daily on AWS S3 and Talend — increased processing capacity by 30% via optimized partitioning and parallel execution.
  • Implemented CI/CD pipelines using Jenkins and Kubernetes for containerized data workflows — reduced deployment time by 40% and improved release reliability.
  • Engineered EDI-based claims ingestion (837/835) with validation, schema standardization, and deduplication — reduced processing errors by 30%.
Mar 2013 – Jul 2016Telkomsel logo

ETL Developer

Telkomsel · India

  • Built large-scale batch ETL pipelines using IBM InfoSphere DataStage — ingested and processed 5–10M+ telecom CDR records daily for customer usage analysis and churn modeling.
  • Performed performance tuning across distributed ETL jobs (parallel processing, partitioning, join optimization) — reduced data latency by ~25% within 4 months.
  • Developed error-handling and recovery mechanisms with job restartability and validation checkpoints — reduced pipeline failures by ~30%.
EDUCATION

Academic background.

MS · May 2026

Advanced Data Analytics

University of North Texas, Denton, TX · GPA 4.0

BE · May 2012

Computer Science & Engineering

Anna University, Chennai, India

ACHIEVEMENTS

Recognition and research.

🏆

AWARD

First Prize — UNT AI in Action 2025

Wildlife Protection AI project recognized at the University of North Texas annual AI research competition.

📄

PUBLICATION

Peer-reviewed Paper · JBAM 2025

Penn, S., Gurunathan, D. "Improving Product Testing with AI for a Card Game Company." Journal of Behavioral and Applied Management. DOI: 10.13140/RG.2.2.33384.64005

🔬

RESEARCH POSTER

Forecasting Hunger Trends — UNT 2025

Research poster on forecasting hunger trends across the United States, published in UNT Scholarly Works 2025.

🌍

VOLUNTEER

LumenIndex — Living Stones Foundation

Building a rural development index for Latin America using World Bank WDI open data, interactive dashboards, and Tableau visualizations.

GET IN TOUCH

Let's build something together.