PV
Available for full-time · Anywhere in the US

Data Scientist& AI Engineer

Prabhath Vipparthi

I’ve spent the last four years shipping ML in production — embedding retrieval and classification pipelines at Scale AI, and churn and demand-forecasting models at HCL Tech before that. Wrapping up my MS in Data Science at NJIT this May.

MRecords processed
+Automated tests
~Symbols / day
Projects shipped
scroll
Stack in production
Python
PyTorch
TensorFlow
Scikit-learn
Hugging Face
Transformers
spaCy
PySpark
Apache Spark
Databricks
dbt
Airflow
DuckDB
PostgreSQL
pgvector
FastAPI
MLflow
Docker
Kubernetes
Azure
AWS
vLLM
Gemini
RAG
SHAP
tree-sitter
Streamlit
Power BI
Python
PyTorch
TensorFlow
Scikit-learn
Hugging Face
Transformers
spaCy
PySpark
Apache Spark
Databricks
dbt
Airflow
DuckDB
PostgreSQL
pgvector
FastAPI
MLflow
Docker
Kubernetes
Azure
AWS
vLLM
Gemini
RAG
SHAP
tree-sitter
Streamlit
Power BI
Python
PyTorch
TensorFlow
Scikit-learn
Hugging Face
Transformers
spaCy
PySpark
Apache Spark
Databricks
dbt
Airflow
DuckDB
PostgreSQL
pgvector
FastAPI
MLflow
Docker
Kubernetes
Azure
AWS
vLLM
Gemini
RAG
SHAP
tree-sitter
Streamlit
Power BI
Python
PyTorch
TensorFlow
Scikit-learn
Hugging Face
Transformers
spaCy
PySpark
Apache Spark
Databricks
dbt
Airflow
DuckDB
PostgreSQL
pgvector
FastAPI
MLflow
Docker
Kubernetes
Azure
AWS
vLLM
Gemini
RAG
SHAP
tree-sitter
Streamlit
Power BI
About

Building ML systems
that work in production.

I'm a Data Scientist and AI Engineer with 4 years of experience building production ML systems across enterprise and AI-native environments.

Currently at Scale AI as an Applied AI Engineer — I prepare training datasets, engineer ML features, build classification models, and implement embedding-based retrieval workflows that reduce manual review effort by 25%.

Before Scale AI, I spent over 3 years at HCL Tech building churn prediction and demand forecasting models for enterprise clients — improving high-risk recall by 18% and cutting forecast error by approximately 15%.

I'm completing my MS in Data Science at NJIT (GPA 3.7, May 2026) and actively looking for my next full-time role.

What I do

ML & Modeling

  • Classification & regression (scikit-learn, PyTorch, TensorFlow)
  • Cross-validation, precision/recall tuning, ROC-AUC
  • Feature engineering, SMOTE, SHAP explainability

NLP & AI Systems

  • Transformers, Hugging Face, spaCy, embeddings
  • RAG pipelines with pgvector and similarity search
  • LLM evaluation, LLM-as-a-Judge, vLLM inference

Data Engineering

  • PySpark, Apache Spark, Databricks at scale
  • dbt, Apache Airflow, DuckDB, medallion architecture
  • ETL pipelines, data quality tests, Apache Arrow

MLOps & Deployment

  • MLflow experiment tracking and model registry
  • Docker, Kubernetes, Azure, AWS, Render
  • Model monitoring, data drift detection, CI/CD
How I work

Three things I care about,
learned the hard way.

01
Reliability

Ship what runs.

Production over prototype. A model that hits 0.94 in a notebook but drifts in the first week isn't shipped — it's abandoned. I care about the version that survives real data, real users, and real Mondays.

02
Interpretability

Explain the model, not just its score.

SHAP values, decision logs, audit trails. Stakeholders don't sign off on ROC-AUC — they sign off on "why did this customer get flagged." Every classifier I ship comes with a plain-English answer to that question.

03
Rigor

Test the pipeline, not just the notebook.

ML fails silently. Wrong join, drifted feature, stale schema — the model still returns a number. 700+ automated tests across my projects because catching it at 3am with a paged alert is worse than catching it in CI.

Selected Work · 04

Featured projects.

01 / 04·NLP / AI Systems

AI Digital Badge Classification — NJIT LDI Capstone

Production AI-assisted classification system automating NJIT's institutional taxonomy across 3 dimensions, validated on 20 real-world held-out submissions with 100% accuracy. Built a 4-layer text extraction pipeline — 130+ lexicon phrase patterns, 44 regex rules, spaCy Bloom's taxonomy matcher, and an LLM stub — handling OBv3 JSON, guided forms, and free-text inputs. Implemented a deterministic 3-stage rule engine (19 rules) with immutable audit logs, plain-English decision explanations, and human-in-the-loop override workflows.

100%
Accuracy
351 passing
Tests
4
Pipeline layers
19
Rules
FastAPIspaCyReact / ViteSQLitePythonGitHub Actions
351-test suite · 100% accuracy on held-out set
02 / 04·AI / RAG

FinSight — Pre-Market Intelligence Platform

Production pre-market briefing system that runs every weekday morning. Ingests ~130 symbols, computes RSI and moving-average indicators, generates commentary with Gemini 2.5 Flash, stores everything in Postgres with pgvector, and emails confirmed subscribers via Resend. A retrieval-augmented Q&A layer answers plain-English questions over the full briefing history with inline date citations.

~130 / day
Symbols
68 passing
Tests
Gemini 2.5
LLM
pgvector
Vector store
FastAPIPostgreSQL · pgvectorGemini 2.5 FlashRAGGitHub ActionsRender · Resend
Live on Render
03 / 04·LLM Alignment

StarCoder2 Self-Alignment Pipeline

Implemented the SelfOSSInstruct methodology from the StarCoder2 paper to generate a TypeScript instruction-tuning dataset. Extracted functions from The Stack v2 using tree-sitter AST parsing and TypeScript compiler type-checking, then ran an S→C→I→R chain (Seed → Concepts → Instructions → Responses) via StarCoder2-3B on a T4 GPU. Produced 448 complete instruction-response pairs from 5,791 type-checked seeds. Filtered final outputs with model-based quality scoring.

5,791
Quality seeds
30,000
Raw files
448
Pairs
vLLM · bs 32
Inference
vLLMHugging Facetree-sitterStarCoder2-3BArrow / DatasetsPython
30k files → 5,791 seeds → 448 pairs
04 / 04·Data Engineering

NYC Taxi Medallion Data Pipeline

End-to-end data engineering pipeline over 5.97M real NYC TLC Yellow Taxi trips (Jan–Feb 2024): raw Parquet → PySpark cleaning → dbt gold marts → 23 data-quality tests → daily Airflow DAG. Built on a DuckDB warehouse with a Streamlit dashboard layer. Analysis surfaced Manhattan as 75% of total revenue ($111.9M of $149.1M).

5.44M
Clean trips
23
DQ tests
$149.1M
Revenue
DuckDB
Warehouse
PySparkdbtDuckDBAirflowStreamlit
5.97M rows · medallion architecture
More work · 4

Other projects.

05 · Analytics

Olist SQL Analysis & Hypothesis Testing

End-to-end SQL analysis on real Olist Brazilian e-commerce data (99,441 orders across 96,096 customers). Six DuckDB window-function queries power a monthly cohort retention matrix. A formal power analysis and two-proportion z-test confirmed a payment-method retention hypothesis as null (p = 0.76) with 80% statistical power on a 0.34 pp minimum detectable effect.

DuckDBstatsmodelsscipypandas
96k customers · 6 SQL queries96k
06 · Machine Learning

Loan Approval Risk Prediction

End-to-end classification pipeline on 20,000 loan applications (36 features, 76.1% rejection baseline). Applied SMOTE on training folds only to avoid data leakage. Compared 6 models — Logistic Regression, Decision Tree, Random Forest, SVM, KNN, and ANN — with GridSearchCV. SHAP identified CreditScore, AnnualIncome, and DebtToIncomeRatio as top predictors, packaged for compliance-ready reporting.

SMOTESHAPScikit-learnTensorFlow
76% rejection baseline · 6 models20k
07 · Big Data

Cryptocurrency Market Analysis on Hadoop

Three distributed MapReduce jobs in Java analyzing 2 GB of historical OHLCV tick data across 100+ cryptocurrency pairs (Binance, Apr–Aug 2024) on a multi-node AWS EC2 Hadoop cluster. Jobs surface volatility rankings, worst-performing assets by open-to-close change, and cumulative volume leaders with peak timestamps.

HadoopMapReduceJavaHDFS
3 MR jobs · multi-node cluster2 GB
08 · Backend

User Management System

FastAPI + PostgreSQL backend with JWT OAuth2 authentication, role-based access control (Admin, Manager, User), and profile-completion tracking. Diagnosed and resolved 5 critical production bugs across CI failures, unique-constraint violations, routing 404s, nested transaction errors, and mocking issues. Added 10 edge-case tests (138 passing) and shipped full CI/CD with GitHub Actions and Docker.

FastAPIPostgreSQLDockerGitHub Actions
RBAC · 3 roles · JWT OAuth2138
Experience · 2 roles

Where I’ve worked.

4y total · Scale AI → HCL Tech

Scale AI

Current
Applied AI Engineer
Jan 2026 – Present
United States
  • Built Python, SQL, and PySpark pipelines processing 10K+ records, cutting manual validation effort by 30%.
  • Prepared training and evaluation datasets and engineered statistical and ML features (Pandas, NumPy, Scikit-learn) for automated anomaly and data-quality detection.
  • Built classification models (Scikit-learn, PyTorch) with cross-validation, precision/recall analysis, and threshold tuning to improve minority-class detection.
  • Integrated Hugging Face transformers and embeddings for text classification and semantic evaluation, reducing manual review activity by 25%.
  • Implemented embedding-based retrieval workflows and Python REST APIs for similarity search and AI/LLM evaluation.
  • Supported model deployment using MLflow, Docker, Kubernetes, Azure, and Databricks.
  • Monitored model quality, data drift, latency, and pipeline failures; optimized PySpark workloads reducing processing time by 28%.
  • Created Power BI dashboards tracking AI/ML performance, data-quality trends, and business KPIs.

HCL Tech

Data Scientist
Jun 2021 – Jul 2024
India
  • Partnered with enterprise business stakeholders to translate retention and growth objectives into churn and forecasting use cases.
  • Conducted EDA and statistical analysis (Python, Pandas, NumPy, SQL) to uncover customer behavior patterns and key churn drivers.
  • Engineered 30+ customer-level features (purchase frequency, recency, transaction value, engagement) for predictive modeling.
  • Built and compared Logistic Regression, Random Forest, and Gradient Boosting classifiers with cross-validation and ROC-AUC, improving high-risk customer recall by 18%.
  • Developed demand forecasting models reducing forecast error by ~15%, supporting operational planning.
  • Processed 5K+ transaction records with PySpark and Databricks, cutting recurring batch runtime by 35%.
  • Automated data and model workflows using Airflow, AWS, and MLflow; delivered Tableau dashboards covering churn segments, forecasts, and KPIs.
Education · Class of 2026

Master of Science in Data Science

New Jersey Institute of Technology · Ying Wu College of Computing
Graduated May 2026 · Newark, NJ · GPA 3.7
Technical Stack · 48 tools

Everything I actually use.

Hover a category to isolate · click a tool to lock focus
7
Categories
48
Tools
4y
In production
Certifications

Credentials & Training

Anthropic Education

AI Fluency for Builders

Covers how to design and build production-ready systems using Claude — prompt engineering, tool use, multi-step pipeline

Jul 2026Verify
Anthropic Education

Introduction to Agent Skills

Hands-on course covering agentic AI design patterns, tool-calling, memory management, and orchestrating multi-step Claud

Jul 2026Verify
DeepLearning.AI · Stanford · Andrew Ng

Supervised Machine Learning: Regression and Classification

Covers linear and logistic regression, gradient descent, regularization, and neural network fundamentals — the core curr

Feb 2025 · CourseraVerify
AI Fluency: Framework & FoundationsAnthropic Education
Verify
Claude Code in ActionAnthropic Education
Claude Platform 101Anthropic Education
Google Cloud Network Engineer ProfessionalCoursera
Verify
AWS Cloud Virtual InternshipAWS Academy · EduSkills · AICTE
Salesforce Administrator Virtual InternshipSmartInternz · AICTE
Verify
Get in touch

Let's work together.

I'm actively looking for full-time roles as a Data Scientist, AI Engineer, ML Engineer, or Data Engineer. Onsite, hybrid, or remote — anywhere in the US.

Available immediately
Open to relocate anywhere in the US
Download Resume