ML Engineer · Data Scientist · IIT (ISM) Dhanbad
I build ML systems that ship. From fraud detection achieving AUC 1.0 on 50K financial transactions to production-grade ETL pipelines, model monitoring with automated drift detection, and full-stack analytics platforms — I take problems from raw data all the way to deployed, measurable impact. 6 live production systems. 32 automated tests. Real business outcomes.
Production ML system on 50,000 credit card transactions. Manually implemented SMOTE in NumPy, trained 3 classifiers achieving AUC 1.0, deployed as a Gunicorn-served Flask REST API with 5-page interactive frontend. Full automated build-train-serve pipeline.
Production MLOps system simulating 4 drift types across 12 months on a credit scoring model. Automated detection via KS Test & PSI, AUC tracking from 0.85→0.72→0.85 post-retraining. Full model feedback loop deployed live.
Three forecasting algorithms (Gradient Boosting, Holt-Winters, Seasonal Naïve) evaluated via walk-forward cross-validation on Rossmann retail data. Best model: MAPE 9.4%. 8-page Flask + Plotly.js dashboard with live 90-day forecast.
End-to-end RFM analysis and K-Means clustering on 50,000 transactions. 7-page interactive dashboard covering segment profiles, revenue breakdown, trend analysis, and strategy recommendations. Built for executive and non-technical audiences.
Full-stack ad campaign analytics platform with JWT auth, role-based access, and 8 analytics pages. Three C++17 engines compiled to WebAssembly: A/B testing (Welch t-test), keyword matching (Porter Stemmer), click fraud detection (Z-score + bot CV). 32 automated tests.
End-to-end migration from Informatica/SSIS flat-file ETL to PySpark. Apache Airflow DAG orchestration, automated data quality checks (Z-score, null rates, referential integrity), RFM scoring, and Parquet output replicating AWS Glue patterns.
End-to-end Python ETL pipeline processing 48K+ listings into an optimised SQLite database. 12 targeted SQL queries uncovering $85M in Manhattan revenue potential and identifying high-demand neighbourhoods like Williamsburg.
Python EDA pipeline on 7K telecom customers. Applied NVF/NCRF scoring to segment users by value and risk, identifying $12M+ revenue at risk from Month-to-Month contracts with retention strategy recommendations.
Applied K-Means clustering to segment individuals by annual income and spending patterns. Identified 5 distinct customer clusters with centroid analysis, comprehensive visualisations, and pattern discovery from unsupervised learning.
Open to ML Engineer, Data Scientist, and Analytics roles where I can ship production systems that make a measurable difference. If you're building something ambitious — let's talk.