ClauseIQ
Agentic RAG for legal document intelligence — analyzes contracts and returns answers with clause-level citations using hybrid retrieval and vector search.
I build AI systems that hold up in production — grounding LLMs so they don't hallucinate, pipelines that fail loudly on bad data, and models evaluated honestly and served behind clean APIs.
SPPU · 8.82 CGPA · A+ Grade
"Grounding LLMs so they don't hallucinate, pipelines that fail loudly on bad data, and ML that's evaluated honestly and served behind an API."
Maharashtra · On-site
Three principles I don't compromise on — they're the difference between a demo and a system you can trust.
LLMs cite their sources. RAG with hybrid retrieval and clause-level citations means answers you can verify — not confident hallucination.
Pipelines fail loudly on bad data. Drift is caught with PSI, KS and Jensen-Shannon before it silently degrades a model in production.
Models evaluated without cheating — proper validation, no leakage — then served behind clean, documented FastAPI endpoints.
Lead the development and delivery of robust data solutions for MNC clients and research institutions, drive business development across verticals, and serve as experienced faculty — training batches of students and interns in Data Science, AI, YOLO, and web development.
Delivered an AI-integrated dashboard for multiple Maharashtra districts as an independent freelance engagement — consolidating district-level data into an interactive decision-support tool with intelligent analytics for administrators.
Standardized data from diverse global formats into structured JSON and designed an optimized AWS S3 ingestion & storage pipeline. Built and deployed enterprise tools — a Fintech Chatbot, an automated Email Verifier, and a Document Processor — while handling server management, backend scaling, and continuous deployment for critical financial services.
Contributed to a Guinness World Record initiative — collecting, cleaning, organizing, and managing data at scale to ensure accuracy and efficiency across the record attempt.
Applied data-science methods to improve solar and renewable-energy systems — boosting processing efficiency by 60% with Python, leading 200 students across data transformation, validation, and storage (50% efficiency gain), and contributing to a world record for the largest folder of self-portraits.
Data analysis and project-management work across guided data-science tasks and deliverables.
Graduated with distinction, specializing in data science, machine learning, and applied AI. Received Graduate Trainee offers from TCS and Wipro — declined both to pursue hands-on, visionary work in AI, Data Science, and ML.
Production-minded AI & data systems — reliable, observable, honest. Every card links to its GitHub source and a one-click "Run in Colab" notebook — no setup needed.
Agentic RAG for legal document intelligence — analyzes contracts and returns answers with clause-level citations using hybrid retrieval and vector search.
Batch + streaming analytics lakehouse on DuckDB with Airflow orchestration — a full ELT platform turning raw retail events into query-ready insight.
Production ML monitoring & drift detection using PSI, KS and Jensen-Shannon — catches data and concept drift before it silently degrades models.
Real-time card-fraud detection on transaction streams with Kafka — low-latency scoring over live event streams.
Telecom churn prediction with a full MLOps lifecycle — MLflow tracking, scikit-learn models and FastAPI serving.
Automated EDA + baseline modelling + an LLM insight-narrative engine that writes human-readable findings from your data.
Multi-model demand forecasting with newsvendor inventory-optimization — bridging forecasting and operations research.
Hybrid semantic product search & recommendations using Reciprocal Rank Fusion over dense + lexical retrieval.
Hybrid LSTM + ARIMA stock-price forecasting on live Yahoo Finance data — blending deep learning with classical time series.
Portfolio management with risk metrics and allocation / performance visualizations built on pandas and NumPy.
Asynchronous email verification with FastAPI and a Streamlit interface for bulk processing.
Exploratory analysis of global terrorism data — surfacing trends, hotspots and patterns through visualization.
Hybrid retrieval-augmented generation that fuses BM25 and dense vectors via Reciprocal Rank Fusion, with a rigorous IR evaluation harness — runs fully offline, no API keys.
End-to-end fraud detection with leakage-safe features, calibrated probabilities and cost-based thresholds — ROC-AUC 0.99 on a time-based split, with a model card and scoring API.
Probabilistic multi-horizon demand forecasting — a Seq2Seq attention LSTM emitting calibrated P10/P50/P90 quantiles that beats a seasonal-naive baseline.
Document-understanding toolkit — hybrid rule+ML NER, section classification and key-value extraction that turns invoices and resumes into structured JSON.
Causal uplift modeling (S/T-learners) that targets persuadable customers, evaluated with Qini curves — it recovers the true treatment effect, not just churn risk.
Streaming feature engineering with windowed aggregations and point-in-time-correct serving — the leakage-safe join that keeps offline and online features in sync.
A framework for tool-using LLM agents — a ReAct reasoning loop, typed tool registry and a deterministic offline backend that makes agents genuinely unit-testable.
Real-time anomaly detection ensembling robust z-score, EWMA, seasonal and isolation-forest detectors, with matched batch and online streaming APIs.
PyTorch CNN for manufacturing surface-defect detection with occlusion-based saliency that localizes the defect — trains on CPU in under a minute.
Two-tower (dual-encoder) retrieval recommender trained with in-batch negatives, evaluated with Recall@K and NDCG — the architecture behind large-scale candidate generation.
Parameter-efficient fine-tuning with LoRA — trains low-rank adapters (<10% of params), merges them back, and matches full fine-tuning at a fraction of the cost.
Clinical-risk prediction built the way healthcare demands — calibrated probabilities, decision-curve analysis, SHAP explanations, and subgroup fairness checks.
A responsible-AI toolkit that audits a lending model — SHAP explanations plus demographic-parity, equalized-odds and disparate-impact metrics, then mitigates the bias.
Per-pixel semantic segmentation with a compact U-Net — Dice/IoU-evaluated defect and lesion masks that decisively beat an intensity-threshold baseline.
A from-scratch Graph Convolutional Network that detects fraud rings on a transaction graph, using network structure to beat a graph-blind model on AUC.
Actuarial pricing done properly — separate claim-frequency and severity models combined into a pure premium, evaluated with Gini and lift curves.
Analysis-first retail BI — RFM & cohort segmentation, market-basket association rules, price-elasticity modelling and rigorous A/B hypothesis testing, with an interactive dashboard.
Full-lifecycle time-series forecasting — a model zoo from seasonal-naive to gradient-boosted, rolling-origin backtesting, conformal prediction intervals and MLflow experiment tracking.
Classical-AI vehicle routing — Clarke-Wright savings, 2-opt, simulated annealing, genetic algorithms, ant-colony optimization and Google OR-Tools, plus A* grid pathfinding.
Responsible credit-default ML — leakage-safe pipelines, probability calibration, SHAP reason codes, a full fairness audit with a model card, served behind a FastAPI endpoint.
Deep-learning visual inspection — a PyTorch CNN (plus transfer learning) for surface-defect detection with mixed-precision training, Grad-CAM heatmaps and ONNX export.
Retrieval-first generative AI — hybrid dense + BM25 retrieval fused with Reciprocal Rank Fusion, MMR diversity re-ranking and a RAGAS-style evaluation harness. Runs fully offline.
Agentic AI framework — planner / researcher / critic agents over a ReAct tool loop, memory and a from-scratch StateGraph engine. Deterministic and unit-tested, no API key needed.
Have a project, a hard data problem, or an idea worth grounding in real data? I'm always glad to collaborate or trade notes on AI & data — reach out and I'll reply.
Email