Anuj Singh

AI Engineer | Full-Stack Developer | Data Scientist

Professional Summary

AI Engineer and Data Scientist with 2+ years building production machine learning models, LLM-powered applications, and full-stack SaaS products. Reduced model deployment time 40% via MLOps (Docker, CI/CD, model versioning). Increased student retention 25% using predictive analytics and automated intervention workflows. Shipped 5 production applications including a SOTA Causal Inference Toolkit (DoWhy, EconML, SCM, DiD) and a RAG-based GovTech platform (LangChain, Pinecone, Hugging Face) achieving 94% eligibility accuracy across 12 languages. Core stack: Python, TypeScript, PyTorch, Causal Inference, React, Next.js, FastAPI, Kubernetes, Docker, AWS.

Technical Skills

Languages: Python, TypeScript, JavaScript, SQL, Go
AI / Machine Learning: Causal Inference (DoWhy, EconML), Synthetic Control Method (SCM), Difference-in-Differences (DiD), Sensitivity Analysis, Large Language Models (LLMs), RAG, NLP, LangChain, LangGraph, Hugging Face, OpenAI API, PyTorch, Scikit-learn, Pandas, NumPy, Computer Vision, OCR
MLOps & Data Engineering: CI/CD Pipelines, Docker, Model Versioning, ETL Pipelines, Data Pipelines, A/B Testing (Bayesian/Sequential), Uplift Modeling, Experiment Tracking
Frontend: React, Next.js, Tailwind CSS, Framer Motion, HTML5, CSS3
Backend & Databases: Node.js, FastAPI, Flask, PostgreSQL, Redis, Supabase, Pinecone, Vector Databases
Cloud & DevOps: Amazon Web Services (AWS), AWS Textract, Docker, Kubernetes, Vercel, GitHub Actions, CI/CD, Linux
Tools & Platforms: Git, GitHub, Figma, Stripe, Jupyter Notebook, VS Code, Agile, Scrum

Professional Experience

Data Scientist — The Boring Education
  • Reduced model deployment time by 40% by implementing MLOps practices (CI/CD, monitoring, versioning) using Docker, Kubernetes, and AWS cloud infrastructure.
  • Increased student retention by 25% by building predictive dropout models (Scikit-learn, Pandas) that triggered automated early-intervention workflows for at-risk learners.
  • Published 3 internal analytics frameworks adopted across product, marketing, and curriculum teams, standardizing experimentation and KPI tracking.
  • Mentored 2 junior data scientists and established code review and documentation standards, improving team velocity and model reproducibility.

Projects

YojanaSetu — AI-Powered GovTech Platform GitHub Live Demo

Next.js, React, TypeScript, Python, FastAPI, OpenAI API, LangChain, Pinecone, PostgreSQL, AWS

  • Architected a RAG-based eligibility engine (LangChain + Pinecone vector database + Hugging Face embeddings) achieving 94% accuracy on citizen eligibility verification across 500+ government scheme documents.
  • Engineered a multi-lingual voice agent (OpenAI Whisper + local LLMs) serving users in 12 regional languages, designed for non-literate populations.
  • Optimized Next.js frontend (code-splitting, lazy-loading) and PostgreSQL queries, cutting average API response time by 40%.
AutoInvoice OCR — B2B SaaS Invoice Extraction GitHub Live Demo

Python, AWS Textract, Tesseract OCR, Next.js, Stripe, Supabase

  • Built a micro-SaaS extracting structured JSON from scanned invoices using Tesseract OCR and AWS Textract with LLM layout parsing, achieving 96% field-level accuracy on non-standard formats.
  • Shipped end-to-end billing integration (Stripe), ERP API connectors, and a confidence-scoring review UI with human-in-the-loop validation.
Agent Parliament — Multi-Agent Decision Simulator GitHub

Next.js, React, TypeScript, OpenAI API, LangChain, Framer Motion

  • Architected a 5-agent debate framework (Optimist, Pessimist, Engineer, Lawyer, User Advocate) with distinct persona prompts covering speed, security, UX, and compliance trade-offs.
  • Built consensus synthesis engine with weighted voting, outputting structured decision documents with confidence scores and actionable items.
Causal Inference Toolkit — Production Library & Dashboard GitHub Docs & Demo

Python, DoWhy, EconML, Streamlit, Sensitivity Analysis, Synthetic Control, DiD, Uplift Modeling, Pytest

  • Engineered a production-ready causal inference toolkit unifying DoWhy identification/refutation and EconML heterogeneous effect estimation (IPW, Doubly Robust, Double ML, Causal Forests).
  • Implemented quasi-experimental engines (Synthetic Control Method with SLSQP weight optimization, 2x2 & TWFE Difference-in-Differences event study panel models) and multi-layer sensitivity analysis (Rosenbaum bounds, Cinelli-Hazlett, E-values).
  • Shipped an interactive zero-code Streamlit web dashboard (`causal-toolkit app`) and automated executive HTML report generator with custom causal graph and TIPS contour visualizers.
DevRel GitHub Agent — Autonomous Issue Triage GitHub Live Demo

TypeScript, Probot, OpenAI API, GitHub Actions

  • Created an autonomous GitHub App (Probot) that triages issues via semantic analysis, auto-labels them, and drafts PRs for trivial fixes — designed to resolve ~25% of routine issues without human intervention.

Publications

  • "Efficient LLM Fine-Tuning" — Peer-reviewed research paper on parameter-efficient fine-tuning techniques (LoRA, QLoRA) for large language models. Cited 15+ times.

Education

Bachelor of Engineering (BE), Computer Science & Engineering — Chandigarh University
  • Developed university's official student portal serving 4,000+ active students daily.

Certifications

  • Project Planning and Machine Learning — University of Colorado Boulder (Coursera, Credential ID: N0O5KQNJM8K3)
  • React Native — Meta (Coursera, Credential ID: MQUI4TAY8HR8)

Interests & Activities

  • Open Source: Active contributor on GitHub with 4+ production-deployed projects and multiple public repositories.
  • Languages: English (Professional Working Proficiency), Hindi (Native).