I engineer production GenAI — RAG, LLM systems & multi-agent platforms.
I'm Ansh Mangukiya, an AI Engineer at Commercient. I spend my days making GenAI work in production — chatbot platforms, fine-tuned models, and inference stacks that real customers depend on.
80%
faster customer onboarding via LLM-generated ERP→CRM connectors
150+
tokens/sec throughput on a custom multi-GPU vLLM inference stack
70%
less engineering support load via a LangGraph conversational agent
60%
reduction in manual ticket triage with fine-tuned domain models
Experience
Production GenAI, shipped daily
AI Engineer · Commercient LLC
Jan 2025 — Present · Hybrid
- Architected a no-code ingestion & chatbot platform that lets enterprise customers launch domain-specific GenAI assistants from their own PDFs, web pages, and internal docs.
- Engineered a custom multi-GPU LLM inference stack on vLLM — advanced batching, KV-cache optimization — sustaining 150+ tokens/sec under concurrent load.
- Designed an LLM-powered SQL view generator that auto-templates ERP-to-CRM data connectors, eliminating 80% of manual integration coding during onboarding.
- Built a conversational ERP↔CRM sync agent with LangGraph so non-technical users configure complex integrations in plain English — cutting engineering support load by 70%.
- Shipped a customer-facing LLM fine-tuning platform (model distillation + RLHF feedback loops); tuned models reduced manual ticket triage by 60%.
- Implemented RAG retrieval pipelines with vector embeddings & semantic search, and bridged .NET enterprise systems with Python AI services via FastAPI microservices — all containerized with Docker + CI/CD on multi-GPU infrastructure.
Projects
Real problems, shipped solutions
All repos on GitHubToolbox
The stack I work with
GenAI & LLMs
Frameworks & Libraries
Vector & Retrieval
Engineering
Cloud & MLOps
Data
Credentials
Education & certifications
Education
B.E. Computer Science
Sarvajanik College of Engineering & Technology, Surat

Based in Surat, India · working worldwide
About
The person behind the code
I'm Ansh. I work as an AI Engineer at Commercient, where I build the systems behind GenAI products — chatbot platforms, fine-tuning pipelines, and the inference infrastructure that keeps them running.
One thing production has taught me: the model is usually the least important part. Good ingestion, clean chunking, proper indexing, and honest evaluation matter more than whatever model topped this week's benchmark.
Where it started
First year of computer science at Sarvajanik College, Surat. My GitHub still has the tic-tac-toe games I built while learning Java — I keep them public as a reminder that everyone starts somewhere.
Learning by shipping
A recommendation engine that matches resumes to companies. An NLP platform. A price-prediction pipeline that scraped its own training data. A chatbot inside a real e-commerce store. None of them stopped at the notebook — each one had to become something a person could actually click.
The jump to production
Joined Commercient as an AI Engineer while still finishing my degree, then graduated with a 9.78/10 CGPA. The first lesson production taught me: a demo only has to work once. A product has to work every time.
Making AI hold up, at Commercient
Still at Commercient — shipping 150+ tokens/sec from a multi-GPU vLLM stack, cutting manual integration coding by 80%, and engineering support load by 70%. None of it came from a bigger model — it came from better retrieval, cleaner pipelines, and honest evaluation. That's the part of this field I like most.
Connect
Let's build something
A collaboration, an opportunity, or just a question about GenAI — drop a message. I reply within 24 hours.