Hello, I'm Ashish

Taking AI from research to production

Models, agents and systems that scale across enterprise workloads.

Experience & Impact

I build research-driven AI that makes an impact in production. From fine‑tuning state‑of‑the‑art models to shipping agentic systems for millions of users, my experience spans the full lifecycle of modern ML.

3M+ requests/day
30B+ parameters served
90% cost reduction
10K+ calls/day summarized

Principal AI/ML Engineer

2023–Present

Leading AI/ML at Vanguard: designing enterprise platforms for LLMs, RAG, and real‑time speech analytics. Driving research to production through fine‑tuning, evaluation, deployment and scaling. Mentoring engineers and shaping AI strategy.

Senior AI/ML Engineer

2020–2023

Built hybrid RAG & agentic systems, optimized inference gateways (EKS + vLLM), and implemented guardrails and evaluation pipelines. Improved retrieval accuracy by 12% and doubled throughput of 30B/8B models using Paged Attention.

AI/ML Engineer

2018–2020

Developed enterprise call‑summarization and speech analytics platform. Delivered 90% cost reduction at 30× higher performance for summarization, supporting tens of millions of calls per month. Delivered MTA NYCT internship projects during graduate studies.

Projects & Systems

Selected work spanning LLM platforms, agentic AI and speech intelligence.

Hybrid RAG Agent Assist

Built a hybrid retrieval‑augmented generation system combining Pinecone, in‑house vector DBs and proprietary LLMs. Improved retrieval accuracy by 12%, delivered 3M+ daily queries and orchestrated agentic flows for advice and support use cases.

  • Powered by Pinecone + bespoke agent orchestrator
  • Expanded adoption across teams via SDKs & secure auth
  • Pinecone customer story

Inference Gateway & Serving

Designed an inference gateway on Kubernetes/EKS leveraging vLLM and Paged Attention. Achieved 2× throughput improvement on 30B & 8B models, integrated safety guardrails and evaluation pipelines, and served 3M+ daily requests.

  • Model‑agnostic hosting & autoscaling
  • Evaluations & guardrail architecture
  • SDKs for internal & external integration

Speech Intelligence Platform

Built a distributed speech intelligence system converting noisy call audio into readable transcripts enriched with journey, sentiment and emotion signals.

  • 10M+ calls/month processed
  • Real‑time ASR & diarization
  • Enriched with journey & sentiment insights

CALM Deployment Framework

Developed CALM—cost‑efficient, model‑agnostic LLM deployment framework. Delivered up to 5× inference cost reductions via dynamic autoscaling, dedicated tenancy and intelligent resource orchestration.

  • Cost‑aware autoscaling & scheduling
  • Production reliability & failover
  • Read the paper

Financial Advisor Augmentation

Trained a domain‑specialized model for financial advisors and planners using SFT, off‑policy & on‑policy distillation and GRPO. Built a multi‑agent system with short‑ and long‑term memory to orchestrate skills such as retirement, estate, investment planning and compliance.

  • Custom eval dataset for quality & compliance
  • Agent orchestration across skills & memory
  • Tax, asset allocation & behavioral coaching

Research & Publications

CALM: Cost‑efficient, model‑Agnostic, Low‑latency Modular LLM Deployment Framework

Presented at the 2026 International Conference on Advances in Artificial Intelligence and Machine Learning (AAIML). CALM addresses the cost and complexity of deploying LLMs in production by integrating dynamic autoscaling, workload‑aware dedicated tenancy and intelligent resource orchestration. Experiments show up to 5× cost reductions compared to managed hosted LLMs under high load. The paper provides practical patterns and guidelines for teams building scalable, reliable and cost‑optimized LLM applications.

Read Publication

Get In Touch

I'm always interested in discussing new ideas, collaborations or speaking opportunities. Feel free to reach out!