Applied AI Engineer — London · Available now

I take AI from prototype to production — and make it trustworthy once it’s there.

3+ years shipping agentic AI, RAG and MLOps for banking, insurance and airline clients at Tata Consultancy Services. MSc Computing (AI & ML), Imperial College London — Distinction.

UK Graduate visa · no sponsorship required

faster query resolution with a production AI agent
60%
less infra deployment time via Terraform + CI/CD
90%
latency overhead for PII & RBAC guardrails
<50ms
faster LLM uncertainty estimation (MSc thesis)
13.7×

01 About

Models are the easy part. Systems are the work.

Most AI demos never meet a real user. The work I care about starts after the notebook: wrapping a model in a service people can depend on, feeding it the right data, putting guardrails around what goes in and what comes out, and watching it closely once it’s live.

For three years at Tata Consultancy Services I did that for insurance, banking and airline clients. I shipped an agentic knowledge assistant that cut query resolution time by 60%, a governance framework that took 90% out of infrastructure release time, and a document pipeline that handles submissions for 2,000+ employees, 90% of them without a human touching them.

“A model is a component. The product is everything around it.”

One question kept coming up in production: how do you know when an LLM is wrong? So I went to Imperial to work on it. My MSc thesis trains small language models to estimate a large model’s uncertainty in a single pass, 13.7× faster than the sampling method it replaces. I bring the same habit to every system I build: measure it, test it, monitor it, and make it cheap enough to run for real.

02 Experience

Three years of shipping AI that people actually use.

  1. – London, UK

    Education

    MSc Computing (AI & ML)

    Imperial College London · Distinction

    Thesis on real-time LLM trustworthiness, supervised by Dr. Matthew Wicker. Head of Technology at Imperial AI Society. Won 1st prize at the Imperial Hackathon.

  2. – India

    Full-time

    Applied AI Engineer

    Tata Consultancy Services

    Production GenAI and MLOps for enterprise clients in insurance, banking and airlines.

    • 60% faster query resolution
    • 90% less deployment time
    • 50k+ docs in vector search
    • 200+ daily queries served
    • 90% straight-through processing
    • 2,000+ employees served

    • Designed and built an agentic AI assistant in production that cut query resolution time by 60%.
    • Built FastAPI REST services behind a ReactJS frontend with OAuth, serving 200+ queries a day.
    • Used PostgreSQL for app data and Milvus vector search over 50,000+ documents. Integrated Azure OpenAI and open-weight LLMs served with vLLM.
    • Covered it with pytest unit and integration tests, then containerised it with Docker and deployed on Azure Kubernetes Service.
    • Agentic AI
    • RAG
    • FastAPI
    • ReactJS
    • OAuth
    • PostgreSQL
    • Milvus
    • Azure OpenAI
    • vLLM
    • pytest
    • Docker
    • AKS

    • Built a reusable AI governance framework with Terraform and GitLab CI/CD to standardise infra and policy releases, cutting deployment time by 90%.
    • Added runtime guardrails for RBAC and PII / sensitive-data checks on both model inputs and outputs, using Azure APIM and Azure Content Safety, with under 50 ms of added latency.
    • Set up Prometheus and Grafana to show service health and application metrics across the deployed environment.
    • Terraform
    • GitLab CI/CD
    • Azure APIM
    • Azure Content Safety
    • Prometheus
    • Grafana

    • Built a Python pipeline using Table Transformers and TrOCR to turn unstructured document images into structured outputs.
    • Reached 90% straight-through processing for 2,000+ employees, so most submissions need no manual review.
    • Python
    • Table Transformers
    • TrOCR

    Also: worked with 4+ engineering and research teams on banking, airline and insurance accounts on system design and code reviews, mentored junior engineers, and demoed each sprint to non-technical client stakeholders. TCS Excellence Award, 2024.

  3. – Mumbai, India

    Internship

    ML Intern

    Vidyalankar Institute of Technology

    • Built and deployed VBot, the official admissions chatbot on VIT’s website.
    • Designed 200+ admission query flows with Dialogflow intents and entities, with rich responses including images, video and links.
    • Dialogflow
    • Conversational AI

03 Selected work

Each project, from problem to result.

Each project is laid out as Situation → Task → Action → Result.

MSc thesis · Imperial · Apr–Sep 2026

0.801AUROC on held-out LLMs
vs 0.698 linear probe

  • 13.7× lower latency
  • 11.5× less compute
  • 13 LLMs, 7B–27B

Amortized Trustworthiness

Small language models that predict when a large one is unsure, in one pass instead of eleven.

SSituation
Semantic entropy is one of the best signals that an LLM is hallucinating. But it needs about 11 sampled generations and entailment clustering for every answer, which is too slow and costly for real-time use.
TTask
Train a small proxy model that predicts a large model’s semantic uncertainty in a single pass and still works on LLMs it has never seen.
AAction
Built a reusable PyTorch dataset pipeline across 13 LLMs, covering stochastic generation, hidden-state extraction and DeBERTa entailment clustering. LoRA fine-tuned Llama-3.2-3B with PEFT on text and hidden-state features. Benchmarked against linear-probe, ridge and representation-alignment baselines using leave-one-LLM-out evaluation.
RResult
0.801 AUROC vs 0.698 (linear probe) and 0.729 (aligned ridge), and 0.791 on unseen Qwen and Gemma models. Transfers from TriviaQA to SQuAD without retraining. Uncertainty goes from 11 LLM passes to one: 13.7× faster, 11.5× cheaper.
  • PyTorch
  • LoRA / PEFT
  • Transformers
  • Llama-3.2-3B
  • DeBERTa
  • Python

ML System Design · Imperial · Jan–Mar 2026

3–4msp90 latency
against a 3-second SLA

  • 0 missed AKI cases
  • 1,600 live messages
  • 4-person team

Real-Time AKI Detection

A fault-tolerant streaming inference service that alerts clinicians to Acute Kidney Injury in milliseconds.

SSituation
Acute Kidney Injury has to be caught while blood-test results are still streaming in. A missed alert is a clinical failure, and a late one is almost as bad.
TTask
Design and ship a real-time system that scores streaming patient data, pages the hospital within a 3-second SLA, and keeps working through failures.
AAction
Designed the system around a LightGBM classifier on a streaming pipeline with error handling and failure recovery. Added pytest unit, integration and end-to-end tests. Deployed with Docker on AKS, with GitLab CI and Prometheus monitoring.
RResult
p90 latency of 3–4 ms against a 3,000 ms budget, with no missed AKI cases across 1,600 patient-data messages.
  • Python
  • LightGBM
  • Docker
  • AKS
  • GitLab CI
  • Prometheus
  • pytest

Client · Production

Agentic Knowledge Hub

60% faster query resolution

An agentic RAG assistant for a European car insurer, searching 50k+ documents and serving 200+ queries a day.

STAR breakdown
S
Support staff spent too long hunting through policy and product documents.
T
Ship a secure, production-grade assistant that answers from internal knowledge.
A
Built an agentic RAG system with FastAPI, ReactJS, OAuth, Milvus and Azure OpenAI/vLLM, running on AKS.
R
Query resolution time dropped by 60%.
  • Agentic AI
  • RAG
  • Milvus
  • FastAPI
  • AKS

Client work, code under NDA

Client · Production

PF Document Pipeline

90% straight-through processing

Document AI that turns scanned forms into structured data for 2,000+ employees.

STAR breakdown
S
Provident Fund submissions arrived as unstructured document images that were checked by hand.
T
Automate extraction so most submissions need no human review.
A
Built a Python pipeline combining Table Transformers for layout with TrOCR for text recognition.
R
90% straight-through processing across 2,000+ employees.
  • Python
  • Table Transformers
  • TrOCR

Client work, code under NDA

Hackathon · Nov 2025

Study Peer-Matching

1st prize, Imperial Hackathon

Recommends study partners from student learning data and behaviour patterns.

STAR breakdown
S
Students struggle to find study partners who match their pace and goals.
T
Build a working matching product within the hackathon window as a team of four.
A
Modelled learning data and behaviour patterns to rank compatible peers, wrapped in a usable app.
R
Won first place.
  • Python
  • scikit-learn
  • Recommender systems

04 Skills

Tools, sorted by how far I’ve taken them.

  • Shipped in production
  • Built with in projects & research
  • Working knowledge

Tags marked ↘ jump to the project or role where I used them.

AI & ML

  • Agentic AI
  • RAG
  • Azure OpenAI
  • vLLM
  • Transformers
  • PyTorch
  • LoRA / PEFT
  • scikit-learn
  • LightGBM
  • LangGraph
  • LangChain
  • MCP

Backend & APIs

  • Python
  • FastAPI
  • REST APIs
  • OAuth
  • ReactJS
  • Flask
  • Django

Data & Search

  • PostgreSQL
  • Milvus
  • SQL
  • NumPy
  • Pandas
  • FAISS
  • SQLite

Cloud, MLOps & Infra

  • Azure
  • Kubernetes / AKS
  • Docker
  • Terraform
  • GitLab CI/CD
  • Azure APIM
  • AWS

Testing & Observability

  • pytest
  • Prometheus
  • Grafana
  • Git
  • GitHub
  • Postman
  • Linux

AI-assisted development

  • Claude Code
  • Codex

05 Education & recognition

Education and awards.

–

MSc Computing (Artificial Intelligence & Machine Learning)

Imperial College London · Distinction

Thesis: Amortized Trustworthiness: Training Small Language Models as real-time cross-model Semantic Uncertainty Proxies for LLMs, supervised by Dr. Matthew Wicker. Modules included Software Engineering for ML Systems, Scalable Systems & Data, ML Systems & Hardware, Formal Methods for Safe AI, Generative AI and Deep Learning.

–

B.E. Computer Engineering

University of Mumbai · CGPA 9.78 / 10

Ranked 1st university-wide in Database Management Systems, Computer Networks, and System Programming & Compiler Construction. Technical Team Member at IEEE-VIT: led seminars and mentored junior students.

Awards

  • 20251st Prize, Imperial Hackathon: study peer-matching app, team of four.
  • 2024TCS Excellence Awards: an individual award for a time-series POC, plus client-nominated team awards for production delivery.

Certifications

  • 2025Microsoft Certified: Azure AI Fundamentals (AI-900)
  • 2024AWS Certified Cloud Practitioner

06 Contact

Let’s build something reliable.

Open to Applied AI, ML Engineering and AI Platform roles in London or remote within the UK. I can start immediately.