Open to ML roles Currently building PageServe ↗

ML engineer working on LLM inference and multi-agent systems. I wrote a vLLM-style serving engine with a paged KV cache, a reinforcement-learning library in C++ that brings its own autograd, and then wrote up every single thing that broke. It was a long list.

nikhil.exe online
Jaipur, IN train/loss0.000

full precision

focus
LLM inference & multi-agent systems
building
PageServe & RLForge
writing
Deep-dives on X~2M impressions since July
based
Jaipur · --:-- IST

01About

The math decides when to stop, not the model's confidence.

temperature

Retrieval scores over self-reported certainty. Pruned weights over bloated checkpoints. Shipped code over benchmark screenshots.

Lately that means going one layer down. I write the scheduler instead of calling generate(), and the autograd instead of importing it. It's slower, it's humbling, and it's the only way I've found to actually understand the thing instead of nodding along to it.

BIT Mesra, studying what doesn't fit in a lecture hall. I write up everything I build, and I'm always maintaining the unbroken eye contact with that single code block like we have beef.

languages
  • Python
  • C++20
  • TypeScript
  • SQL
modeling
  • PyTorch
  • HuggingFace
  • PEFT / LoRA
  • scikit-learn
inference
  • Continuous batching
  • Paged KV cache
  • Pruning
  • ONNX
  • Ollama
agents · rag
  • LangGraph
  • ChromaDB
  • FAISS
  • RAGAS
  • LangSmith
infra
  • FastAPI
  • Docker
  • AWS
  • Redis
  • OpenTelemetry
$ nikhil --infov2026
location
Jaipur, India
college
BIT Mesra
focus
Multi-agent systems
cf
1400+ Specialist
builds
SIH '24 · IIIT-D '23
reading
The Humans, Matt Haig
status
Open to ML roles

Ask this page

BM25 over this page's own text, right in your browser. No LLM, no API, no making things up. If the score is too low, it says so.

built on HiveMind for the rule: refuse when the retrieval score is too low ↗ CodeLens for the idea: offline search by intent (embeddings there, BM25 here) ↗

Right now

live, while you read
Jaipur · --:-- IST Probably building something Replies usually within a day

02Experience

Where it runs in production.

Notebooks are cheap. These are the places where my code had to survive real users, real traffic and someone else's pager.

  1. 2025

    HireBuddy

    Software Engineer, ML

    • Built and deployed a resume-to-JD matching system on RAG + Transformers, and retired the legacy regex parsers for good.
    • Designed an NLP ingestion pipeline processing 10K+ resumes a day for automated shortlisting; shipped the APIs on FastAPI, Docker and AWS.
    • Set up a deterministic eval stack with RAGAS and LangSmith so improvements were measured, not vibed.
    +35%shortlist accuracy
    −28%chain latency
    +40%user engagement
  2. Founder

    DevPath

    Live product · dev-path.site

    Structured learning for developers stuck in tutorial hell. Curated roadmaps and one focused task a day, so the hardest decision is already made for you. Designed, built and shipped solo. Free to start.

  3. Research

    IIT Kharagpur

    Model compression collaboration

    Structured pruning of U-Net for biomedical segmentation: 97.3% fewer parameters, 92% fewer FLOPs, IoU still above 0.95 on MoNuSeg. Turns out most of the network was just along for the ride.

03Work10 projects

Rebuilt from first principles. No black boxes I didn't open.

Two systems I wrote end to end because reading the library source stopped being enough. One serves LLMs. The other trains agents.

LLM inference & serving

★ 97 ⑂ 10

PageServe

A vLLM-style inference engine, grown from a naive generate() loop in 11 phases.

Continuous batching, a paged KV cache over a physical tensor pool, prefill/decode budgets, and CPU swap, so under load it slows down politely instead of OOMing. Along the way I found a cache bug quietly charging 1.1 ms per token, and evicted it.

−98.5%
mean TTFT at 4-way load
1,418 → 20.9 ms
−26.5%
wall-clock latency
vs sequential
11
build phases,
each runnable
kv pool · 64 blocks × 16 tokens live
pageserve-phi.vercel.app open ↗

Reinforcement learning · C++20

★ 11

RLForge

A reinforcement-learning library with no import torch, down to its own autograd.

Tensor engine, autograd, optimizers, Q-learning, DQN and PPO in C++20 with zero dependencies, including the diamond-graph gradient case that most from-scratch engines get wrong without telling anyone.

Environments, tensor/autograd and DQN by me. PPO, threading and CUDA/BLAS backends co-built with aprv10.

148/148
tests passing
warnings as errors
0
runtime
dependencies
−100→3
GridWorld eval return
3 = proven optimal
backward pass · diamond dependency tensor.cpp
x ax · 2 bx · 3 ca + b lossmean(c) forward ↑ grad ↓ ∂x += 2/n ∂x += 3/n → 5/n ✓
$ cmake --build build -j && ctest
100% tests passed, 0 failed out of 148
$ ./train --env gridworld --steps 20000
eval return -100.0 → 3.0 # optimal
rl-forge.vercel.app open ↗
  • A 5-agent pipeline (Planner → Researcher → Critic → Writer → Evaluator) that loops until a deterministic confidence threshold is met, computed from mean retrieval score, not the LLM's self-report. The same rule "Ask this page" uses up there.

    Agents & RAGFastAPI · Redis · ChromaDB · OpenTelemetry · Docker

  • A VS Code extension that chunks code by AST with tree-sitter, embeds it locally through Ollama, and jumps you to the exact line from a plain-English query. No internet, no API keys. It reindexes on save and beat plain grep on every query in a 50-query test.

    ToolsTypeScript · tree-sitter · Ollama · VS Code API

  • Retrieval by LLM-guided tree traversal instead of embeddings. I wanted to know whether the vector-DB default was actually necessary, so I built the version without one.

    Agents & RAG

  • Fine-tuned a 767M summariser with 1.57M trainable params. Full fine-tuning fell apart on unseen domains; LoRA didn't.

    ModelsPyTorch · HuggingFace · PEFT

  • Research collaboration with IIT Kharagpur. Structured pruning of U-Net for biomedical segmentation: 97.3% fewer parameters, 92% fewer FLOPs, IoU still above 0.95 on MoNuSeg.

    Research

  • The full Transformer from the paper, multi-head attention to training loop. Not "used a framework."

    ModelsPyTorch

  • Skin-lesion classification across 7 classes with a 57:1 imbalance. Focal loss so the rare classes count, Grad-CAM so you can see what it looked at, ONNX export so it runs anywhere.

    Models

  • A senior engineer code-reviews your brain. Six-stage multimodal pipeline, real models, no affirmations.

    Tools

04Lab5 toys

Ideas I keep explaining, so now you can poke them.

Small toys for the ideas I explain to anyone who stands still long enough. Everything runs in your browser, nothing phones home.

Speculative decoding how fast inference cheats, honestly

A small draft model guesses 4 tokens ahead. The big model checks all of them in one forward pass, keeps the matching prefix, and fixes the first miss. Same output as normal decoding, fewer expensive passes.

target passes
…
tokens / pass
…
expected (1−αk+1)/(1−α)
…
vs one token per pass
…

The KV cache stores the ____

Temperature

Low T sharpens the distribution toward the top token; high T flattens it. Drag, then sample.

loss …

Learning rate

Gradient descent on a stretched bowl. Too small crawls, too large zig-zags, past 0.167 it diverges.

hover a token

causal mask · illustrative weights

Causal attention

Each token only attends to what came before it. Hover a word to see where it looks.

Tokenizer, the step before BPE type anything

GPT-2 splits text on words, numbers, punctuation and leading spaces before a single merge happens. This is that exact split, live. Real BPE then lands a little higher on rare words, which is why your name costs more than "the".

0 pre-tokens0 characters0 characters per token

05Writing20 articles

I build it, then I write it down. Usually in that order.

One part at a time. Scheduler traces, memory bugs, and the math that has to be exactly right or nothing works.

20 articles~2M impressions since July~120K views on inference posts~20K on RL posts

All articles on X ↗

What I cannot create, I do not understand.
Richard Feynman

06Contact

If you've made it this far, you might as well say hi.

Open to ML engineering roles in inference, serving and agent infrastructure. Email is fastest. DMs on X work too, I check them more than I should.

It's --:-- in Jaipur (IST)

tabular Q-learning, running live in your browser Even my agent is figuring out how to reach me. episode 0steps …ε 1.00

Elsewhere, I'm on X, where 3.1K people put up with my threads, on GitHub, where the commits are louder than the threads, and on LinkedIn for the version of me that wears a collar. is what I look like away from a terminal, and is the thing I built for people stuck in tutorial hell.

Writing · visible only to you

Add an article

Checked against the server once, then remembered in this browser only.