Founder & AI engineer · Germany & India

I build the thing,
then I ship it.

I’m Nikhilvarma, a founder and AI engineer. I build AI products end to end. Ganymede decides which delinquent borrower is worth an agent-minute, then coaches the call that follows. Witness turns a photograph of a returned gearbox part into an ISO failure record linked to the batch that made it. FirstChair reads how ChatGPT, Gemini and Perplexity describe a law firm against its competitors. Before all of it I spent eighteen months inside a US fintech, rebuilding a monolith into event-driven services for 500+ concurrent users and going from data engineer to lead developer. Six products are live below, with seven investigations and a peer-reviewed paper behind them.

Built with
  • Python×10
  • Vercel×6
  • pandas×4
  • scikit-learn×4
  • TypeScript×4
  • Next.js×3
  • LLM orchestration×2
  • Supabase×2

Shipped & live · 6

Things I’ve shipped

Not prototypes and not screenshots. Six products running in production right now, each with the build written up: what it does, how it is put together, and what each decision cost.

01 · Ganymede

01 · Shipped & live · 2026

Ganymede

Which delinquent borrower is worth an agent-minute, and what to say once the call starts.

A collections team has more accounts in arrears than it has hours to call them. Ganymede decides which accounts are worth calling, then helps the agent through the call that follows. It ranks by the money a call is expected to recover rather than by the probability the borrower defaults. Those two orderings are not the same, and the difference is most of the value. The outcome of each call becomes the label that retrains the model that picked it.

The questionA risk score sorts borrowers by how likely they are to get worse. A collections team does not need that list. It needs the list of calls that recover the most money per agent-minute. How much does the difference between those two lists actually cost?

How it’s put together

  1. 01

    Panel

    • Loan-level servicing history
    • Monthly borrower panel
    • Contact events by channel

    Real dates, real delinquency transitions.

  2. 02

    Score

    • L1 trajectory: does this worsen
    • L2 self-cure: does it fix itself
    • L4 promise kept
    • Calibration, then reason codes

    Two questions, not one.

  3. 03

    Allocate

    • Uplift over self-cure
    • Weighted by exposure
    • Under a capacity constraint
    • Do not contact is a scored action

    Expected value, never probability.

  4. 04

    Coach

    • Tier 1 renders under a millisecond
    • Tier 2 waits for the next pause
    • Branches on borrower state

    Two tiers, set by the measured gap.

  5. 05

    Learn

    • Every decision logs its arm and propensity
    • Promise resolved against payment
    • Drift monitors on input and score
One loop, two lenses. The risk lens picks the conversation. The coach lens shapes it. The outcome of that conversation is the label both lenses retrain on. Split them and each half degrades: a queue nobody knows how to work, or advice with no idea who it is talking to.

Rough sketch

Allocator studio

Agent desk

Evidence

The allocator studio leads with both queues side by side, because the comparison is the finding. A single ranked list would hide it.

The architecture diagram and the wireframe are drawn wide.See them on the write-up →

What happens, step by step

  1. Build the panel. Loan-level servicing history becomes a monthly borrower panel. Calendar dates are the point: a timing feature from a source without real dates is refused at the feature layer rather than caught in review.
  2. Score two questions, not one. L1 asks whether the account worsens over the next ninety days. L2 asks whether it recovers without contact. An agent reads those numbers at face value, so calibration is the gate rather than AUC.
  3. Rank by money, not by risk. The allocator maximises expected recovered value per agent-minute: uplift over self-cure, weighted by exposure, under the capacity the team actually has. Probability never sorts the queue.
  4. Coach inside the measured gap. 328 inter-turn gaps from a real call give a median of 479 ms and a p25 of 292 ms. The budget is 300 ms because that is what the distribution allows. A hint that misses the gap arrives after the moment it was for.
  5. Log the decision, then resolve it. Every score, hint and override is written down with its experiment arm and its propensity. When the payment arrives or does not, the promise resolves and both lenses retrain on the result.

A decision that shaped it

Rank by expected valueinstead of Rank by probability of default

Risk-ranking calls a 1,289 euro account early and never reaches a 1.93 million euro one anywhere in the capacity sweep. Twelve agent-minutes cost more than the whole uplift on the small account is worth. At 15% capacity the value ordering recovers 59% more using roughly half the contacts. At 60% the edge falls to 2.4%, because with enough agents to call everyone the ordering stops mattering. The gain lives exactly where the constraint is real.

  • Python
  • Polars
  • LightGBM
  • scikit-learn
  • Calibration
  • Uplift modelling
  • LLM orchestration
  • Vercel
02 · Witness

02 · Shipped & live · 2026

Witness

A photograph of a returned gearbox part becomes an ISO failure record linked to the batch that made it.

A returned gearbox part arrives with a complaint and almost nothing else. Someone photographs it, writes one sentence in a spreadsheet, and puts the part in a bin. When the same damage appears on a later batch, nobody can prove it, because the first record said "worn" and the second said "pitting" and neither cited a standard. Witness makes the record the product. A photograph becomes a record with a damage mode, a clause number, a severity, a cause and a link to the batch that made the part. The model suggests. The inspector decides.

The questionThe failure was never the hard part. The record was. How do you turn a photograph and a one-line complaint into evidence a warranty claim can stand on?

How it’s put together

  1. 01

    Intake

    • EXIF read
    • Provenance scored
    • Perceptual hash against the tenant's assets
    • A re-sent photo is caught as a duplicate

    Before anything is believed.

  2. 02

    Enrol

    • The tenant's own known-good photos
    • A coreset is trained from them
    • Threshold calibrated leave-one-image-out

    Per part family, per tenant.

  3. 03

    Detect

    • PatchCore-style memory bank
    • Nearest-neighbour distance
    • Anomaly score and heat map
    • No database credentials in the function

    Abnormal against normal, nothing more.

  4. 04

    Classify

    • ISO 15243 for bearings
    • ISO 10825 for gear teeth
    • An off-list code is discarded
    • Confidence gate at 0.70

    A forced choice, never free text.

  5. 05

    Decide

    • High confidence files itself
    • Low confidence queues for review
    • Inspector confirms or corrects
    • Model suggestion kept beside it, immutable

    The human call is authoritative.

  6. 06

    Attribute

    • Mode, severity and cause as three axes
    • Finding linked to batch and supplier
    • Warranty report
    • Label export as model and human pairs
The confidence gate is the one place the software decides whether a person is needed. Everything before it is a pipeline stage that does one job and hands off. Everything after it is a human decision that the system records and never overwrites.

Rough sketch

Review queue

Finding detail

Insights cube

The review queue is the product's centre of gravity. It is the only screen where a person changes what the system believes.

The architecture diagram and the wireframe are drawn wide.See them on the write-up →

What happens, step by step

  1. Read the photograph before trusting it. EXIF gives a provenance score. A perceptual hash is checked against the tenant's other assets, so the same part photographed twice and sent twice is recognised rather than counted twice.
  2. Learn what good looks like. Enrolment is the tenant's own photographs of undamaged parts, per part family. There is no synthetic defect data anywhere in the system, because a generated defect teaches the model the generator.
  3. Ask only whether the part is abnormal. Stage one is a memory bank of normal patch features scored by nearest-neighbour distance. It answers one question. It does not name the damage.
  4. Force the answer onto a published standard. Stage two picks one code from the part family's ISO catalogue. It cannot invent a code, because an invented one is discarded rather than stored. The finding carries a clause number, so it traces back to the document.
  5. Send the unsure ones to a person. Above the gate the record files itself. Below it the record waits in the review queue. The threshold is visible on the page rather than buried in a config file.
  6. Record the mechanism, not only the surface. The same pit can come from fatigue at end of life or from a contaminant dent that started it early. The damage mode looks identical. The batch-level action does not, so attribution is a separate axis.
  7. Close the loop. Every confirmed record becomes a model-and-human pair in the label export. The training data is the inspector's work, collected as a by-product of doing it.

A decision that shaped it

Classify against ISO 15243 and ISO 10825instead of A house taxonomy of damage names

A house taxonomy is a private opinion. A warranty claim needs a clause number from a published standard, and two inspectors using the same catalogue disagree far less than two inspectors writing prose. The standard also makes the record portable to a customer who never saw this tool.

  • TypeScript
  • Next.js
  • React
  • Supabase
  • Postgres RLS
  • Python
  • NumPy
  • Computer vision
  • Vercel

The other four, in brief

Data & research · 7

Questions worth measuring

Every figure below came out of the analysis, including the one that found nothing.

Currently

Studying, building, and looking

M.Sc. Big Data & Business Analytics at FOM Hochschule, through August 2027. Shipping products alongside it. Open to data, software and AI engineering roles anywhere in Germany, and to full-time work in India.