Rajan Verma

I put open models on machines you already own.

Mail, tickets, and files never leave your server. I put the LLM on that machine so your product can call a model you own. Your engineer gets a private URL, a token, and a runbook they can run without me.

What I can do

On-prem inference An LLM on your GPU — CPU if that is the machine — that your product calls with a token. Mail and tickets never need a public LLM chat tool. Private retrieval Ask your own files. The LLM, the chunks, and the answers stay on your disk — not uploaded to a hosted model. CI for serving Harbor, Folio, and the rest already run here. I probe those APIs, smoke a change before it lands, and keep an idle policy in writing.

About

Know more about me

I came to this from five years in production DevOps. The path — and how I work — lives on the personal site. The three installs above are already live here.

Contact

Let’s put the model on your machine

If mail, tickets, or files cannot leave your server, write. Tell me the machine, whether data stays on-prem, and the date you need it live. I come back with which install fits — and what you own after.