About
Software engineer building practical AI workflows — document pipelines, assistants, agents — with the production discipline that comes from working on payments.
Bio
Most LLM demos fall over the first week they meet real data: a PDF the extractor hasn't seen, a question the docs don't answer, an agent that retries and does the thing twice. The work I care about is the part after the demo — the evals, the confidence gates, the human sign-off, the idempotent side effects — that turns a promising prototype into something a business can actually run on.
That instinct comes from my day job. At Petpooja I work on offline payments infrastructure — card terminals, QR, settlement and reconciliation — for restaurants and retail stores across India. In payments, a retried request can't be allowed to move money twice, and “probably correct” isn't a state. I hold AI systems to the same bar.
I came to all of it through the web — TypeScript, React, Node — and I have preferred tools but no loyalties. The job is to find the solution that fits the problem, and most of the interesting decisions are about what not to automate.
Expertise
AI systems
- Structured extraction
- RAG & hybrid retrieval
- Tool-calling agents
- Evals & regression testing
- Guardrails & human-in-the-loop
- Prompt & context design
- Cost & latency budgeting
Payments & systems
- Offline payment rails (EDC, QR)
- Settlement & reconciliation
- Idempotent webhooks
- Payment orchestration
- Postgres performance
- Zero-downtime schema evolution
Stack
How I build with AI
AI is a tool in my editor and a component in my products, not a bullet on my résumé. Here's the honest version.
In production
I've shipped LLM-backed features: conversational assistants, retrieval over documents, structured extraction and summarization pipelines, and tool-calling agents. The interesting work is rarely the prompt — it's the evals, the fallbacks, and deciding what the model is never allowed to decide.
What a workflow needs before I'd call it done
- An eval set built from real cases — and a regression run before any prompt or model change.
- A confidence signal, and a human review path for everything below the threshold.
- Idempotent side effects, so a retry never sends the email or moves the money twice.
- A cost and latency budget per request, measured, not estimated.
- An honest failure mode: “I don't know” beats a fluent wrong answer.
Where it doesn't help
Anything that needs to know why the system is built the way it is. Data modelling. Naming. Deciding what not to build. It will confidently produce plausible code that quietly does the wrong thing — and catching that is still the job. So I read everything it writes.
What I actually use
- multi-file changes, refactors, and reading unfamiliar code fast
- in-editor work and line-level iteration
- design discussion, debugging, and research
Experience
Software Engineer
Petpooja- Offline payments infrastructure: EDC card terminals, static and dynamic QR, settlement and reconciliation.
- LLM-backed features in production — assistants, document retrieval, extraction pipelines, tool-calling agents.
- Integrations with lending, wallet and POS partners; idempotent webhooks and fleet-scale polling.
Outside work
Chess, cricket, and travel. I watch more chess streams than I'd like to admit — there's a lot to learn about the game from people thinking out loud.