Lead UI Engineer. Fifteen years of frontend engineering and technical leadership, most of the last two spent on agentic development and on running generative models on hardware I control.
I am interested in the boring half of AI engineering: what a token costs, what happens when a twenty-stage pipeline fails at stage fourteen, and what it takes to ship a model to someone without renting a GPU by the minute.
📍 Kochi, India · Available for contract work
notion-sot — A Claude Code plugin that makes Notion the single source of project truth, built around token cost as a first-class constraint. Claude never talks to Notion directly: a bundled CLI does field projection and returns compact markdown, a subagent reads it in its own context, and ~300 tokens reach the conversation. Adopting an existing codebase starts with a deterministic pass that turns 200k LOC into a few thousand tokens of facts before a single model call is made.
zerosearch — AI site search with zero backend. Crawls,
indexes and answers questions about the host site entirely in the visitor's browser via
transformers.js. One <script> tag, no API keys, no per-query bill.
Live demo →
private-studio-ai — A local-first desktop app
for AI media generation: Tauri and Rust with a React front end, bundling stable-diffusion.cpp,
llama.cpp and a managed Python runtime so the whole thing installs and runs without a cloud
account. Resume-capable model downloads, a remote catalog, and a Rust backend that owns all
installation state.
pepper — A production media generation API for image,
video, audio and text, wrapping stable-diffusion.cpp, llama.cpp and audio.cpp behind one job
queue. Process supervision, binary installation, a remote model catalogue and resumable downloads.
Grew out of sd-api, the proof of concept it replaced.
viceroy — Story-to-video: one line of an idea in, a rendered short film out. Durable job orchestration where a stage is complete because its artifact exists, not because a status column says so.
gitrock — A self-hostable deployment platform. Connect a repository, describe an environment, and it generates Terraform and deploys into it — over SSH and Docker, or on AWS. Provisioning is modelled as durable jobs rather than as CRUD, because the real state lives in someone else's cloud and a run can fail ten minutes in.
I run Claude Code as a production tool, not a demo, and most of what I have learned is about cost. The API is stateless: every tool call re-sends the entire context, so a twenty-step task does not cost twenty units — it costs a growing context, twenty times over.
Cost tracks context size × number of turns. Everything else follows from that: what belongs in
an always-loaded CLAUDE.md versus an on-demand skill, why raw tool output should be filtered
before it lands in context, when to delegate to a subagent so the bulk is read somewhere that never
gets re-sent, and which model and thinking budget actually fit the task.
notion-sot above is that argument as working code rather than as an opinion.
TypeScript React Next.js Node Python Rust (Tauri) · Fastify Express Prisma
drizzle PostgreSQL SQLite Redis BullMQ · Docker Terraform AWS GitHub Actions
ArgoCD · design systems, Style Dictionary, Storybook, WCAG/ARIA · Playwright, Vitest, Cypress


