
MCP Developer & AI Systems Engineer
Ketan Shukla
I build the full stack around the Model Context Protocol: servers that expose tools, hosts that run the loop, gates that ask first, crews that delegate, and ledgers that refuse the bill.
Built in public
Five projects. One protocol.
A progressive stack of the Model Context Protocol, from the first server to a cost-aware host that refuses to overspend.
MCP Server — Tools, Resources, and Prompts
A complete, working Model Context Protocol server built with Next.js and deployed to Vercel, with a hand-written route handler and a live playground.
- 4 MCP tools, 1 resource, 1 prompt
- Live playground + Streamable HTTP endpoint
- Zod input validation with guard rails
- Mermaid diagrams and executable examples in the repo
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}MCP Host — The Agent Loop
The other half of the protocol: a host that reads plain English, picks tools across multiple MCP servers, runs them, and loops until it has an answer.
- Hand-written agent loop with explicit stop_reason handling
- MCP schema → Claude API translation layer
- Multi-server namespacing to avoid tool collisions
- Prompt caching that cuts input cost by ~47%
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}The Agent That Asks First
Same loop, plus a notebook, an "are you sure?", a report card, and a rewind button — so an agent can own a dangerous tool and a human can still sleep at night.
- Host-owned approval gate — never trusts server hints
- Postgres persistence for messages, pending calls, and traces
- Regression eval suite that deliberately regresses to stay honest
- Replay UI to watch any past run step by step
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}One Agent That Hires Help
A sub-agent is a tool that happens to think. This repo adds spawn_agent and measures when a crew is cheaper, better, or both compared to a lone agent.
- Local spawn_agent tool that reuses the same loop
- One approval queue for an entire tree of workers
- partial_results and paused_children for safe resume
- Token-cost comparison showing 0.56× price at scale
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}The Host That Owns The Wallet
Sampling turns an MCP server into a client that asks your host to think for it. This repo adds a spend gate and a ledger so your host says no before the bill arrives.
- Implements MCP sampling with stateless retry flow
- Cost ceiling that refuses before any model call
- Ledger rows for allowed and refused draws, tagged by server
- $0 regression suite replayed from stored traces
{
"jsonrpc": "2.0",
"id": 1,
"method": "tools/list"
}About
I build end-to-end systems and ship them in public.
My work sits at the intersection of AI infrastructure and production engineering. Every project in this series is a deployed, open-source system with live endpoints, regression suites, and measured tradeoffs — not slide decks.
I also ship products outside the AI space. I founded Metronagon Media, a book-cover and series-branding studio, and I am a published author of 22 books across 3 series. That same obsession with clarity, finish, and long-term craft shows up in my engineering work.
If you are hiring for AI platform, agent infrastructure, or developer tooling, I can take a system from idea to deployed endpoint — and document the path from prototype to production.
Ship fast
Next.js + Vercel from day one. Live URLs, not slide decks.
Build safe
Approval gates, cost ceilings, evals, and replay as first-class features.
Document the work
READMEs with diagrams, numbers, and honest tradeoffs that survive handoff.
Think end-to-end
Server, host, gate, crew, wallet — the whole protocol stack.
Stack
What I actually ship with
89 tools and practices across 9 areas, in the order they matter for the work I want. Everything here has been used in something deployed — nothing is aspirational.
AI & Agentic Systems
14
- LLM application development
- Model Context Protocol
- MCP servers, hosts & clients
- Tools, resources, prompts, sampling
- Agent loops & tool-use orchestration
- Sub-agent delegation
- Human-in-the-loop approval gates
- LLM evaluation & eval suites
- Trace replay
- Retrieval-augmented generation
- Embeddings & semantic search
- Cross-encoder reranking
- Prompt engineering
- Token & cost governance
AI Platforms & APIs
7
- Anthropic Claude API
- Tool use & sampling
- Prompt caching
- OpenAI GPT & GPT Image
- OpenAI embeddings
- Azure Neural TTS
- Windsurf · Cursor · Claude Code
Evaluated LangChain and LlamaIndex; agent loops built directly against provider SDKs.
Front End
10
- React 19
- Next.js 16 (App Router, Turbopack)
- TypeScript
- JavaScript
- Tailwind CSS v4
- MDX
- Server-side rendering
- Responsive design
- SEO
- Accessibility
Back End & Python
10
- Python
- Node.js
- REST APIs
- Pandas
- NumPy
- SQLAlchemy
- ETL pipeline design
- Great Expectations
- Cheerio
- Monorepos (pnpm, Turborepo)
ML & Retrieval
10
- PyTorch
- scikit-learn
- sentence-transformers
- HuggingFace models
- BM25 & TF-IDF
- Reciprocal rank fusion
- Cross-encoder rerankers
- K-means & MMR
- IR metrics (recall, MRR, nDCG)
- Model evaluation & threshold tuning
Applied to retrieval and routing rather than training. Every item here is used in a deployed repo with a measured result, not a tutorial.
Data & Storage
10
- PostgreSQL
- Supabase (Auth, Postgres, Realtime)
- pgvector
- HNSW vector indexing
- Row Level Security
- Google OAuth (PKCE)
- Oracle
- Sybase
- SQLite
- Schema design & query optimisation
Cloud & DevOps
11
- Vercel
- GitHub Actions
- CI/CD
- Git / GitHub
- Stripe API (embedded checkout)
- Inngest · Trigger.dev
- Durable background execution
- Release packaging & versioning
- Preview deployments
- Zero-downtime releases
- Rollback & recovery
Testing & Quality
8
- Vitest
- pytest
- unittest
- Mocked model & network clients
- Regression & eval suites
- Contract validation
- Root-cause debugging
- Production support
Enterprise & Legacy
9
- Microsoft .NET
- Visual C#
- Visual Basic
- BizTalk Server
- Pipelines, maps, orchestrations
- XML / XSD
- ESRI ArcGIS & ArcMap
- Client/server architecture
- Windows Server
Message orchestration across systems you do not control is the same problem an agent loop solves.