Work with me
I help teams ship AI systems that survive contact with production — the same agent
architectures, local-LLM setups, and security trade-offs I write about here, applied to
your stack. If you found this site through a benchmark or a deep-dive, this is the
hands-on version.
14+ yrs Senior Staff engineer — AI/ML, RAG, agents, full-stack
4M+ Search impressions/quarter on this site’s AI engineering writing
220+ Published deep-dives, benchmarks & comparisons — all reproducible
Real systems Everything here is run on shipped, operated code — not demos
How I can help
Architecture & production-readiness review
A focused review of your AI agent, RAG, or LLM system before (or after) it hits production: failure modes, cost and latency budgets, eval coverage, fallback chains, and the security surface. You get a written report with prioritized fixes — the same standards I apply to systems I operate myself.
Best for: teams 2–6 weeks from shipping, or debugging a system that misbehaves in prod.
Local LLM & hardware strategy
Should you run open-weight models on your own hardware, and on what? Grounded in the benchmark data this site is known for — Apple Silicon vs NVIDIA vs AMD ROCm, quantization trade-offs, inference servers, and the real total cost against API pricing.
Best for: teams weighing local inference vs API spend, or planning a hardware purchase.
Advisory call
A single working session on a specific question — agent control flow, prompt-injection surface, eval design, model selection, cost blow-ups. Come with a concrete problem; leave with a concrete plan.
Best for: a decision you need a second senior opinion on this week.
Book an intro call
Prefer async? [email protected]
— include what you're building and where it hurts.