Staff AI Engineer | Making agents work
I build agentic AI systems that run in production. Before that, ten years shipping product at B2B SaaS companies: point of sale at Lightspeed, terabyte-scale inspection data and geospatial tooling at Zeitview, cap table software at Pulley. Getting an agent to work on the ten-thousandth call, inside a real business workflow, is the same discipline: retries, idempotency, partial failure, state that survives a restart, and a cost per call you can defend. WHAT I WORK ON Staff AI Engineer at Relevance AI, building arg.ai, an AI workspace where teams and agents share the same files, chat, and compute. Most of my work sits below the product: - Agent turns run for hours on Cloudflare Workflows and Durable Objects. I work on journal replay, retry backoff, hibernation, and the state resets a deploy causes mid-run. - Agents write and run real code in container sandboxes. I own their lifecycle, isolation, and how much damage one can do. - Tools reach the agent over MCP. I build and run the server, the tool schemas, the auth, the transports. - Chat streams over a WebSocket turn protocol with model routing, reasoning-effort control, and tool dispatch. I track time to first token. TOOLING SO AGENTS CAN DO THE WORK I also build what AI coding agents need to ship real changes and prove them: - Ephemeral cloud environments that boot a full production-shaped stack on demand, so agents build and verify without my laptop. - Browser automation so an agent exercises the real UI end to end. - CLI tools that drive the real agent over the real transport, so backend changes get checked against the running system. - Eval and trace pipelines: error clustering and regression triage. - Recorded demos and diffs that show a change working, so a reviewer can watch it. AGENTS AND LLM SYSTEMS MCP servers and clients: tool schemas, transports, auth · Context engineering: retrieval, compaction, memory, cache-aware prompt layout · Sandboxing and computer-use safety: isolation, approval gates, kill switches, prompt injection · Evals: golden sets, offline and online, LLM-as-judge, regression gates in CI · Observability: tracing, span-level cost and latency, drift alerting · Retrieval and memory: pgvector, Qdrant, hybrid search · LoRA/QLoRA, DPO THE REST OF THE STACK Cloudflare Workers, Durable Objects, Workflows, Containers, R2 · TypeScript/Node, Python, Rails · AWS, GCP, Kubernetes, Terraform · Postgres, MongoDB, Redis, PostGIS · Next.js, React, React Native Happy to talk with anyone building AI automations, agent infrastructure, sandboxes, or evals. tejbirwason@gmail.com