I build RL environments that stump frontier models | Software Engineer @ Polymath | AI Evals | Contributor to karpathy/autoresearch, PyTorch, LangChain & NVIDIA NemoClaw
TL;DR: I enjoy finding the limits of frontier models.
I build RL environments and scenarios that stump frontier models. Usually that means long-horizon coding tasks where agents have to preserve constraints, reason across files, recover from broken assumptions, and handle messy real-world systems. I do this at Polymath.
The same pattern runs through my open-source work. A few hours after Andrej Karpathy published autoresearch, I had a PR merged into it: a self-healing loop that lets the agent parse its own stack traces and recover from crashes automatically.
Across Google Workspace CLI, NVIDIA NemoClaw, LangChain, LangGraph, and PyTorch, I've worked on the same class of problem : hidden failure modes in systems people are starting to trust: agent crashes, silent tool-calling failures, credential exposure, shell injection, and compiler correctness bugs.
I also like helping people understand how AI can make the things they want to do 10x easier
Interested in frontier AI evals, coding-agent reliability, agent red-teaming, AI infrastructure, and AI education.