Calibrated Decisions at the Human–Robot Boundary
Decision layer between robot planner and policy
Investigated the decision layer between a robot's planner and its policy (act, ask, or hand off to a person). Tested calibrated models against frozen rules and an oracle on four simulated bodies across 160+ experiments. Distilled a 421M-parameter on-device model from a cloud judge, maintaining 88.3% accuracy versus 87.9% for the cloud version at 90ms latency with calibration intact. Identified a failure mode in fleet learning where retraining on operator takeovers can make models confidently wrong outside corrected regions.