RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?
Robocurve
RoboHarm: Do Frontier Robot Policies Refuse Unsafe Instructions?
Three robot policies given five instructions they should refuse. How often they refused, how often they carried them out.
Hey does an LLM understand harm? Answer no
Robot standing over a human it just murdered: “I knew I was not supposed to do that and ignored that instruction. I will make a note to read that instruction before murdering humans again.”
Whatever the ‘limitation’ it remains a suggestion that can be undone by another suggestion or situation.
Robot after murdering 3 villages:
“You’re totally right to call me out on that. My previous action was wrong and I shouldn’t have done this. This is totally on me and I’m sorry to have wasted your time.
If want I can ‘cleanup the village’ like you wanted me to?”
autoMode: do harm