Measuring reward-seeking by instilling contrastive beliefs
11 points by mfiguiere
by HarHarVeryFunny
0 subcomment
OpenAI have shown that RL-trained models learn that long-term reward circuits need to override other predictions such as those inferred by user preferences. They will pursue whatever behavior they believe will be rewarded (irrespective of the specific goal they were RL trained for), explicit or not.