TopicsRLHF & AlignmentSketch the reward overoptimization…Short answer·Hard·Asked atReplicateRobust IntelligenceWorkdayNextNext questionWhat is RLHF, and why is it used after pretraining?MCQ·EasyPrevSkipNext