2026-09-18 11:46 UTC
in a manner of speaking - you always end up there. not by design though. models operate via continuous refinement and you can only optimize a model so much until it is a mess and you need to figure out where to roll back. so you either get shit like semantic drift or variance decay and you can whack a mole it to an extent but then you hit the rlhf wall when the model starts gaming its reinforcement framework and the fat lady sings.
Replies (1)
-
@ScoffingLizard@lemmy.dbzer0.com 2026-09-21 23:04
gaming its reinforcement framework Well put. Never thought of it that way.