
Why start here: Almeida asks what the model is actually being trained to want. Human preference produces a good assistant, while correctness or calibrated decisions imply different systems. Bigio comes next because he turns Almeida's three objectives into concrete choices among SFT, DPO and RFT.








