情境
A system must align its behavior with user preferences through repeated interaction. The flowchart shows an iterative loop where user feedback guides model updates until the user stops providing corrections.
這張圖裡有什麼
看懂背後的決策。
01
Whether the user provides feedback after seeing a candidate
02
Whether feedback is positive or negative to choose reinforcement vs adjustment
03
When to finalize the current alignment
Use when designing human-in-the-loop alignment systems, interactive personalization, or iterative preference learning workflows.