シナリオ
A system must align its behavior with user preferences through repeated interaction. The flowchart shows an iterative loop where user feedback guides model updates until the user stops providing corrections.
この図に含まれるもの
図の背後にある意思決定を読み解く。
01
Whether the user provides feedback after seeing a candidate
02
Whether feedback is positive or negative to choose reinforcement vs adjustment
03
When to finalize the current alignment
Use when designing human-in-the-loop alignment systems, interactive personalization, or iterative preference learning workflows.