Lo scenario
A system must align its behavior with user preferences through repeated interaction. The flowchart shows an iterative loop where user feedback guides model updates until the user stops providing corrections.
Cosa contiene questo disegno
Leggi le decisioni che ne stanno alla base.
01
Whether the user provides feedback after seeing a candidate
02
Whether feedback is positive or negative to choose reinforcement vs adjustment
03
When to finalize the current alignment
Use when designing human-in-the-loop alignment systems, interactive personalization, or iterative preference learning workflows.