O cenário
A system must align its behavior with user preferences through repeated interaction. The flowchart shows an iterative loop where user feedback guides model updates until the user stops providing corrections.
O que há neste desenho
Entenda as decisões por trás dele.
01
Whether the user provides feedback after seeing a candidate
02
Whether feedback is positive or negative to choose reinforcement vs adjustment
03
When to finalize the current alignment
Use when designing human-in-the-loop alignment systems, interactive personalization, or iterative preference learning workflows.