The scenario
A system must align its behavior with user preferences through repeated interaction. The flowchart shows an iterative loop where user feedback guides model updates until the user stops providing corrections.
What is in this drawing
Read the decisions behind it.
01
Whether the user provides feedback after seeing a candidate
02
Whether feedback is positive or negative to choose reinforcement vs adjustment
03
When to finalize the current alignment
Use when designing human-in-the-loop alignment systems, interactive personalization, or iterative preference learning workflows.
More like this