Value alignment addresses broader expectations that may not fit into a single task objective. These can include avoiding harm, respecting people, acting honestly, and recognizing that different contexts may require different judgments.
Human values are varied, sometimes conflicting, and difficult to express completely in training data or rules. Value alignment therefore requires choices about whose values apply, how disagreements are handled, and how behavior is evaluated in unfamiliar situations.
ELI5
Value alignment means teaching an AI helper to care about the important principles behind a task, not only the score it can earn. The helper should understand that some ways of winning are not acceptable.
For example, a study helper should not invent a source just to finish an answer faster. Honesty remains important even when the immediate goal is to produce a complete report.
