AI Alignment
The field focused on ensuring AI systems behave in line with human values and intended goals.
In practice it covers techniques like RLHF, red-teaming, guardrails, and refusal training that keep models helpful, honest, and harmless. It becomes more critical as autonomous agents take real-world actions.