PromptHub

Alignment Tax

Quick Definition

Performance cost of aligning AI with human preferences

Full Definition

The performance cost incurred when aligning AI models to follow human preferences and safety guidelines.

Examples

capability reduction after RLHF, safety vs helpfulness tradeoff

Related Terms

rlhf ai-alignment fine-tuning

More Large Language Models Terms