Large Language Models
169 views
Alignment Tax
Quick Definition
Performance cost of aligning AI with human preferences
Full Definition
The performance cost incurred when aligning AI models to follow human preferences and safety guidelines.
Examples
capability reduction after RLHF, safety vs helpfulness tradeoff
Related Terms
rlhf
ai-alignment
fine-tuning
More Large Language Models Terms
Jailbreaking
Circumventing AI safety restrictions through prompt techniques
Vector Database
Specialized database for high-dimensional embedding queries
Multi-Head Attention
Parallel attention operations concatenated for richer representations
BERT
Google's bidirectional transformer for context understanding
Flash Attention
Memory-efficient attention using GPU SRAM block computation
Prompt Caching
Storing key-value states for repeated prompt prefixes