Large Language Models
33 views
Alignment Tax
Quick Definition
Performance cost of aligning AI with human preferences
Full Definition
The performance cost incurred when aligning AI models to follow human preferences and safety guidelines.
Examples
capability reduction after RLHF, safety vs helpfulness tradeoff
Related Terms
rlhf
ai-alignment
fine-tuning
More Large Language Models Terms
RoPE
Position encoding using rotation matrices for relative positions
Tree-of-Thought
Framework exploring multiple reasoning paths for optimal solutions
SentencePiece
Language-agnostic tokenizer operating on raw text bytes
Grounded Generation
Generating outputs faithful to provided source documents
Guardrails LLM
Frameworks for monitoring and controlling LLM I/O
BPE
Subword tokenization algorithm merging frequent character pairs