Large Language Models
21 views
Reward Model
Quick Definition
Model scoring LLM outputs based on human preferences
Full Definition
A model trained to score LLM outputs based on human preference data for RLHF training pipelines.
Examples
RLHF pipeline, preference ranking, reward shaping
Related Terms
rlhf
reward-modeling
ai-alignment
More Large Language Models Terms
Context Window
Maximum tokens a language model can process at once
Jailbreaking
Circumventing AI safety restrictions through prompt techniques
RAG
Enhancing LLMs by retrieving relevant external knowledge
Tree-of-Thought
Framework exploring multiple reasoning paths for optimal solutions
Alignment Tax
Performance cost of aligning AI with human preferences
Constitutional AI
Training AI to follow explicit principles for consistency