Artificial Intelligence
143 views
Reward Modeling
Quick Definition
Training models to predict human preferences for RLHF
Full Definition
Training a model to predict human preferences and provide reward signals for reinforcement learning from human feedback.
Examples
RLHF training pipeline, preference learning, alignment
Related Terms
rlhf
ai-alignment
reinforcement-learning
More Artificial Intelligence Terms
OCR
Technology converting images of text into machine-readable text
Pruning
Removing redundant neural network connections to reduce size
Feature Importance
Quantifying each input features contribution to predictions
Superintelligence
Hypothetical AI surpassing all human cognitive abilities
Natural Language Processing
AI field enabling computers to understand human language
Meta-Learning
Designing models that learn how to learn new tasks