Artificial Intelligence
25 views
Reward Modeling
Quick Definition
Training models to predict human preferences for RLHF
Full Definition
Training a model to predict human preferences and provide reward signals for reinforcement learning from human feedback.
Examples
RLHF training pipeline, preference learning, alignment
Related Terms
rlhf
ai-alignment
reinforcement-learning
More Artificial Intelligence Terms
Few-Shot Learning
Learning to perform tasks from a small number of examples
AI Chip
Specialized hardware for accelerating AI computation
Model Distillation
Creating smaller models that replicate larger model behavior
Zero-Shot Learning
Performing tasks on unseen classes without training examples
AI Benchmark
Standardized tests measuring AI model performance
Generative AI
AI systems that create new content