Fine-tuning GPT-2 from human preferences
Researchers at OpenAI fine-tuned a large language model, GPT-2, using human feedback for various tasks. They found that the model learned to prioritize copying text from the input when summarizing, as this was preferred by external labelers. This approach required significant human effort, with 60k labels needed for summarization and only 5k for simpler tasks. The goal is to develop safety techniques that can effectively communicate with humans, a key step in understanding hu
Researchers at OpenAI fine-tuned a large language model, GPT-2, using human feedback for various tasks. They found that the model learned to prioritize copying text from the input when summarizing, as this was preferred by external labelers. This approach required significant human effort, with 60k labels needed for summarization and only 5k for simpler tasks. The goal is to develop safety techniques that can effectively communicate with humans, a key step in understanding human values.
---
Why it matters: This research matters because it highlights the challenges of aligning AI models with human preferences, particularly when those preferences are unclear or conflicting. Understanding how to fine-tune models using human feedback is crucial for developing trustworthy and effective AI systems that can communicate effectively with humans.
Source: https://openai.com/index/fine-tuning-gpt-2
This article was originally published at: https://openai.com/index/fine-tuning-gpt-2