Learning from human preferences
Researchers at OpenAI have developed an algorithm that infers human preferences from simple feedback. By comparing two proposed behaviors and identifying the preferred one, the algorithm aims to reduce the need for humans to write complex goal functions. This could lead to safer AI systems by minimizing the risk of undesirable behavior caused by small errors in goal specification.
Researchers at OpenAI have developed an algorithm that infers human preferences from simple feedback. By comparing two proposed behaviors and identifying the preferred one, the algorithm aims to reduce the need for humans to write complex goal functions. This could lead to safer AI systems by minimizing the risk of undesirable behavior caused by small errors in goal specification.
---
Why it matters: This matters because it addresses a critical challenge in building safe AI: specifying goals that accurately reflect human preferences without introducing unintended consequences. By developing algorithms that can infer these preferences, researchers can create more robust and reliable AI systems.
Source: https://openai.com/index/learning-from-human-preferences
This article was originally published at: https://openai.com/index/learning-from-human-preferences