Evolution strategies as a scalable alternative to reinforcement learning
Evolution strategies, a long-established optimization technique, has been found to perform similarly to reinforcement learning methods in various test...
Evolution strategies, a long-established optimization technique, has been found to perform similarly to reinforcement learning methods in various test...
Researchers at OpenAI have developed a method for training AI models to imitate complex tasks with just one example. This approach, called One-Shot Im...
OpenAI has launched Distill, a journal focused on effectively communicating machine learning research. The publication aims to present novel and exist...
Researchers at OpenAI have developed a system where artificial agents can create and use their own language. This is achieved through a process called...
Researchers at OpenAI have found that multi-agent populations can develop a form of compositional language, where agents learn to combine words and sy...
OpenAI researchers have developed a new type of model called temporal segment models, which can predict the future behavior of complex systems. These ...
OpenAI has developed a new technique for training AI models called third-person imitation learning. This method involves having the model learn from d...
Machine learning models can be tricked into making mistakes using 'adversarial examples', which are intentionally designed inputs that exploit the mod...
Researchers at OpenAI have explored the vulnerability of neural network policies to adversarial attacks. These attacks involve manipulating inputs to ...
OpenAI has grown to a team of 45 people, working together to advance artificial intelligence capabilities. They are exploring novel ideas, developing ...
Researchers at OpenAI have introduced a modified version of the PixelCNN model, called PixelCNN++. The new model incorporates discretized logistic mix...
Reinforcement learning algorithms can fail due to a common issue known as 'misspecified reward functions'. This occurs when the reward function, which...