AI

Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study

A recent study benchmarked 14 large language models (LLMs) on their accuracy in cosmetic chemistry and skin health. The results showed poor performance overall, with most LLMs struggling with quantitative reasoning and structural identification tasks. While they could provide general skincare advice, their responses lacked technical depth, making them unreliable sources of information for consumers. The study suggests that LLMs are not yet suitable for public use in this area
A recent study benchmarked 14 large language models (LLMs) on their accuracy in cosmetic chemistry and skin health. The results showed poor performance overall, with most LLMs struggling with quantitative reasoning and structural identification tasks. While they could provide general skincare advice, their responses lacked technical depth, making them unreliable sources of information for consumers. The study suggests that LLMs are not yet suitable for public use in this area without further improvements to their training data and algorithmic reasoning. --- Why it matters: This matters because AI chatbots are increasingly being used by consumers for skincare advice, but the accuracy of these models is uncertain. Engineers and researchers working on AI need to understand the limitations of current LLMs in order to improve them and make them safe for public use. Source: https://arxiv.org/abs/2608.14631

This article was originally published at: https://arxiv.org/abs/2608.14631