Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning
Researchers have proposed a new method to evaluate the quality of autoformalization tasks, which inv...
Researchers have proposed a new method to evaluate the quality of autoformalization tasks, which inv...
Researchers have developed a framework called Cognitive Pairwise Comparison Classification Model Sel...
Researchers Nicolas Boizard and colleagues investigate whether distilling reasoning traces from stro...
Researchers have proposed a new framework called StruProKGR for efficiently reasoning over sparse kn...
Researchers have developed MedRAGChecker, a tool to verify the accuracy of claim...
Researchers have developed SlidesGen-Bench, a benchmark for evaluating automated...
Researchers have developed a method to improve the accuracy of AI-generated diag...
Researchers propose a three-layer framework for building trustworthy AI systems ...