Skip to main navigation Skip to search Skip to main content

ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty

  • Qing ZONG
  • , Zhaowei WANG
  • , Xiyu REN
  • , Tianshi ZHENG
  • , Yangqiu SONG

Research output: Chapter in Book/Conference Proceeding/ReportConference Paper published in a bookpeer-review

Abstract

The rapid development of LLMs has sparked extensive research into their factual knowledge. Current works find that LLMs fall short on questions around low-frequency entities. However, such proofs are unreliable since the questions can differ not only in entity frequency but also in difficulty themselves. So we introduce COMPARISONQA benchmark, containing 283K abstract questions, each instantiated by a pair of high-frequency and low-frequency entities. It ensures a controllable comparison to study the role of knowledge frequency in the performance of LLMs. Because the difference between such a pair is only the entity with different frequencies. In addition, we use both correctness and uncertainty to develop a two-round method to evaluate LLMs' knowledge robustness. It aims to avoid possible semantic shortcuts which is a serious problem of current QA study. Experiments reveal that LLMs, including GPT-4o, exhibit particularly low robustness regarding low-frequency knowledge. Besides, we find that uncertainty can be used to effectively identify high-quality and shortcut-free questions while maintaining the data size. Based on this, we propose an automatic method to select such questions to form a subset called COMPARISONQA-Hard, containing only hard low-frequency questions.
Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics: ACL 2025
EditorsWanxiang Che, Joyce Nabende, Ekaterina Shutova, Mohammad Taher Pilehvar
Place of PublicationVienna, Austria
PublisherAssociation for Computational Linguistics (ACL)
Pages4101–4117
Number of pages17
ISBN (Electronic)9798891762565
DOIs
Publication statusPublished - Jul 2025
EventThe 63rd Annual Meeting of the Association for Computational Linguistics - Vienna, Austria
Duration: 27 Jul 20251 Aug 2025

Publication series

NameProceedings of the Annual Meeting of the Association for Computational Linguistics
PublisherAssociation for Computational Linguistics
ISSN (Electronic)0736-587X

Conference

ConferenceThe 63rd Annual Meeting of the Association for Computational Linguistics
Country/TerritoryAustria
CityVienna
Period27/07/251/08/25

Bibliographical note

Publisher Copyright:
© 2025 Association for Computational Linguistics.

Fingerprint

Dive into the research topics of 'ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty'. Together they form a unique fingerprint.

Cite this