Extension of Language Model to Solve Inconsistency, Incompleteness, and Short Query in the Collection of Cultural Heritage

Authors

  • K. L. Tan Department of Computer Science, Faculty of Art, Computing and Creative Industry,Sultan Idris Education University, Tanjong Malim, Perak Darul Ridzuan 35900, Malaysia.
  • C. K Lim Department of Computer Science, Faculty of Art, Computing and Creative Industry,Sultan Idris Education University, Tanjong Malim, Perak Darul Ridzuan 35900, Malaysia.

Keywords:

Cultural heritage, Information retrieval, Language Model,

Abstract

With the explosive growth of online information such as email messages, news articles, and scientific literature, many institutions and museums are converting their cultural collections from physical data to digital format. However, this conversion results in the issues of inconsistency and incompleteness. Besides, the usage of inaccurate keywords also results in short query problem. Most of the time, the inconsistency and incompleteness are caused by the aggregation fault in annotating a document itself while the short query problem is caused by naive user who has prior knowledge and experience in cultural heritage domain. In this paper, we presented an approach to solve the problem of inconsistency, incompleteness and short query by incorporating the Term Similarity Matrix into the Language Model. Our approach is tested on the Cultural Heritage in CLEF (CHiC) collection, which consists of short queries and documents. The results show that the proposed approach is effective and has improved the accuracy in retrieval time.

Downloads

Published

2017-09-15

How to Cite

Tan, K. L., & Lim, C. K. (2017). Extension of Language Model to Solve Inconsistency, Incompleteness, and Short Query in the Collection of Cultural Heritage. Journal of Telecommunication, Electronic and Computer Engineering (JTEC), 9(2-12), 119–123. Retrieved from https://jtec.utem.edu.my/jtec/article/view/2780