Combination of Cosine Similarity Method and Conditional Probability for Plagiarism Detection in the Thesis Documents Vector Space Model


  • Ristu Saptono Department of Informatics,Universitas Sebelas Maret, Surakarta, Indonesia.
  • Heri Prasetyo Department of Informatics,Universitas Sebelas Maret, Surakarta, Indonesia.
  • Ade Irawan Department of Informatics,Universitas Sebelas Maret, Surakarta, Indonesia.


Conditional Probability, Cosine Similarity, Plagiarism, Vector Space Model,


Plagiarism is one of negative impact derived from the internet growth. It can take place in various place, one of the examples is higher education environment. Plagiarism can cause many disadvantageous to other parties. So, there must be a detection system to avoid this kind of bad thing. In this proposed research, there will be made a plagiarism detection system by implementing Vector Space Model (VSM). Cosine Similarity used to make the rank of the paragraphs based on the formed angle from query vector and collection vector. The number of the taken words from the query paragraph will be derived from the calculation of the conditional probability value. After testing phase has been finished, there will be a conclusion that VSM can be implemented in the system. There are 10 testing paragraphs that compared with the collection paragraphs. The best result shows from threshold 0.3 for the conditional probability and 0.2 for cosine similarity with 54.28% for the average precision and 100% for the average recall.


Download data is not yet available.




How to Cite

Saptono, R., Prasetyo, H., & Irawan, A. (2018). Combination of Cosine Similarity Method and Conditional Probability for Plagiarism Detection in the Thesis Documents Vector Space Model. Journal of Telecommunication, Electronic and Computer Engineering (JTEC), 10(2-4), 139–143. Retrieved from

Most read articles by the same author(s)