{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,2,21]],"date-time":"2026-02-21T19:17:54Z","timestamp":1771701474098,"version":"3.50.1"},"reference-count":32,"publisher":"Wiley","issue":"13","license":[{"start":{"date-parts":[[2019,6,23]],"date-time":"2019-06-23T00:00:00Z","timestamp":1561248000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":["onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Concurrency and Computation"],"published-print":{"date-parts":[[2020,7,10]]},"abstract":"<jats:title>Summary<\/jats:title><jats:p>Plagiarism is becoming an increasingly serious problem in academic environment. In this paper, we deal with a specific kind of plagiarism: source code plagiarism. In this case, there is no software available for detecting plagiarism on a larger scale (hundreds of student submissions every year). We propose algorithms for source code parsing and processing as a part of a complex system for plagiarism detection. A source code vectorization using characteristic vectors is a vital piece of the whole process, and k\u2010means algorithm helps with the classification and clustering of vectors. Student assignments are submitted regularly, and any plagiarism detection system needs to handle them as they come. For this reason, we propose a modified incremental k\u2010means algorithm and a method for determining the number of clusters. We also consider methods for vector search among clusters and suggest the use of conditional entropy to select the important vector elements used in the search algorithm. Our results show how the proposed algorithms and methods work on real student submissions.<\/jats:p>","DOI":"10.1002\/cpe.5416","type":"journal-article","created":{"date-parts":[[2019,6,23]],"date-time":"2019-06-23T23:34:37Z","timestamp":1561332877000},"update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":6,"title":["Searching source code fragments using incremental clustering"],"prefix":"10.1002","volume":"32","author":[{"given":"Michal","family":"\u010eura\u010d\u00edk","sequence":"first","affiliation":[{"name":"Faculty of Management Science and Informatics University of Zilina  Zilina Slovakia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Emil","family":"Kr\u0161\u00e1k","sequence":"additional","affiliation":[{"name":"Faculty of Management Science and Informatics University of Zilina  Zilina Slovakia"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-8747-9194","authenticated-orcid":false,"given":"Patrik","family":"Hrk\u00fat","sequence":"additional","affiliation":[{"name":"Faculty of Management Science and Informatics University of Zilina  Zilina Slovakia"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2019,6,23]]},"reference":[{"key":"e_1_2_11_2_1","doi-asserted-by":"publisher","DOI":"10.1109\/TE.2010.2046664"},{"key":"e_1_2_11_3_1","doi-asserted-by":"crossref","unstructured":"\u010eura\u010d\u00edkM Kr\u0161\u00e1kE Hrk\u00fatP.Issues with the detection of plagiarism in programming courses on a larger scale. Paper presented at: 2018 16th International Conference on Emerging eLearning Technologies and Applications (ICETA);2018;Star\u00fd Smokovec Slovakia.","DOI":"10.1109\/ICETA.2018.8572260"},{"key":"e_1_2_11_4_1","doi-asserted-by":"publisher","DOI":"10.1080\/07294360.2016.1161602"},{"key":"e_1_2_11_5_1","unstructured":"KravjarJ.SK antiplag is bearing fruit. In: Proceedings of the International Conference on Plagiarism Across Europe and Beyond;2017;Brno Czech Republic."},{"key":"e_1_2_11_6_1","doi-asserted-by":"crossref","unstructured":"JiangL MisherghiG SuZ GlonduS.Deckard: scalable and accurate tree\u2010based detection of code clones. In: Proceedings of the 29th International Conference on Software Engineering;2007;Minneapolis MN.","DOI":"10.1109\/ICSE.2007.30"},{"key":"e_1_2_11_7_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.infsof.2006.10.017"},{"key":"e_1_2_11_8_1","unstructured":"MaleticJI ValluriN.Automatic software clustering via latent semantic analysis. In: Proceedings of the 14th IEEE International Conference on Automated Software Engineering;1999;Cocoa Beach FL."},{"key":"e_1_2_11_9_1","unstructured":"DovalD MancoridisS MitchellBS.Automatic clustering of software systems using a genetic algorithm. In: Proceedings of the 9th International Workshop Software Technology and Engineering Practice;1999;Pittsburgh PA."},{"key":"e_1_2_11_10_1","unstructured":"RousidisD TjortjisC.Clustering data retrieved from Java source code to support software maintenance: a case study. Paper presented at: 9th European Conference on Software Maintenance and Reengineering;2005;Manchester UK."},{"key":"e_1_2_11_11_1","doi-asserted-by":"crossref","unstructured":"LukinsSK KraftNA EtzkornLA.Source code retrieval for bug localization using latent dirichlet allocation. Paper presented at: 2008 15th Working Conference on Reverse Engineering;2008;Antwerp Belgium.","DOI":"10.1109\/WCRE.2008.33"},{"key":"e_1_2_11_12_1","doi-asserted-by":"crossref","unstructured":"MancoridisS MitchellBS ChenY GansnerER.Bunch: a clustering tool for the recovery and maintenance of software system structures. In: Proceedings of the IEEE International Conference on Software Maintenance (ICSM);1999;Oxford UK.","DOI":"10.1109\/ICSM.1999.792498"},{"key":"e_1_2_11_13_1","unstructured":"MancoridisS MitchellBS RorresC ChenY GansnerER.Using automatic clustering to produce high\u2010level system organizations of source code. In: Proceedings of the 6th International Workshop on Program Comprehension (IWPC);1998;Ischia Italy."},{"key":"e_1_2_11_14_1","unstructured":"\u010eura\u010d\u00edkM Kr\u0161\u00e1kE Hrk\u00fatP.Using concepts of text based plagiarism detection in source code plagiarism analysis. In: Proceedings of the International Conference on Plagiarism Across Europe and Beyond;2017;Brno Czech Republic."},{"key":"e_1_2_11_15_1","doi-asserted-by":"publisher","DOI":"10.1093\/comjnl\/bxh119"},{"key":"e_1_2_11_16_1","doi-asserted-by":"publisher","DOI":"10.1504\/IJBIDM.2008.020514"},{"key":"e_1_2_11_17_1","unstructured":"EsterM KriegelH\u2010P SanderJ XuX.A density\u2010based algorithm for discovering clusters in large spatial databases with noise. In:SimoudisE HanJ FayyadUM eds. Proceedings of the Second International Conference on Knowledge Discovery and Data Mining;1996;Portland OR."},{"key":"e_1_2_11_18_1","doi-asserted-by":"publisher","DOI":"10.1109\/TC.2011.223"},{"key":"e_1_2_11_19_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.proeng.2017.06.024"},{"key":"e_1_2_11_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-319-95522-3_6"},{"key":"e_1_2_11_21_1","doi-asserted-by":"crossref","unstructured":"ZhaoJ XiaK FuY CuiB.An AST\u2010based code plagiarism detection algorithm. Paper presented at: 2015 10th International Conference on Broadband and Wireless Computing Communication and Applications (BWCCA);2015;Krak\u00f3w Poland.","DOI":"10.1109\/BWCCA.2015.52"},{"key":"e_1_2_11_22_1","unstructured":"HuangA.Similarity measures for text document clustering. In: Proceedings of the 6th New Zealand Computer Science Research Student Conference (NZCSRSC);2008;Christchurch New Zealand."},{"key":"e_1_2_11_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-0-12-385022-5.00015-4"},{"key":"e_1_2_11_24_1","doi-asserted-by":"crossref","unstructured":"SculleyD.Web\u2010scale k\u2010means clustering. In: Proceedings of the 19th International Conference on World Wide Web;2010;Raleigh NC.","DOI":"10.1145\/1772690.1772862"},{"key":"e_1_2_11_25_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.chemolab.2012.11.006"},{"key":"e_1_2_11_26_1","first-page":"451","volume-title":"Readings in Multimedia Computing and Networking","author":"Berchtold S","year":"2001"},{"key":"e_1_2_11_27_1","unstructured":"LinH\u2010Y.A compact index structure with high data retrieval efficiency. Paper presented at: 2008 International Conference on Service Systems and Service Management;2008;Melbourne Australia."},{"key":"e_1_2_11_28_1","doi-asserted-by":"crossref","unstructured":"FeldmanD SchmidtM SohlerC.Turning big data into tiny data: constant\u2010size coresets for k\u2010means PCA and projective clustering. In: Proceedings of the 24th Annual ACM\u2010SIAM Symposium on Discrete Algorithms;2013;New Orleans LA.","DOI":"10.1137\/1.9781611973105.103"},{"key":"e_1_2_11_29_1","doi-asserted-by":"publisher","DOI":"10.1109\/21.97458"},{"key":"e_1_2_11_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/18.119732"},{"issue":"17","key":"e_1_2_11_31_1","first-page":"18","article-title":"Source code plagiarism Detection \u2018SCPDet\u2018: a review","volume":"105","author":"Gondaliya TP","year":"2014","journal-title":"Int J Comput Appl"},{"key":"e_1_2_11_32_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.jss.2014.10.013"},{"key":"e_1_2_11_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patrec.2004.04.007"}],"container-title":["Concurrency and Computation: Practice and Experience"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fcpe.5416","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.5416","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/cpe.5416","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/cpe.5416","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,9,6]],"date-time":"2023-09-06T05:53:18Z","timestamp":1693979598000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/10.1002\/cpe.5416"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2019,6,23]]},"references-count":32,"journal-issue":{"issue":"13","published-print":{"date-parts":[[2020,7,10]]}},"alternative-id":["10.1002\/cpe.5416"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/cpe.5416","archive":["Portico"],"relation":{},"ISSN":["1532-0626","1532-0634"],"issn-type":[{"value":"1532-0626","type":"print"},{"value":"1532-0634","type":"electronic"}],"subject":[],"published":{"date-parts":[[2019,6,23]]},"assertion":[{"value":"2018-11-19","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-05-21","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2019-06-23","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}],"article-number":"e5416"}}