{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,10,17]],"date-time":"2025-10-17T13:38:35Z","timestamp":1760708315562},"reference-count":36,"publisher":"Wiley","issue":"10","license":[{"start":{"date-parts":[[2011,7,15]],"date-time":"2011-07-15T00:00:00Z","timestamp":1310688000000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":["J. Am. Soc. Inf. Sci."],"published-print":{"date-parts":[[2011,10]]},"abstract":"<jats:title>Abstract<\/jats:title><jats:p>The progressive increase of information content has recently made it necessary to create a system for automatic classification of documents. In this article, a system is presented for the categorization of multiclass Farsi documents that requires fewer training examples and can help to compensate the shortcoming of the standard training dataset. The new idea proposed in the present article is based on extending the feature vector by adding some words extracted from a thesaurus and then filtering the new feature vector by applying secondary feature selection to discard inappropriate features. In fact, a phase of secondary feature selection is applied to choose more appropriate features among the features added from a thesaurus to enhance the effect of using a thesaurus on the efficiency of the classifier. To evaluate the proposed system, a corpus is gathered from the Farsi Wikipedia website and some articles in the <jats:italic>Hamshahri<\/jats:italic> newspaper, the <jats:italic>Roshd<\/jats:italic> periodical, and the <jats:italic>Soroush<\/jats:italic> magazine. In addition to studying the role of a thesaurus and applying secondary feature selection, the effect of a various number of categories, size of the training dataset, and average number of words in the test data also are examined. As the results indicate, classification efficiency improves by applying this approach, especially when available data is not sufficient for some text categories.<\/jats:p>","DOI":"10.1002\/asi.21592","type":"journal-article","created":{"date-parts":[[2011,7,15]],"date-time":"2011-07-15T14:39:19Z","timestamp":1310740759000},"page":"2055-2066","source":"Crossref","is-referenced-by-count":3,"title":["Improving Farsi multiclass text classification using a thesaurus and two\u2010stage feature selection"],"prefix":"10.1002","volume":"62","author":[{"given":"Nooshin","family":"Maghsoodi","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mohammad Mehdi","family":"Homayounpour","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"311","published-online":{"date-parts":[[2011,7,15]]},"reference":[{"key":"e_1_2_12_2_1","first-page":"821","article-title":"Theoretical foundations of the potential function method in pattern recognition learning","volume":"25","author":"Aizerman A.","year":"1964","journal-title":"Automation and Remote Control"},{"key":"e_1_2_12_3_1","first-page":"245","volume-title":"Proceedings of the Second Workshop on Persian Language and Computer","author":"Arabsorkhi M.","year":"2006"},{"key":"e_1_2_12_4_1","first-page":"383","volume-title":"Proceedings of the 13th International Computer Conference of Computer Society of Iran","author":"Basiri M.E.","year":"2008"},{"key":"e_1_2_12_5_1","volume-title":"100 million word Farsi corpus (Tech. Rep.)","author":"Bijankhan M.","year":"2008"},{"key":"e_1_2_12_6_1","first-page":"385","volume-title":"Proceedings of the 4th International Conference on Data Mining (DMIN'2008)","author":"Bina B.","year":"2008"},{"key":"e_1_2_12_7_1","first-page":"149","volume-title":"Workshop on Text\u2010Based Information Retrieval at the 27th German Conference on Artificial Intelligence (TIR\u201004)","author":"Bloehdorn S.","year":"2004"},{"key":"e_1_2_12_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/S0004-3702(97)00063-5"},{"key":"e_1_2_12_9_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijar.2008.10.006"},{"key":"e_1_2_12_10_1","unstructured":"Dehdari J.(2008).Perstem Stemmer (Version 0.9.7) [Computer Program]. Retrieved fromhttps:\/\/2.zoppoz.workers.dev:443\/http\/www.ling.ohio\u2010state.edu\/\u223cjonsafari\/persian_nlp.html"},{"key":"e_1_2_12_11_1","doi-asserted-by":"publisher","DOI":"10.1002\/asi.10409"},{"key":"e_1_2_12_12_1","doi-asserted-by":"publisher","DOI":"10.1145\/288627.288651"},{"key":"e_1_2_12_13_1","first-page":"210","volume-title":"Proceedings of the First Workshop on Farsi Language and Computer","author":"Fararuy J.","year":"2004"},{"key":"e_1_2_12_14_1","first-page":"1048","volume-title":"Proceedings of the 19th International Joint Conference on Artificial Intelligence (IJCIA\u201005)","author":"Gabrilovich E.","year":"2005"},{"key":"e_1_2_12_15_1","doi-asserted-by":"publisher","DOI":"10.1162\/153244303322753616"},{"key":"e_1_2_12_16_1","first-page":"61","volume-title":"Proceedings of the Semantic Web Workshop at the 26th International Conference on Research and Development in Information Retrieval (ACM SIGIR '03)","author":"Hotho A.","year":"2003"},{"key":"e_1_2_12_17_1","first-page":"1023","article-title":"A theoretic and empirical research of cluster indexing for Mandarin Chinese full text document","volume":"24","author":"Huang Y.L.","year":"1998","journal-title":"Journal of Library and Information Science"},{"key":"e_1_2_12_18_1","first-page":"656","volume-title":"Proceedings of the 23rd IASTED International Multi\u2010Conference on Artificial Intelligence and Applications","author":"Ino Y.","year":"2005"},{"key":"e_1_2_12_19_1","first-page":"137","volume-title":"Proceedings of the 10th European Conference on Machine Learning (ECML)","author":"Joachims T.","year":"1998"},{"key":"e_1_2_12_20_1","doi-asserted-by":"publisher","DOI":"10.1016\/B978-1-55860-335-6.50023-4"},{"key":"e_1_2_12_21_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2004.08.006"},{"key":"e_1_2_12_22_1","first-page":"81","volume-title":"Proceedings of the Third Annual Symposium on Document Analysis and Information Retrieval (SDAIR\u201094)","author":"Lewis D.D.","year":"1994"},{"key":"e_1_2_12_23_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2005.10.006"},{"key":"e_1_2_12_24_1","first-page":"41","volume-title":"Proceedings of the Workshop on Learning for Text Categorization (AAAI'98)","author":"McCallum A.","year":"1998"},{"key":"e_1_2_12_25_1","doi-asserted-by":"publisher","DOI":"10.1109\/NLPKE.2009.5313761"},{"key":"e_1_2_12_26_1","unstructured":"Roget's Thesaurus. (n.d.). Retrieved from:https:\/\/2.zoppoz.workers.dev:443\/http\/www.rain.org\/\u223ckarpeles\/rogfrm.html"},{"key":"e_1_2_12_27_1","doi-asserted-by":"publisher","DOI":"10.1145\/361219.361220"},{"key":"e_1_2_12_28_1","doi-asserted-by":"publisher","DOI":"10.1145\/215206.215365"},{"key":"e_1_2_12_29_1","doi-asserted-by":"publisher","DOI":"10.1145\/505282.505283"},{"key":"e_1_2_12_30_1","doi-asserted-by":"publisher","DOI":"10.1109\/IFCSTA.2009.167"},{"key":"e_1_2_12_31_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4757-2440-0"},{"key":"e_1_2_12_32_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10115-008-0152-4"},{"key":"e_1_2_12_33_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ipm.2006.09.011"},{"key":"e_1_2_12_34_1","first-page":"317","volume-title":"Proceedings of the Fourth annual Symposium on Document Analysis and Information Retrieval (SDAIR\u201095)","author":"Wiener E.","year":"1995"},{"key":"e_1_2_12_35_1","doi-asserted-by":"publisher","DOI":"10.1007\/978-1-4471-2099-5_2"},{"key":"e_1_2_12_36_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1009982220290"},{"key":"e_1_2_12_37_1","first-page":"412","volume-title":"Proceedings of 14th International Conference on Machine Learning (ICML\u201097)","author":"Yang Y.","year":"1997"}],"container-title":["Journal of the American Society for Information Science and Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/api.wiley.com\/onlinelibrary\/tdm\/v1\/articles\/10.1002%2Fasi.21592","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/asi.21592","content-type":"application\/pdf","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/full-xml\/10.1002\/asi.21592","content-type":"application\/xml","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/pdf\/10.1002\/asi.21592","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2023,8,24]],"date-time":"2023-08-24T07:49:46Z","timestamp":1692863386000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/onlinelibrary.wiley.com\/doi\/10.1002\/asi.21592"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2011,7,15]]},"references-count":36,"journal-issue":{"issue":"10","published-print":{"date-parts":[[2011,10]]}},"alternative-id":["10.1002\/asi.21592"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/asi.21592","archive":["Portico"],"relation":{},"ISSN":["1532-2882","1532-2890"],"issn-type":[{"value":"1532-2882","type":"print"},{"value":"1532-2890","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,7,15]]}}}