{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T05:05:57Z","timestamp":1750309557645,"version":"3.41.0"},"reference-count":46,"publisher":"Association for Computing Machinery (ACM)","issue":"3","license":[{"start":{"date-parts":[[2025,3,22]],"date-time":"2025-03-22T00:00:00Z","timestamp":1742601600000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62376145, U24A20323, 62376087, and 62406184"],"award-info":[{"award-number":["62376145, U24A20323, 62376087, and 62406184"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Science and Technology Innovation Talent Team of Shanxi Province","award":["202204051002016"],"award-info":[{"award-number":["202204051002016"]}]},{"name":"Key Technologies Program of Taihang Laboratory in Shanxi Province","award":["THYFJSZX-24010700"],"award-info":[{"award-number":["THYFJSZX-24010700"]}]}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Trans. Knowl. Discov. Data"],"published-print":{"date-parts":[[2025,4,30]]},"abstract":"<jats:p>\n            Clustering is a fundamental technique widely used for exploring the inherent data structure. Many studies indicate that an appropriate feature representation can effectively improve clustering performance. However, the existing feature representation methods are based on correlation to select or extract features, which makes it hard to deal with spurious correlations. The spurious correlations mislead the correlation-based methods to consider features that have no causal relationship as being correlative, which limits the clustering performance and feature interpretability. To tackle this issue, inspired by causal learning, we propose a new joint optimization\n            <jats:bold>Clustering Method with Causal Feature Embedding (CM-CaFE)<\/jats:bold>\n            , which utilizes the causality of features to learn more discriminative representation for clustering. Specifically, to eliminate spurious correlations among features, we first employ any state-of-the-art Markov blanket learning method to learn an undirected causal graph. Next, we extract the maximal fully connected causal subgraphs from the learned undirected causal graph and propose an approach to merge them to generate the causal matrix. Based on the causal matrix, we present an objective function that consists of a clustering loss term and a causal matrix fitting term to learn a causal transformation matrix. The causal transformation matrix is utilized to map the original data into a new space for clustering. Finally, we comprehensively compare the proposed method with some state-of-the-art clustering approaches on several datasets to demonstrate the effectiveness and interpretability of the proposed method.\n          <\/jats:p>","DOI":"10.1145\/3717068","type":"journal-article","created":{"date-parts":[[2025,2,18]],"date-time":"2025-02-18T16:00:00Z","timestamp":1739894400000},"page":"1-23","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":0,"title":["CM-CaFE: A Clustering Method with Causality-based Feature Embedding"],"prefix":"10.1145","volume":"19","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-7623-452X","authenticated-orcid":false,"given":"Xuechun","family":"Jing","sequence":"first","affiliation":[{"name":"The School of Computer and Information Technology, Shanxi University, Taiyuan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-1111-8461","authenticated-orcid":false,"given":"Fuyuan","family":"Cao","sequence":"additional","affiliation":[{"name":"The School of Computer and Information Technology, Shanxi University, Taiyuan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-2442-4572","authenticated-orcid":false,"given":"Kui","family":"Yu","sequence":"additional","affiliation":[{"name":"The School of Computer and Information, Hefei University of Technology, Hefei, China"}],"role":[{"role":"author","vocabulary":"crossref"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-5887-9327","authenticated-orcid":false,"given":"Jiye","family":"Liang","sequence":"additional","affiliation":[{"name":"The School of Computer and Information Technology, Shanxi University, Taiyuan, China"}],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"320","published-online":{"date-parts":[[2025,3,22]]},"reference":[{"key":"e_1_3_1_2_1","doi-asserted-by":"publisher","DOI":"10.1002\/wics.101"},{"issue":"3","key":"e_1_3_1_3_1","first-page":"110","article-title":"Feature selection for clustering: A review","volume":"21","author":"Alelyani Salem","year":"2016","unstructured":"Salem Alelyani, Jiliang Tang, and Huan Liu. 2016. Feature selection for clustering: A review. Encyclopedia of Database Systems 21, 3 (2016), 110\u2013121.","journal-title":"Encyclopedia of Database Systems"},{"key":"e_1_3_1_4_1","doi-asserted-by":"publisher","DOI":"10.4086\/toc.2012.v008a006"},{"key":"e_1_3_1_5_1","first-page":"1027","volume-title":"Proceedings of the Annual ACM-SIAM Symposium on Discrete Algorithms","author":"Arthur David","year":"2007","unstructured":"David Arthur and Sergei Vassilvitskii. 2007. K-means++: The advantages of careful seeding. In Proceedings of the Annual ACM-SIAM Symposium on Discrete Algorithms, 1027\u20131035."},{"key":"e_1_3_1_6_1","first-page":"21","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Cai Jinyu","year":"2022","unstructured":"Jinyu Cai, Jicong Fan, Wenzhong Guo, Shiping Wang, Yunhe Zhang, and Zhao Zhang. 2022a. Efficient deep embedded subspace clustering. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 21\u201330."},{"key":"e_1_3_1_7_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2005.198"},{"key":"e_1_3_1_8_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2021.108386"},{"key":"e_1_3_1_9_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2889949"},{"key":"e_1_3_1_10_1","doi-asserted-by":"publisher","DOI":"10.1109\/TIP.2023.3263102"},{"key":"e_1_3_1_11_1","doi-asserted-by":"publisher","DOI":"10.1109\/TAI.2022.3150264"},{"issue":"9","key":"e_1_3_1_12_1","doi-asserted-by":"crossref","first-page":"6108","DOI":"10.1109\/TNNLS.2021.3133337","article-title":"Toward unique and unbiased causal effect estimation from data with hidden variables","volume":"34","author":"Cheng Debo","year":"2022","unstructured":"Debo Cheng, Jiuyong Li, Lin Liu, Kui Yu, Thuc Duy Le, and Jixue Liu. 2022b. Toward unique and unbiased causal effect estimation from data with hidden variables. IEEE Transactions on Neural Networks and Learning Systems 34, 9 (2022), 6108\u20136120.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_1_13_1","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-022-00445-z"},{"key":"e_1_3_1_14_1","first-page":"226","volume-title":"Proceedings of the International Conference on Knowledge Discovery and Data Mining","author":"Ester Martin","year":"1996","unstructured":"Martin Ester, Hans-Peter Kriegel, J\u00f6rg Sander, and Xiaowei Xu. 1996. A density-based algorithm for discovering clusters in large spatial databases with noise. In Proceedings of the International Conference on Knowledge Discovery and Data Mining, 226\u2013231."},{"key":"e_1_3_1_15_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.spasta.2022.100621"},{"key":"e_1_3_1_16_1","doi-asserted-by":"publisher","DOI":"10.24963\/ijcai.2017\/243"},{"key":"e_1_3_1_17_1","first-page":"1","volume-title":"Proceedings of the IEEE International Conference on Progress in Informatics and Computing","author":"Han Mingyue","year":"2021","unstructured":"Mingyue Han and Yinglin Wang. 2021. A survey on the identification of causal relation in texts. In Proceedings of the IEEE International Conference on Progress in Informatics and Computing, 1\u20137."},{"issue":"1","key":"e_1_3_1_18_1","first-page":"100","article-title":"Algorithm AS 136: A k-means clustering algorithm","volume":"28","author":"Hartigan John A.","year":"1979","unstructured":"John A. Hartigan and Manchek A. Wong. 1979. Algorithm AS 136: A k-means clustering algorithm. Journal of the Royal Statistical Society 28, 1 (1979), 100\u2013108.","journal-title":"Journal of the Royal Statistical Society"},{"key":"e_1_3_1_19_1","doi-asserted-by":"publisher","DOI":"10.1109\/tfuzz.2012.2201485"},{"key":"e_1_3_1_20_1","doi-asserted-by":"publisher","DOI":"10.1007\/s13042-022-01557-z"},{"key":"e_1_3_1_21_1","doi-asserted-by":"publisher","DOI":"10.1109\/ICPR.2014.272"},{"key":"e_1_3_1_22_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2021.3049633"},{"key":"e_1_3_1_23_1","doi-asserted-by":"publisher","DOI":"10.1080\/03610928008827904"},{"key":"e_1_3_1_24_1","doi-asserted-by":"crossref","first-page":"846","DOI":"10.1145\/3534678.3539407","volume-title":"Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining","author":"Leiber Collin","year":"2022","unstructured":"Collin Leiber, Lena G. M. Bauer, Michael Neumayr, Claudia Plant, and Christian B\u00f6hm. 2022. The DipEncoder: Enforcing multimodality in autoencoders. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 846\u2013856."},{"issue":"12","key":"e_1_3_1_25_1","doi-asserted-by":"crossref","first-page":"13848","DOI":"10.1109\/TCYB.2021.3109066","article-title":"An integrated cluster detection, optimization, and interpretation approach for financial data","volume":"52","author":"Li Tie","year":"2022","unstructured":"Tie Li, Gang Kou, Yi Peng, and Philip S. Yu. 2022. An integrated cluster detection, optimization, and interpretation approach for financial data. IEEE Transactions on Cybernetics 52, 12 (2022), 13848\u201313861.","journal-title":"IEEE Transactions on Cybernetics"},{"key":"e_1_3_1_26_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2019.2955388"},{"issue":"2","key":"e_1_3_1_27_1","first-page":"940","article-title":"Clustering high-dimensional data via feature selection","volume":"79","author":"Liu Tianqi","year":"2022","unstructured":"Tianqi Liu, Yu Lu, Biqing Zhu, and Hongyu Zhao. 2022. Clustering high-dimensional data via feature selection. Biometrics 79, 2 (2022), 940\u2013950.","journal-title":"Biometrics"},{"key":"e_1_3_1_28_1","first-page":"344","volume-title":"Proceedings of the 7th International Conference on Disruptive Technologies (ICDT)","author":"Mishra Annu","year":"2023","unstructured":"Annu Mishra, Pankaj Gupta, and Peeyush Tewari. 2023. Biomedical image segmentation using integrated FCM clustering modified with regularized level set method. In Proceedings of the 7th International Conference on Disruptive Technologies (ICDT), 344\u2013348."},{"key":"e_1_3_1_29_1","doi-asserted-by":"publisher","DOI":"10.1177\/0013164484441003"},{"key":"e_1_3_1_30_1","doi-asserted-by":"crossref","first-page":"512","DOI":"10.1145\/2623330.2623611","volume-title":"Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining","author":"Nguyen Xuan Vinh","year":"2014","unstructured":"Xuan Vinh Nguyen, Jeffrey Chan, Simone Romano, and James Bailey. 2014. Effective global approaches for mutual information based feature selection. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 512\u2013521."},{"key":"e_1_3_1_31_1","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2022.3155450"},{"key":"e_1_3_1_32_1","doi-asserted-by":"publisher","DOI":"10.1017\/CBO9780511803161"},{"key":"e_1_3_1_33_1","first-page":"14855","volume-title":"Proceedings of the 33rd Advances in Neural Information Processing Systems","author":"Pei Shenfei","year":"2020","unstructured":"Shenfei Pei, Feiping Nie, Rong Wang, and Xuelong Li. 2020. Efficient clustering based on a unified view of K-means and ratio-cut. In Proceedings of the 33rd Advances in Neural Information Processing Systems, 14855\u201314866."},{"key":"e_1_3_1_34_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.ijar.2006.06.008"},{"key":"e_1_3_1_35_1","doi-asserted-by":"publisher","DOI":"10.1023\/A:1025667309714"},{"key":"e_1_3_1_36_1","first-page":"4644","volume-title":"Proceedings of the International Conference on Neural Information Processing Systems","author":"Shishkin Alexander","year":"2016","unstructured":"Alexander Shishkin, Anastasia Bezzubtseva, Alexey Drutsa, Ilia Shishkov, Ekaterina Gladkikh, Gleb Gusev, and Pavel Serdyukov. 2016. Efficient high-order interaction-aware feature selection based on conditional mutual information. In Proceedings of the International Conference on Neural Information Processing Systems, 4644\u20134652."},{"issue":"47","key":"e_1_3_1_37_1","first-page":"1393","article-title":"Feature selection via dependence maximization","volume":"13","author":"Song Le","year":"2012","unstructured":"Le Song, Alex Smola, Arthur Gretton, Justin Bedo, and Karsten Borgwardt. 2012. Feature selection via dependence maximization. Journal of Machine Learning Research 13, 47 (2012), 1393\u20131434.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_1_38_1","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2017.2777489"},{"key":"e_1_3_1_39_1","first-page":"78205","volume-title":"Proceedings of the 36th Advances in Neural Information Processing Systems","author":"Wu Zhengxuan","year":"2023","unstructured":"Zhengxuan Wu, Atticus Geiger, Thomas Icard, Christopher Potts, and Noah Goodman. 2023. Interpretability at scale: Identifying causal mechanisms in alpaca. In Proceedings of the 36th Advances in Neural Information Processing Systems, 78205\u201378226."},{"key":"e_1_3_1_40_1","first-page":"478","volume-title":"Proceedings of the International Conference on Machine Learning","author":"Xie Junyuan","year":"2016","unstructured":"Junyuan Xie, Ross Girshick, and Ali Farhadi. 2016. Unsupervised deep embedding for clustering analysis. In Proceedings of the International Conference on Machine Learning, 478\u2013487."},{"key":"e_1_3_1_41_1","doi-asserted-by":"publisher","DOI":"10.1145\/3436891"},{"key":"e_1_3_1_42_1","doi-asserted-by":"publisher","DOI":"10.1109\/TCYB.2020.3034157"},{"key":"e_1_3_1_43_1","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2018.2877660"},{"key":"e_1_3_1_44_1","first-page":"10148","volume-title":"Proceedings of the 33rd Advances in Neural Information Processing Systems","author":"Zhang Zhiyue","year":"2020","unstructured":"Zhiyue Zhang, Kenneth Lange, and Jason Xu. 2020b. Simple and scalable sparse k-means clustering via feature ranking. In Proceedings of the 33rd Advances in Neural Information Processing Systems, 10148\u201310160."},{"key":"e_1_3_1_45_1","doi-asserted-by":"publisher","DOI":"10.1023\/b:mach.0000027785.44527.d6"},{"issue":"4","key":"e_1_3_1_46_1","doi-asserted-by":"crossref","first-page":"1073","DOI":"10.1109\/TFUZZ.2021.3052362","article-title":"Robust jointly sparse fuzzy clustering with neighborhood structure preservation","volume":"30","author":"Zhou Jie","year":"2021","unstructured":"Jie Zhou, Witold Pedrycz, Can Gao, Zhihui Lai, Jun Wan, and Zhong Ming. 2021. Robust jointly sparse fuzzy clustering with neighborhood structure preservation. IEEE Transactions on Fuzzy Systems 30, 4 (2021), 1073\u20131087.","journal-title":"IEEE Transactions on Fuzzy Systems"},{"key":"e_1_3_1_47_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2017.01.016"}],"container-title":["ACM Transactions on Knowledge Discovery from Data"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3717068","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3717068","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,19]],"date-time":"2025-06-19T01:18:55Z","timestamp":1750295935000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3717068"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2025,3,22]]},"references-count":46,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2025,4,30]]}},"alternative-id":["10.1145\/3717068"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3717068","relation":{},"ISSN":["1556-4681","1556-472X"],"issn-type":[{"type":"print","value":"1556-4681"},{"type":"electronic","value":"1556-472X"}],"subject":[],"published":{"date-parts":[[2025,3,22]]},"assertion":[{"value":"2024-03-27","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-02-08","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2025-03-22","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}