{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,2]],"date-time":"2026-06-02T11:56:30Z","timestamp":1780401390123,"version":"3.54.1"},"reference-count":179,"publisher":"Association for Computing Machinery (ACM)","issue":"7","license":[{"start":{"date-parts":[[2024,4,9]],"date-time":"2024-04-09T00:00:00Z","timestamp":1712620800000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.acm.org\/publications\/policies\/copyright_policy#Background"}],"funder":[{"name":"National Key R&D Project","award":["2021YFC3320301"],"award-info":[{"award-number":["2021YFC3320301"]}]},{"DOI":"10.13039\/501100001809","name":"National Natural Science Foundation of China","doi-asserted-by":"crossref","award":["62171325"],"award-info":[{"award-number":["62171325"]}],"id":[{"id":"10.13039\/501100001809","id-type":"DOI","asserted-by":"crossref"}]},{"name":"Hubei Key R&D","award":["2022BAA033"],"award-info":[{"award-number":["2022BAA033"]}]},{"name":"CAAI-Huawei MindSpore Open Fund"}],"content-domain":{"domain":["dl.acm.org"],"crossmark-restriction":true},"short-container-title":["ACM Comput. Surv."],"published-print":{"date-parts":[[2024,7,31]]},"abstract":"<jats:p>\n            Large-scale datasets have played a crucial role in the advancement of computer vision. However, they often suffer from problems such as class imbalance, noisy labels, dataset bias, or high resource costs, which can inhibit model performance and reduce trustworthiness. With the advocacy of data-centric research, various data-centric solutions have been proposed to solve the dataset problems mentioned above. They improve the quality of datasets by re-organizing them, which we call dataset refinement. In this survey, we provide a comprehensive and structured overview of recent advances in dataset refinement for problematic computer vision datasets.\n            <jats:xref ref-type=\"fn\">\n              <jats:sup>1<\/jats:sup>\n            <\/jats:xref>\n            Firstly, we summarize and analyze the various problems encountered in large-scale computer vision datasets. Then, we classify the dataset refinement algorithms into three categories based on the refinement process: data sampling, data subset selection, and active learning. In addition, we organize these dataset refinement methods according to the addressed data problems and provide a systematic comparative description. We point out that these three types of dataset refinement have distinct advantages and disadvantages for dataset problems, which informs the choice of the data-centric method appropriate to a particular research objective. Finally, we summarize the current literature and propose potential future research topics.\n          <\/jats:p>","DOI":"10.1145\/3627157","type":"journal-article","created":{"date-parts":[[2023,10,10]],"date-time":"2023-10-10T11:30:22Z","timestamp":1696937422000},"page":"1-34","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/crossmark-policy","source":"Crossref","is-referenced-by-count":14,"title":["A Survey of Dataset Refinement for Problems in Computer Vision Datasets"],"prefix":"10.1145","volume":"56","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-1273-9043","authenticated-orcid":false,"given":"Zhijing","family":"Wan","sequence":"first","affiliation":[{"name":"National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-5016-587X","authenticated-orcid":false,"given":"Zhixiang","family":"Wang","sequence":"additional","affiliation":[{"name":"Graduate School of Information Science and Technology, The University of Tokyo, Tokyo, Japan"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0009-0009-7260-4077","authenticated-orcid":false,"given":"Cheukting","family":"Chung","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-3846-9157","authenticated-orcid":false,"given":"Zheng","family":"Wang","sequence":"additional","affiliation":[{"name":"National Engineering Research Center for Multimedia Software, Institute of Artificial Intelligence, School of Computer Science, Wuhan University, Wuhan, China"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"320","published-online":{"date-parts":[[2024,4,9]]},"reference":[{"key":"e_1_3_2_2_2","article-title":"Variance reduction in SGD by distributed importance sampling","author":"Alain Guillaume","year":"2015","unstructured":"Guillaume Alain, Alex Lamb, Chinnadhurai Sankar, Aaron Courville, and Yoshua Bengio. 2015. Variance reduction in SGD by distributed importance sampling. arXiv preprint arXiv:1511.06481 (2015).","journal-title":"arXiv preprint arXiv:1511.06481"},{"key":"e_1_3_2_3_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.knosys.2021.106771"},{"key":"e_1_3_2_4_2","first-page":"11816","article-title":"Gradient based sample selection for online continual learning","volume":"32","author":"Aljundi Rahaf","year":"2019","unstructured":"Rahaf Aljundi, Min Lin, Baptiste Goujaud, and Yoshua Bengio. 2019. Gradient based sample selection for online continual learning. Advances in Neural Information Processing Systems (NIPS) 32 (2019), 11816\u201311825.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_5_2","first-page":"335","volume-title":"British Machine Vision Conference (BMVC)","author":"Arazo Eric","year":"2021","unstructured":"Eric Arazo, Diego Ortego, Paul Albert, Noel E. O\u2019Connor, and Kevin McGuinness. 2021. How important is importance sampling for deep budgeted training?. In British Machine Vision Conference (BMVC). BMVA Press, 335."},{"key":"e_1_3_2_6_2","unstructured":"Devansh Arpit Stanis\u0142aw Jastrz\u0229bski Nicolas Ballas David Krueger Emmanuel Bengio Maxinder S. Kanwal Tegan Maharaj Asja Fischer Aaron Courville Yoshua Bengio and Simon Lacoste-Julien. 2017. A closer look at memorization in deep networks. In International Conference on Machine Learning (ICML\u201917) Vol. 70. PMLR Sydney NSW 233\u2013242."},{"key":"e_1_3_2_7_2","article-title":"Deep batch active learning by diverse, uncertain gradient lower bounds","author":"Ash Jordan T.","year":"2019","unstructured":"Jordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2019. Deep batch active learning by diverse, uncertain gradient lower bounds. arXiv preprint arXiv:1906.03671 (2019).","journal-title":"arXiv preprint arXiv:1906.03671"},{"key":"e_1_3_2_8_2","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-022-09744-2"},{"key":"e_1_3_2_9_2","first-page":"209","volume-title":"International Conference on Machine Learning (ICML)","volume":"37","author":"Bachem Olivier","year":"2015","unstructured":"Olivier Bachem, Mario Lucic, and Andreas Krause. 2015. Coresets for nonparametric estimation-the case of DP-means. In International Conference on Machine Learning (ICML), Vol. 37. PMLR, Lille, France, 209\u2013217."},{"key":"e_1_3_2_10_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01270-0_28"},{"key":"e_1_3_2_11_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00976"},{"key":"e_1_3_2_12_2","doi-asserted-by":"publisher","DOI":"10.1145\/1553374.1553380"},{"key":"e_1_3_2_13_2","article-title":"Are we done with ImageNet?","author":"Beyer Lucas","year":"2020","unstructured":"Lucas Beyer, Olivier J. H\u00e9naff, Alexander Kolesnikov, Xiaohua Zhai, and A\u00e4ron van den Oord. 2020. Are we done with ImageNet? arXiv preprint arXiv:2006.07159 (2020).","journal-title":"arXiv preprint arXiv:2006.07159"},{"key":"e_1_3_2_14_2","first-page":"9","volume-title":"NIPS Workshop on Analyzing Networks and Learning with Graphs","author":"Bilgic Mustafa","year":"2009","unstructured":"Mustafa Bilgic and Lise Getoor. 2009. Link-based active learning. In NIPS Workshop on Analyzing Networks and Learning with Graphs, Vol. 4. 9."},{"key":"e_1_3_2_15_2","article-title":"Semantic redundancies in image-classification datasets: The 10% you don\u2019t need","author":"Birodkar Vighnesh","year":"2019","unstructured":"Vighnesh Birodkar, Hossein Mobahi, and Samy Bengio. 2019. Semantic redundancies in image-classification datasets: The 10% you don\u2019t need. arXiv preprint arXiv:1901.11409 (2019).","journal-title":"arXiv preprint arXiv:1901.11409"},{"key":"e_1_3_2_16_2","first-page":"14879","article-title":"Coresets via bilevel optimization for continual learning and streaming","volume":"33","author":"Borsos Zal\u00e1n","year":"2020","unstructured":"Zal\u00e1n Borsos, Mojmir Mutny, and Andreas Krause. 2020. Coresets via bilevel optimization for continual learning and streaming. Advances in Neural Information Processing Systems (NIPS) 33 (2020), 14879\u201314890.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_17_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.606"},{"key":"e_1_3_2_18_2","first-page":"698","volume-title":"International Conference on Machine Learning (ICML)","volume":"80","author":"Campbell Trevor","year":"2018","unstructured":"Trevor Campbell and Tamara Broderick. 2018. Bayesian coreset construction via greedy iterative geodesic ascent. In International Conference on Machine Learning (ICML), Vol. 80. PMLR, Stockholm, Sweden, 698\u2013706."},{"key":"e_1_3_2_19_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR42600.2020.00571"},{"key":"e_1_3_2_20_2","first-page":"4750","volume-title":"Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR)","author":"Cazenavette George","year":"2022","unstructured":"George Cazenavette, Tongzhou Wang, Antonio Torralba, Alexei A. Efros, and Jun-Yan Zhu. 2022. Dataset distillation by matching training trajectories. In Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, New Orleans, LA, USA, 4750\u20134759."},{"key":"e_1_3_2_21_2","first-page":"1002","article-title":"Active bias: Training more accurate neural networks by emphasizing high variance samples","volume":"30","author":"Chang Haw-Shiuan","year":"2017","unstructured":"Haw-Shiuan Chang, Erik Learned-Miller, and Andrew McCallum. 2017. Active bias: Training more accurate neural networks by emphasizing high variance samples. Advances in Neural Information Processing Systems (NIPS) 30 (2017), 1002\u20131012.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_22_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.953"},{"key":"e_1_3_2_23_2","first-page":"1062","volume-title":"International Conference on Machine Learning (ICML)","volume":"97","author":"Chen Pengfei","year":"2019","unstructured":"Pengfei Chen, Ben Ben Liao, Guangyong Chen, and Shengyu Zhang. 2019. Understanding and utilizing deep neural networks trained with noisy labels. In International Conference on Machine Learning (ICML), Vol. 97. PMLR, Long Beach, California, USA, 1062\u20131070."},{"key":"e_1_3_2_24_2","article-title":"Learning with instance-dependent label noise: A sample sieve approach","author":"Cheng Hao","year":"2020","unstructured":"Hao Cheng, Zhaowei Zhu, Xingyu Li, Yifei Gong, Xing Sun, and Yang Liu. 2020. Learning with instance-dependent label noise: A sample sieve approach. arXiv preprint arXiv:2010.02347 (2020).","journal-title":"arXiv preprint arXiv:2010.02347"},{"key":"e_1_3_2_25_2","article-title":"Selection via proxy: Efficient data selection for deep learning","author":"Coleman Cody","year":"2019","unstructured":"Cody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman, Peter Bailis, Percy Liang, Jure Leskovec, and Matei Zaharia. 2019. Selection via proxy: Efficient data selection for deep learning. arXiv preprint arXiv:1906.11829 (2019).","journal-title":"arXiv preprint arXiv:1906.11829"},{"key":"e_1_3_2_26_2","article-title":"A meta-learning approach to one-step active learning","author":"Contardo Gabriella","year":"2017","unstructured":"Gabriella Contardo, Ludovic Denoyer, and Thierry Arti\u00e8res. 2017. A meta-learning approach to one-step active learning. arXiv preprint arXiv:1706.08334 (2017).","journal-title":"arXiv preprint arXiv:1706.08334"},{"key":"e_1_3_2_27_2","doi-asserted-by":"publisher","DOI":"10.1080\/00401706.1980.10486199"},{"key":"e_1_3_2_28_2","doi-asserted-by":"publisher","DOI":"10.1109\/SIBGRAPI51738.2020.00010"},{"key":"e_1_3_2_29_2","doi-asserted-by":"publisher","DOI":"10.1109\/TGRS.2010.2072929"},{"key":"e_1_3_2_30_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206848"},{"key":"e_1_3_2_31_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2017.02.033"},{"key":"e_1_3_2_32_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.neucom.2022.04.083"},{"key":"e_1_3_2_33_2","unstructured":"Chris Drummond and Robert C. Holte. 2003. C4. 5 class imbalance and cost sensitivity: Why under-sampling beats oversampling. In Workshop on Learning from Imbalanced Datasets II Vol. 11. Citeseer 1\u20138."},{"key":"e_1_3_2_34_2","doi-asserted-by":"publisher","DOI":"10.1214\/17-AOS1679"},{"key":"e_1_3_2_35_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2511748"},{"key":"e_1_3_2_36_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2013.33"},{"key":"e_1_3_2_37_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.cviu.2022.103552"},{"key":"e_1_3_2_38_2","article-title":"Learning what data to learn","author":"Fan Yang","year":"2017","unstructured":"Yang Fan, Fei Tian, Tao Qin, Jiang Bian, and Tie-Yan Liu. 2017. Learning what data to learn. arXiv preprint arXiv:1702.08635 (2017).","journal-title":"arXiv preprint arXiv:1702.08635"},{"key":"e_1_3_2_39_2","article-title":"Learning how to active learn: A deep reinforcement learning approach","author":"Fang Meng","year":"2017","unstructured":"Meng Fang, Yuan Li, and Trevor Cohn. 2017. Learning how to active learn: A deep reinforcement learning approach. arXiv preprint arXiv:1708.02383 (2017).","journal-title":"arXiv preprint arXiv:1708.02383"},{"key":"e_1_3_2_40_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1007\/978-3-7908-2151-2","volume-title":"Facility Location: Concepts, Models, Algorithms and Case Studies","author":"Farahani Reza Zanjirani","year":"2009","unstructured":"Reza Zanjirani Farahani and Masoud Hekmatfar. 2009. Facility Location: Concepts, Models, Algorithms and Case Studies. Springer Science & Business Media, 1\u2013545."},{"key":"e_1_3_2_41_2","first-page":"23","volume-title":"Core-Sets: Updated Survey","author":"Feldman Dan","year":"2020","unstructured":"Dan Feldman. 2020. Core-Sets: Updated Survey. Springer International Publishing, Cham, 23\u201344."},{"key":"e_1_3_2_42_2","first-page":"1","article-title":"Hybrid machine learning methods combined with computer vision approaches to estimate biophysical parameters of pastures","author":"Franco Victor Rezende","year":"2022","unstructured":"Victor Rezende Franco, Marcos Cicarini Hott, Ricardo Guimar\u00e3es Andrade, and Leonardo Goliatt. 2022. Hybrid machine learning methods combined with computer vision approaches to estimate biophysical parameters of pastures. Evolutionary Intelligence (2022), 1\u201314.","journal-title":"Evolutionary Intelligence"},{"key":"e_1_3_2_43_2","volume-title":"Submodular Functions and Optimization","author":"Fujishige Satoru","year":"2005","unstructured":"Satoru Fujishige. 2005. Submodular Functions and Optimization. Elsevier Science."},{"key":"e_1_3_2_44_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-020-00257-z"},{"key":"e_1_3_2_45_2","article-title":"ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness","author":"Geirhos Robert","year":"2018","unstructured":"Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. 2018. ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness. arXiv preprint arXiv:1811.12231 (2018).","journal-title":"arXiv preprint arXiv:1811.12231"},{"key":"e_1_3_2_46_2","first-page":"2242","volume-title":"International Conference on Machine Learning (ICML)","volume":"97","author":"Ghorbani Amirata","year":"2019","unstructured":"Amirata Ghorbani and James Zou. 2019. Data Shapley: Equitable valuation of data for machine learning. In International Conference on Machine Learning (ICML), Vol. 97. PMLR, Long Beach, California, USA, 2242\u20132251."},{"key":"e_1_3_2_47_2","article-title":"Towards understanding deep learning from noisy labels with small-loss criterion","author":"Gui Xian-Jin","year":"2021","unstructured":"Xian-Jin Gui, Wei Wang, and Zhang-Hao Tian. 2021. Towards understanding deep learning from noisy labels with small-loss criterion. arXiv preprint arXiv:2106.09291 (2021).","journal-title":"arXiv preprint arXiv:2106.09291"},{"key":"e_1_3_2_48_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-12423-5_14"},{"key":"e_1_3_2_49_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01249-6_9"},{"key":"e_1_3_2_50_2","first-page":"802","article-title":"Active instance sampling via matrix partition","volume":"23","author":"Guo Yuhong","year":"2010","unstructured":"Yuhong Guo. 2010. Active instance sampling via matrix partition. Advances in Neural Information Processing Systems 23 (2010), 802\u2013810.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_51_2","first-page":"8536","article-title":"Co-teaching: Robust training of deep neural networks with extremely noisy labels","volume":"31","author":"Han Bo","year":"2018","unstructured":"Bo Han, Quanming Yao, Xingrui Yu, Gang Niu, Miao Xu, Weihua Hu, Ivor Tsang, and Masashi Sugiyama. 2018. Co-teaching: Robust training of deep neural networks with extremely noisy labels. Advances in Neural Information Processing Systems (NIPS) 31 (2018), 8536\u20138546.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"issue":"5","key":"e_1_3_2_52_2","first-page":"2223","article-title":"SlimML: Removing non-critical input data in large-scale iterative machine learning","volume":"33","author":"Han Rui","year":"2019","unstructured":"Rui Han, Chi Harold Liu, Shilin Li, Lydia Y. Chen, Guoren Wang, Jian Tang, and Jieping Ye. 2019. SlimML: Removing non-critical input data in large-scale iterative machine learning. IEEE Transactions on Knowledge and Data Engineering 33, 5 (2019), 2223\u20132236.","journal-title":"IEEE Transactions on Knowledge and Data Engineering"},{"key":"e_1_3_2_53_2","doi-asserted-by":"publisher","DOI":"10.1007\/s00454-006-1271-x"},{"key":"e_1_3_2_54_2","article-title":"Deep active learning with adaptive acquisition","author":"Hau\u00dfmann Manuel","year":"2019","unstructured":"Manuel Hau\u00dfmann, Fred A. Hamprecht, and Melih Kandemir. 2019. Deep active learning with adaptive acquisition. arXiv preprint arXiv:1906.11471 (2019).","journal-title":"arXiv preprint arXiv:1906.11471"},{"key":"e_1_3_2_55_2","doi-asserted-by":"publisher","DOI":"10.1109\/TKDE.2008.239"},{"key":"e_1_3_2_56_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.90"},{"key":"e_1_3_2_57_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00342"},{"key":"e_1_3_2_58_2","first-page":"4080","article-title":"Coresets for scalable Bayesian logistic regression","volume":"29","author":"Huggins Jonathan","year":"2016","unstructured":"Jonathan Huggins, Trevor Campbell, and Tamara Broderick. 2016. Coresets for scalable Bayesian logistic regression. Advances in Neural Information Processing Systems (NIPS) 29 (2016), 4080\u20134088.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_59_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v32i1.11595"},{"key":"e_1_3_2_60_2","unstructured":"Angela H. Jiang Daniel L.-K. Wong Giulio Zhou David G. Andersen Jeffrey Dean Gregory R. Ganger Gauri Joshi Michael Kaminksy Michael Kozuch Zachary C. Lipton and Padmanabhan Pillai. 2019. Accelerating deep learning by focusing on the biggest losers. arXiv preprint arXiv:1910.00762 (2019)."},{"key":"e_1_3_2_61_2","first-page":"2304","volume-title":"International Conference on Machine Learning (ICML)","volume":"80","author":"Jiang Lu","year":"2018","unstructured":"Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei. 2018. MentorNet: Learning data-driven curriculum for very deep neural networks on corrupted labels. In International Conference on Machine Learning (ICML), Vol. 80. PMLR, Stockholm, Sweden, 2304\u20132313."},{"key":"e_1_3_2_62_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2009.5206627"},{"key":"e_1_3_2_63_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-87237-3_1"},{"key":"e_1_3_2_64_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00945"},{"key":"e_1_3_2_65_2","first-page":"2525","volume-title":"International Conference on Machine Learning (ICML)","volume":"80","author":"Katharopoulos Angelos","year":"2018","unstructured":"Angelos Katharopoulos and Fran\u00e7ois Fleuret. 2018. Not all samples are created equal: Deep learning with importance sampling. In International Conference on Machine Learning (ICML), Vol. 80. PMLR, Stockholm, Sweden, 2525\u20132534."},{"issue":"4","key":"e_1_3_2_66_2","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1145\/3343440","article-title":"A systematic review on imbalanced data challenges in machine learning: Applications and solutions","volume":"52","author":"Kaur Harsurinder","year":"2019","unstructured":"Harsurinder Kaur, Husanbir Singh Pannu, and Avleen Kaur Malhi. 2019. A systematic review on imbalanced data challenges in machine learning: Applications and solutions. ACM Computing Surveys (CSUR) 52, 4 (2019), 1\u201336.","journal-title":"ACM Computing Surveys (CSUR)"},{"key":"e_1_3_2_67_2","doi-asserted-by":"publisher","DOI":"10.1109\/WACV.2019.00142"},{"key":"e_1_3_2_68_2","first-page":"669","volume-title":"International Conference on Artificial Intelligence and Statistics","volume":"108","author":"Kawaguchi Kenji","year":"2020","unstructured":"Kenji Kawaguchi and Haihao Lu. 2020. Ordered SGD: A new stochastic optimization framework for empirical risk minimization. In International Conference on Artificial Intelligence and Statistics, Vol. 108. PMLR, Palermo, Sicily, Italy, 669\u2013679."},{"key":"e_1_3_2_69_2","doi-asserted-by":"crossref","unstructured":"Douwe Kiela Max Bartolo Yixin Nie Divyansh Kaushik Atticus Geiger Zhengxuan Wu Bertie Vidgen Grusha Prasad Amanpreet Singh Pratik Ringshia Zhiyi Ma Tristan Thrush Sebastian Riedel Zeerak Waseem Pontus Stenetorp Robin Jia Mohit Bansal Christopher Potts and Adina Williams. 2021. Dynabench: Rethinking benchmarking in NLP. arXiv preprint arXiv:2104.14337 (2021).","DOI":"10.18653\/v1\/2021.naacl-main.324"},{"key":"e_1_3_2_70_2","unstructured":"Krishnateja Killamsetty Guttu Sai Abhishek Aakriti Lnu Ganesh Ramakrishnan Alexandre Evfimievski Lucian Popa and Rishabh Iyer. 2022. Automata: Gradient based data subset selection for compute-efficient hyper-parameter tuning. Advances in Neural Information Processing Systems (NIPS) 35 (2022) 28721\u201328733."},{"key":"e_1_3_2_71_2","first-page":"5464","volume-title":"International Conference on Machine Learning (ICML)","volume":"139","author":"Killamsetty Krishnateja","year":"2021","unstructured":"Krishnateja Killamsetty, S. Durga, Ganesh Ramakrishnan, Abir De, and Rishabh Iyer. 2021. Grad-Match: Gradient matching based data subset selection for efficient deep model training. In International Conference on Machine Learning (ICML), Vol. 139. PMLR, 5464\u20135474."},{"key":"e_1_3_2_72_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i9.16988"},{"key":"e_1_3_2_73_2","first-page":"14488","article-title":"Retrieve: Coreset selection for efficient and robust semi-supervised learning","volume":"34","author":"Killamsetty Krishnateja","year":"2021","unstructured":"Krishnateja Killamsetty, Xujiang Zhao, Feng Chen, and Rishabh Iyer. 2021. Retrieve: Coreset selection for efficient and robust semi-supervised learning. Advances in Neural Information Processing Systems (NIPS) 34 (2021), 14488\u201314501.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_74_2","first-page":"1885","volume-title":"International Conference on Machine Learning (ICML)","volume":"70","author":"Koh Pang Wei","year":"2017","unstructured":"Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. In International Conference on Machine Learning (ICML), Vol. 70. PMLR, Sydney, NSW, Australia, 1885\u20131894."},{"key":"e_1_3_2_75_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i9.21264"},{"key":"e_1_3_2_76_2","unstructured":"Alex Krizhevsky and Geoffrey Hinton. 2009. Learning multiple layers of features from tiny images. Computer Science Department University of Toronto Tech. Rep. 1 (01 2009)."},{"key":"e_1_3_2_77_2","first-page":"1189","article-title":"Self-paced learning for latent variable models","volume":"23","author":"Kumar M.","year":"2010","unstructured":"M. Kumar, Benjamin Packer, and Daphne Koller. 2010. Self-paced learning for latent variable models. Advances in Neural Information Processing Systems (NIPS) 23 (2010), 1189\u20131197.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_78_2","article-title":"Are all training examples equally valuable?","author":"Lapedriza Agata","year":"2013","unstructured":"Agata Lapedriza, Hamed Pirsiavash, Zoya Bylinskii, and Antonio Torralba. 2013. Are all training examples equally valuable? arXiv preprint arXiv:1311.6510 (2013).","journal-title":"arXiv preprint arXiv:1311.6510"},{"key":"e_1_3_2_79_2","first-page":"1078","volume-title":"International Conference on Machine Learning (ICML)","volume":"119","author":"Bras Ronan Le","year":"2020","unstructured":"Ronan Le Bras, Swabha Swayamdipta, Chandra Bhagavatula, Rowan Zellers, Matthew Peters, Ashish Sabharwal, and Yejin Choi. 2020. Adversarial filters of dataset biases. In International Conference on Machine Learning (ICML), Vol. 119. PMLR, 1078\u20131088."},{"key":"e_1_3_2_80_2","doi-asserted-by":"publisher","DOI":"10.1109\/5.726791"},{"key":"e_1_3_2_81_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.01091"},{"key":"e_1_3_2_82_2","first-page":"13","volume-title":"ACM SIGIR Forum","author":"Lewis David D.","year":"1995","unstructured":"David D. Lewis. 1995. A sequential algorithm for training text classifiers: Corrigendum and additional data. In ACM SIGIR Forum. ACM New York, NY, USA, 13\u201319."},{"key":"e_1_3_2_83_2","doi-asserted-by":"publisher","DOI":"10.1145\/3136625"},{"key":"e_1_3_2_84_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-01231-1_32"},{"key":"e_1_3_2_85_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2019.00980"},{"key":"e_1_3_2_86_2","doi-asserted-by":"publisher","DOI":"10.1038\/s42256-022-00516-1"},{"key":"e_1_3_2_87_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.324"},{"key":"e_1_3_2_88_2","doi-asserted-by":"publisher","DOI":"10.1145\/3510414"},{"key":"e_1_3_2_89_2","doi-asserted-by":"publisher","DOI":"10.1109\/TPAMI.2015.2456899"},{"key":"e_1_3_2_90_2","article-title":"Online batch selection for faster training of neural networks","author":"Loshchilov Ilya","year":"2015","unstructured":"Ilya Loshchilov and Frank Hutter. 2015. Online batch selection for faster training of neural networks. arXiv preprint arXiv:1511.06343 (2015).","journal-title":"arXiv preprint arXiv:1511.06343"},{"key":"e_1_3_2_91_2","first-page":"7145","volume-title":"International Conference on Machine Learning (ICML)","volume":"139","author":"Lu Yucheng","year":"2021","unstructured":"Yucheng Lu, Youngsuk Park, Lifan Chen, Yuyang Wang, Christopher De Sa, and Dean Foster. 2021. Variance reduced training with stratified sampling for forecasting models. In International Conference on Machine Learning (ICML), Vol. 139. PMLR, 7145\u20137155."},{"key":"e_1_3_2_92_2","article-title":"Curriculum loss: Robust learning and generalization against label corruption","author":"Lyu Yueming","year":"2019","unstructured":"Yueming Lyu and Ivor W. Tsang. 2019. Curriculum loss: Robust learning and generalization against label corruption. arXiv preprint arXiv:1905.10045 (2019).","journal-title":"arXiv preprint arXiv:1905.10045"},{"key":"e_1_3_2_93_2","first-page":"960","article-title":"Decoupling \u201dwhen to update\u201d from \u201dhow to update\u201d","volume":"30","author":"Malach Eran","year":"2017","unstructured":"Eran Malach and Shai Shalev-Shwartz. 2017. Decoupling \u201dwhen to update\u201d from \u201dhow to update\u201d. Advances in Neural Information Processing Systems (NIPS) 30 (2017), 960\u2013970.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_94_2","unstructured":"Baharan Mirzasoleiman Ashwinkumar Badanidiyuru and Amin Karbasi. 2016. Fast constrained submodular maximization: Personalized data summarization. In International Conference on Machine Learning (ICML) Vol. 48. JMLR.org New York City NY 1358\u20131367."},{"key":"e_1_3_2_95_2","first-page":"11465","article-title":"Coresets for robust training of deep neural networks against noisy labels","volume":"33","author":"Mirzasoleiman Baharan","year":"2020","unstructured":"Baharan Mirzasoleiman, Kaidi Cao, and Jure Leskovec. 2020. Coresets for robust training of deep neural networks against noisy labels. Advances in Neural Information Processing Systems (NIPS) 33 (2020), 11465\u201311477.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_96_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v36i7.20755"},{"key":"e_1_3_2_97_2","doi-asserted-by":"publisher","DOI":"10.1007\/BF01588971"},{"key":"e_1_3_2_98_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-010-9156-z"},{"key":"e_1_3_2_99_2","doi-asserted-by":"publisher","DOI":"10.3115\/1118693.1118701"},{"key":"e_1_3_2_100_2","article-title":"Self: Learning to filter noisy labels with self-ensembling","author":"Nguyen Duc Tam","year":"2019","unstructured":"Duc Tam Nguyen, Chaithanya Kumar Mummadi, Thi Phuong Nhung Ngo, Thi Hoai Phuong Nguyen, Laura Beggel, and Thomas Brox. 2019. Self: Learning to filter noisy labels with self-ensembling. arXiv preprint arXiv:1910.01842 (2019).","journal-title":"arXiv preprint arXiv:1910.01842"},{"key":"e_1_3_2_101_2","doi-asserted-by":"publisher","DOI":"10.1145\/1015330.1015349"},{"key":"e_1_3_2_102_2","doi-asserted-by":"publisher","DOI":"10.1613\/jair.1.12125"},{"key":"e_1_3_2_103_2","article-title":"Learning with confident examples: Rank pruning for robust classification with noisy labels","author":"Northcutt Curtis G.","year":"2017","unstructured":"Curtis G. Northcutt, Tailin Wu, and Isaac L. Chuang. 2017. Learning with confident examples: Rank pruning for robust classification with noisy labels. arXiv preprint arXiv:1705.01936 (2017).","journal-title":"arXiv preprint arXiv:1705.01936"},{"key":"e_1_3_2_104_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-010-9165-y"},{"key":"e_1_3_2_105_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00077"},{"key":"e_1_3_2_106_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.01192"},{"key":"e_1_3_2_107_2","first-page":"20596","article-title":"Deep learning on a data diet: Finding important examples early in training","volume":"34","author":"Paul Mansheej","year":"2021","unstructured":"Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. 2021. Deep learning on a data diet: Finding important examples early in training. Advances in Neural Information Processing Systems (NIPS) 34 (2021), 20596\u201320607.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_108_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2019.2957003"},{"key":"e_1_3_2_109_2","first-page":"17044","article-title":"Identifying mislabeled data using the area under the margin ranking","volume":"33","author":"Pleiss Geoff","year":"2020","unstructured":"Geoff Pleiss, Tianyi Zhang, Ethan Elenberg, and Kilian Q. Weinberger. 2020. Identifying mislabeled data using the area under the margin ranking. Advances in Neural Information Processing Systems 33 (2020), 17044\u201317056.","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_3_2_110_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.01053"},{"key":"e_1_3_2_111_2","first-page":"17848","volume-title":"International Conference on Machine Learning (ICML)","volume":"162","author":"Pooladzandi Omead","year":"2022","unstructured":"Omead Pooladzandi, David Davini, and Baharan Mirzasoleiman. 2022. Adaptive second order coresets for data-efficient machine learning. In International Conference on Machine Learning (ICML), Vol. 162. PMLR, Baltimore, Maryland, USA, 17848\u201317869."},{"key":"e_1_3_2_112_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICIP.2017.8297020"},{"key":"e_1_3_2_113_2","first-page":"1","volume-title":"International Conference on Learning Representations","author":"Ravi Sachin","year":"2017","unstructured":"Sachin Ravi and Hugo Larochelle. 2017. Optimization as a model for few-shot learning. In International Conference on Learning Representations. OpenReview.net, Palais des Congr\u00e8s Neptune, Toulon, France, 1\u201311. https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=rJY0-Kcll"},{"key":"e_1_3_2_114_2","unstructured":"Sachin Ravi and Hugo Larochelle. 2018. Meta-learning for batch mode active learning. (2018). https:\/\/2.zoppoz.workers.dev:443\/https\/openreview.net\/forum?id=r1PsGFJPz"},{"key":"e_1_3_2_115_2","first-page":"4175","volume-title":"Advances in Neural Information Processing Systems (NIPS)","author":"Ren Jiawei","year":"2020","unstructured":"Jiawei Ren, Cunjun Yu, Shunan Sheng, Xiao Ma, Haiyu Zhao, Shuai Yi, and Hongsheng Li. 2020. Balanced meta-softmax for long-tailed visual recognition. In Advances in Neural Information Processing Systems (NIPS), Vol. 33. Curran Associates, Inc., 4175\u20134186."},{"key":"e_1_3_2_116_2","first-page":"4334","volume-title":"International Conference on Machine Learning (ICML)","volume":"80","author":"Ren Mengye","year":"2018","unstructured":"Mengye Ren, Wenyuan Zeng, Bin Yang, and Raquel Urtasun. 2018. Learning to reweight examples for robust deep learning. In International Conference on Machine Learning (ICML), Vol. 80. PMLR, Stockholm, Sweden, 4334\u20134343."},{"key":"e_1_3_2_117_2","doi-asserted-by":"publisher","DOI":"10.1145\/3472291"},{"key":"e_1_3_2_118_2","first-page":"2940","volume-title":"International Conference on Machine Learning (ICML)","volume":"70","author":"Ritter Samuel","year":"2017","unstructured":"Samuel Ritter, David G. T. Barrett, Adam Santoro, and Matt M. Botvinick. 2017. Cognitive psychology for deep neural networks: A shape bias case study. In International Conference on Machine Learning (ICML), Vol. 70. PMLR, Sydney, NSW, Australia, 2940\u20132949."},{"key":"e_1_3_2_119_2","volume-title":"Proceedings of MACLEAN: MAChine Learning for EArth ObservatioN Workshop co-located with the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML\/PKDD 2020)","volume":"2766","author":"Ruzicka V\u00edt","year":"2020","unstructured":"V\u00edt Ruzicka, Stefano D\u2019Aronco, Jan D. Wegner, and Konrad Schindler. 2020. Deep active learning in remote sensing for data efficient change detection. In Proceedings of MACLEAN: MAChine Learning for EArth ObservatioN Workshop co-located with the European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML\/PKDD 2020), Vol. 2766. RWTH Aachen University."},{"key":"e_1_3_2_120_2","first-page":"8346","volume-title":"International Conference on Machine Learning (ICML)","volume":"119","author":"Sagawa Shiori","year":"2020","unstructured":"Shiori Sagawa, Aditi Raghunathan, Pang Wei Koh, and Percy Liang. 2020. An investigation of why overparameterization exacerbates spurious correlations. In International Conference on Machine Learning (ICML), Vol. 119. PMLR, 8346\u20138356."},{"key":"e_1_3_2_121_2","doi-asserted-by":"publisher","DOI":"10.1007\/s10107-016-1030-6"},{"key":"e_1_3_2_122_2","doi-asserted-by":"publisher","DOI":"10.1145\/3381831"},{"key":"e_1_3_2_123_2","article-title":"Active learning for convolutional neural networks: A core-set approach","author":"Sener Ozan","year":"2017","unstructured":"Ozan Sener and Silvio Savarese. 2017. Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489 (2017).","journal-title":"arXiv preprint arXiv:1708.00489"},{"key":"e_1_3_2_124_2","article-title":"A geometric approach to active learning for convolutional neural networks","volume":"7","author":"Sener Ozan","year":"2017","unstructured":"Ozan Sener and Silvio Savarese. 2017. A geometric approach to active learning for convolutional neural networks. arXiv preprint arXiv:1708.00489 7 (2017).","journal-title":"arXiv preprint arXiv:1708.00489"},{"key":"e_1_3_2_125_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-031-01560-1"},{"key":"e_1_3_2_126_2","doi-asserted-by":"publisher","DOI":"10.1145\/130385.130417"},{"key":"e_1_3_2_127_2","first-page":"5739","volume-title":"International Conference on Machine Learning (ICML)","volume":"97","author":"Shen Yanyao","year":"2019","unstructured":"Yanyao Shen and Sujay Sanghavi. 2019. Learning with bad training data via iterative trimmed loss minimization. In International Conference on Machine Learning (ICML), Vol. 97. PMLR, Long Beach, California, USA, 5739\u20135748."},{"key":"e_1_3_2_128_2","doi-asserted-by":"publisher","DOI":"10.1186\/s40537-019-0197-0"},{"key":"e_1_3_2_129_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2016.89"},{"key":"e_1_3_2_130_2","first-page":"1917","article-title":"Meta-Weight-Net: Learning an explicit mapping for sample weighting","volume":"32","author":"Shu Jun","year":"2019","unstructured":"Jun Shu, Qi Xie, Lixuan Yi, Qian Zhao, Sanping Zhou, Zongben Xu, and Deyu Meng. 2019. Meta-Weight-Net: Learning an explicit mapping for sample weighting. Advances in Neural Information Processing Systems (NIPS) 32 (2019), 1917\u20131928.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_131_2","article-title":"CMW-Net: Learning a class-aware sample weighting mapping for robust deep learning","author":"Shu Jun","year":"2022","unstructured":"Jun Shu, Xiang Yuan, Deyu Meng, and Zongben Xu. 2022. CMW-Net: Learning a class-aware sample weighting mapping for robust deep learning. arXiv preprint arXiv:2202.05613 (2022).","journal-title":"arXiv preprint arXiv:2202.05613"},{"key":"e_1_3_2_132_2","first-page":"5907","volume-title":"International Conference on Machine Learning (ICML)","volume":"97","author":"Song Hwanjun","year":"2019","unstructured":"Hwanjun Song, Minseok Kim, and Jae-Gil Lee. 2019. Selfie: Refurbishing unclean samples for robust deep learning. In International Conference on Machine Learning (ICML), Vol. 97. PMLR, Long Beach, California, USA, 5907\u20135915."},{"key":"e_1_3_2_133_2","article-title":"How does early stopping help generalization against label noise?","author":"Song Hwanjun","year":"2019","unstructured":"Hwanjun Song, Minseok Kim, Dongmin Park, and Jae-Gil Lee. 2019. How does early stopping help generalization against label noise? arXiv preprint arXiv:1911.08059 (2019).","journal-title":"arXiv preprint arXiv:1911.08059"},{"key":"e_1_3_2_134_2","doi-asserted-by":"publisher","DOI":"10.1145\/3447548.3467222"},{"key":"e_1_3_2_135_2","first-page":"1","article-title":"Learning from noisy labels with deep neural networks: A survey","author":"Song Hwanjun","year":"2022","unstructured":"Hwanjun Song, Minseok Kim, Dongmin Park, Yooju Shin, and Jae-Gil Lee. 2022. Learning from noisy labels with deep neural networks: A survey. IEEE Transactions on Neural Networks and Learning Systems (2022), 1\u201319.","journal-title":"IEEE Transactions on Neural Networks and Learning Systems"},{"key":"e_1_3_2_136_2","article-title":"Energy and policy considerations for deep learning in NLP","author":"Strubell Emma","year":"2019","unstructured":"Emma Strubell, Ananya Ganesh, and Andrew McCallum. 2019. Energy and policy considerations for deep learning in NLP. arXiv preprint arXiv:1906.02243 (2019).","journal-title":"arXiv preprint arXiv:1906.02243"},{"key":"e_1_3_2_137_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.97"},{"key":"e_1_3_2_138_2","first-page":"9923","volume-title":"International Conference on Machine Learning (ICML)","volume":"139","author":"Sun Ming","year":"2021","unstructured":"Ming Sun, Haoxuan Dou, Baopu Li, Junjie Yan, Wanli Ouyang, and Lei Cui. 2021. AutoSampling: Search for effective data sampling schedules. In International Conference on Machine Learning (ICML), Vol. 139. PMLR, virtual, 9923\u20139933."},{"key":"e_1_3_2_139_2","doi-asserted-by":"publisher","DOI":"10.1145\/3394171.3413978"},{"key":"e_1_3_2_140_2","article-title":"Dataset cartography: Mapping and diagnosing datasets with training dynamics","author":"Swayamdipta Swabha","year":"2020","unstructured":"Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020. Dataset cartography: Mapping and diagnosing datasets with training dynamics. arXiv preprint arXiv:2009.10795 (2020).","journal-title":"arXiv preprint arXiv:2009.10795"},{"key":"e_1_3_2_141_2","first-page":"37","volume-title":"A Deeper Look at Dataset Bias","author":"Tommasi Tatiana","year":"2017","unstructured":"Tatiana Tommasi, Novi Patricia, Barbara Caputo, and Tinne Tuytelaars. 2017. A Deeper Look at Dataset Bias. Springer, Cham, 37\u201355."},{"key":"e_1_3_2_142_2","article-title":"An empirical study of example forgetting during deep neural network learning","author":"Toneva Mariya","year":"2018","unstructured":"Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J. Gordon. 2018. An empirical study of example forgetting during deep neural network learning. arXiv preprint arXiv:1812.05159 (2018).","journal-title":"arXiv preprint arXiv:1812.05159"},{"key":"e_1_3_2_143_2","doi-asserted-by":"publisher","DOI":"10.1023\/A:1023052124951"},{"key":"e_1_3_2_144_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2011.5995347"},{"issue":"4","key":"e_1_3_2_145_2","first-page":"363","article-title":"Core vector machines: Fast SVM training on very large data sets","volume":"6","author":"Tsang Ivor W.","year":"2005","unstructured":"Ivor W. Tsang, James T. Kwok, Pak-Ming Cheung, and Nello Cristianini. 2005. Core vector machines: Fast SVM training on very large data sets. Journal of Machine Learning Research 6, 4 (2005), 363\u2013392.","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_3_2_146_2","doi-asserted-by":"publisher","DOI":"10.1109\/RADIOELEKTRONIKA54537.2022.9764909"},{"key":"e_1_3_2_147_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR.2018.00914"},{"key":"e_1_3_2_148_2","doi-asserted-by":"publisher","DOI":"10.1109\/TNNLS.2017.2699783"},{"key":"e_1_3_2_149_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICTAI.2018.00017"},{"key":"e_1_3_2_150_2","doi-asserted-by":"publisher","DOI":"10.1007\/978-3-030-58568-6_43"},{"key":"e_1_3_2_151_2","article-title":"Dataset distillation","author":"Wang Tongzhou","year":"2018","unstructured":"Tongzhou Wang, Jun-Yan Zhu, Antonio Torralba, and Alexei A. Efros. 2018. Dataset distillation. arXiv preprint arXiv:1811.10959 (2018).","journal-title":"arXiv preprint arXiv:1811.10959"},{"issue":"9","key":"e_1_3_2_152_2","first-page":"4555","article-title":"A survey on curriculum learning","volume":"44","author":"Wang Xin","year":"2021","unstructured":"Xin Wang, Yudong Chen, and Wenwu Zhu. 2021. A survey on curriculum learning. IEEE Transactions on Pattern Analysis and Machine Intelligence 44, 9 (2021), 4555\u20134576.","journal-title":"IEEE Transactions on Pattern Analysis and Machine Intelligence"},{"key":"e_1_3_2_153_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2019.00512"},{"key":"e_1_3_2_154_2","first-page":"5139","article-title":"E2-Train: Training state-of-the-art CNNs with over 80% energy savings","volume":"32","author":"Wang Yue","year":"2019","unstructured":"Yue Wang, Ziyu Jiang, Xiaohan Chen, Pengfei Xu, Yang Zhao, Yingyan Lin, and Zhangyang Wang. 2019. E2-Train: Training state-of-the-art CNNs with over 80% energy savings. Advances in Neural Information Processing Systems (NIPS) 32 (2019), 5139\u20135151.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_155_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v34i04.6103"},{"key":"e_1_3_2_156_2","first-page":"1954","volume-title":"International Conference on Machine Learning (ICML)","volume":"37","author":"Wei Kai","year":"2015","unstructured":"Kai Wei, Rishabh Iyer, and Jeff Bilmes. 2015. Submodularity in data subset selection and active learning. In International Conference on Machine Learning (ICML), Vol. 37. PMLR, Lille, France, 1954\u20131963."},{"key":"e_1_3_2_157_2","first-page":"1","article-title":"Data collection and quality challenges in deep learning: A data-centric AI perspective","author":"Whang Steven Euijong","year":"2023","unstructured":"Steven Euijong Whang, Yuji Roh, Hwanjun Song, and Jae-Gil Lee. 2023. Data collection and quality challenges in deep learning: A data-centric AI perspective. The VLDB Journal (2023), 1\u201323.","journal-title":"The VLDB Journal"},{"key":"e_1_3_2_158_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00918"},{"key":"e_1_3_2_159_2","first-page":"21382","article-title":"A topological filter for learning with label noise","volume":"33","author":"Wu Pengxiang","year":"2020","unstructured":"Pengxiang Wu, Songzhu Zheng, Mayank Goswami, Dimitris Metaxas, and Chao Chen. 2020. A topological filter for learning with label noise. Advances in Neural Information Processing Systems (NIPS) 33 (2020), 21382\u201321393.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_160_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00013"},{"key":"e_1_3_2_161_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00740"},{"key":"e_1_3_2_162_2","doi-asserted-by":"publisher","DOI":"10.1016\/j.patcog.2018.12.027"},{"key":"e_1_3_2_163_2","doi-asserted-by":"publisher","DOI":"10.1145\/2964284.2967213"},{"key":"e_1_3_2_164_2","article-title":"Online coreset selection for rehearsal-based continual learning","author":"Yoon Jaehong","year":"2021","unstructured":"Jaehong Yoon, Divyam Madaan, Eunho Yang, and Sung Ju Hwang. 2021. Online coreset selection for rehearsal-based continual learning. arXiv preprint arXiv:2106.01085 (2021).","journal-title":"arXiv preprint arXiv:2106.01085"},{"key":"e_1_3_2_165_2","doi-asserted-by":"publisher","DOI":"10.1145\/3225058.3225069"},{"key":"e_1_3_2_166_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR52688.2022.00017"},{"key":"e_1_3_2_167_2","first-page":"7164","volume-title":"International Conference on Machine Learning (ICML)","volume":"97","author":"Yu Xingrui","year":"2019","unstructured":"Xingrui Yu, Bo Han, Jiangchao Yao, Gang Niu, Ivor Tsang, and Masashi Sugiyama. 2019. How does disagreement help generalization against label corruption?. In International Conference on Machine Learning (ICML), Vol. 97. PMLR, Long Beach, California, USA, 7164\u20137173."},{"key":"e_1_3_2_168_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00344"},{"key":"e_1_3_2_169_2","doi-asserted-by":"publisher","DOI":"10.1145\/3446776"},{"key":"e_1_3_2_170_2","doi-asserted-by":"publisher","DOI":"10.1609\/aaai.v35i4.16444"},{"key":"e_1_3_2_171_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00953"},{"key":"e_1_3_2_172_2","doi-asserted-by":"publisher","DOI":"10.1109\/CVPR46437.2021.00239"},{"key":"e_1_3_2_173_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV.2017.578"},{"key":"e_1_3_2_174_2","doi-asserted-by":"publisher","DOI":"10.1109\/ICCV48922.2021.00786"},{"key":"e_1_3_2_175_2","article-title":"Deep long-tailed learning: A survey","author":"Zhang Yifan","year":"2021","unstructured":"Yifan Zhang, Bingyi Kang, Bryan Hooi, Shuicheng Yan, and Jiashi Feng. 2021. Deep long-tailed learning: A survey. arXiv preprint arXiv:2110.04596 (2021).","journal-title":"arXiv preprint arXiv:2110.04596"},{"issue":"2","key":"e_1_3_2_176_2","first-page":"3","article-title":"Dataset condensation with gradient matching","volume":"1","author":"Zhao Bo","year":"2021","unstructured":"Bo Zhao, Konda Reddy Mopuri, and Hakan Bilen. 2021. Dataset condensation with gradient matching. 9th International Conference on Learning Representations, ICLR 2021, Virtual Event 1, 2 (2021), 3.","journal-title":"9th International Conference on Learning Representations, ICLR 2021, Virtual Event"},{"key":"e_1_3_2_177_2","article-title":"Accelerating minibatch stochastic gradient descent using stratified sampling","author":"Zhao Peilin","year":"2014","unstructured":"Peilin Zhao and Tong Zhang. 2014. Accelerating minibatch stochastic gradient descent using stratified sampling. arXiv preprint arXiv:1405.3080 (2014).","journal-title":"arXiv preprint arXiv:1405.3080"},{"key":"e_1_3_2_178_2","article-title":"Coverage-centric coreset selection for high pruning rates","author":"Zheng Haizhong","year":"2022","unstructured":"Haizhong Zheng, Rui Liu, Fan Lai, and Atul Prakash. 2022. Coverage-centric coreset selection for high pruning rates. arXiv preprint arXiv:2210.15809 (2022).","journal-title":"arXiv preprint arXiv:2210.15809"},{"key":"e_1_3_2_179_2","first-page":"8602","article-title":"Curriculum learning by dynamic instance hardness","volume":"33","author":"Zhou Tianyi","year":"2020","unstructured":"Tianyi Zhou, Shengjie Wang, and Jeffrey Bilmes. 2020. Curriculum learning by dynamic instance hardness. Advances in Neural Information Processing Systems (NIPS) 33 (2020), 8602\u20138613.","journal-title":"Advances in Neural Information Processing Systems (NIPS)"},{"key":"e_1_3_2_180_2","first-page":"27287","volume-title":"International Conference on Machine Learning (ICML)","author":"Zhou Xiao","year":"2022","unstructured":"Xiao Zhou, Renjie Pi, Weizhong Zhang, Yong Lin, Zonghao Chen, and Tong Zhang. 2022. Probabilistic bilevel coreset selection. In International Conference on Machine Learning (ICML). PMLR, Baltimore, Maryland, USA, 27287\u201327302."}],"container-title":["ACM Computing Surveys"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3627157","content-type":"unspecified","content-version":"vor","intended-application":"text-mining"},{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/pdf\/10.1145\/3627157","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,6,18]],"date-time":"2025-06-18T23:57:05Z","timestamp":1750291025000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/dl.acm.org\/doi\/10.1145\/3627157"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,4,9]]},"references-count":179,"journal-issue":{"issue":"7","published-print":{"date-parts":[[2024,7,31]]}},"alternative-id":["10.1145\/3627157"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1145\/3627157","relation":{},"ISSN":["0360-0300","1557-7341"],"issn-type":[{"value":"0360-0300","type":"print"},{"value":"1557-7341","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,4,9]]},"assertion":[{"value":"2022-10-14","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2023-09-26","order":1,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-04-09","order":2,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}