{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,7,20]],"date-time":"2026-07-20T23:11:47Z","timestamp":1784589107191,"version":"3.55.0"},"reference-count":116,"publisher":"Wiley","issue":"5","license":[{"start":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T00:00:00Z","timestamp":1720742400000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/http\/onlinelibrary.wiley.com\/termsAndConditions#vor"}],"content-domain":{"domain":["bera-journals.onlinelibrary.wiley.com"],"crossmark-restriction":true},"short-container-title":["Brit J Educational Tech"],"published-print":{"date-parts":[[2024,9]]},"abstract":"<jats:sec>\n                    <jats:title>Abstract<\/jats:title>\n                    <jats:p>Large language models (LLMs) are increasingly adopted in educational contexts to provide personalized support to students and teachers. The unprecedented capacity of LLM\u2010based applications to understand and generate natural language can potentially improve instructional effectiveness and learning outcomes, but the integration of LLMs in education technology has renewed concerns over algorithmic bias, which may exacerbate educational inequalities. Building on prior work that mapped the traditional machine learning life cycle, we provide a framework of the LLM life cycle from the initial development of LLMs to customizing pre\u2010trained models for various applications in educational settings. We explain each step in the LLM life cycle and identify potential sources of bias that may arise in the context of education. We discuss why current measures of bias from traditional machine learning fail to transfer to LLM\u2010generated text (eg, tutoring conversations) because text encodings are high\u2010dimensional, there can be multiple correct responses, and tailoring responses may be pedagogically desirable rather than unfair. The proposed framework clarifies the complex nature of bias in LLM applications and provides practical guidance for their evaluation to promote educational equity.<\/jats:p>\n                  <\/jats:sec>\n                  <jats:sec>\n                    <jats:label\/>\n                    <jats:p>\n                      <jats:boxed-text content-type=\"box\" position=\"anchor\">\n                        <jats:caption>\n                          <jats:title>Practitioner notes<\/jats:title>\n                        <\/jats:caption>\n                        <jats:p>\n                          What is already known about this topic\n                          <jats:list list-type=\"bullet\">\n                            <jats:list-item>\n                              <jats:p>The life cycle of traditional machine learning (ML) applications which focus on predicting labels is well understood.<\/jats:p>\n                            <\/jats:list-item>\n                            <jats:list-item>\n                              <jats:p>Biases are known to enter in traditional ML applications at various points in the life cycle, and methods to measure and mitigate these biases have been developed and tested.<\/jats:p>\n                            <\/jats:list-item>\n                            <jats:list-item>\n                              <jats:p>Large language models (LLMs) and other forms of generative artificial intelligence (GenAI) are increasingly adopted in education technologies (EdTech), but current evaluation approaches are not specific to the domain of education.<\/jats:p>\n                            <\/jats:list-item>\n                          <\/jats:list>\n                        <\/jats:p>\n                        <jats:p>\n                          What this paper adds\n                          <jats:list list-type=\"bullet\">\n                            <jats:list-item>\n                              <jats:p>A holistic perspective of the LLM life cycle with domain\u2010specific examples in education to highlight opportunities and challenges for incorporating natural language understanding (NLU) and natural language generation (NLG) into EdTech.<\/jats:p>\n                            <\/jats:list-item>\n                            <jats:list-item>\n                              <jats:p>Potential sources of bias are identified in each step of the LLM life cycle and discussed in the context of education.<\/jats:p>\n                            <\/jats:list-item>\n                            <jats:list-item>\n                              <jats:p>A framework for understanding where to expect potential harms of LLMs for students, teachers, and other users of GenAI technology in education, which can guide approaches to bias measurement and mitigation.<\/jats:p>\n                            <\/jats:list-item>\n                          <\/jats:list>\n                        <\/jats:p>\n                        <jats:p>\n                          Implications for practice and\/or policy\n                          <jats:list list-type=\"bullet\">\n                            <jats:list-item>\n                              <jats:p>Education practitioners and policymakers should be aware that biases can originate from a multitude of steps in the LLM life cycle, and the life cycle perspective offers them a heuristic for asking technology developers to explain each step to assess the risk of bias.<\/jats:p>\n                            <\/jats:list-item>\n                            <jats:list-item>\n                              <jats:p>Measuring the biases of systems that use LLMs in education is more complex than with traditional ML, in large part because the evaluation of natural language generation is highly context\u2010dependent (eg, what counts as good feedback on an assignment varies).<\/jats:p>\n                            <\/jats:list-item>\n                            <jats:list-item>\n                              <jats:p>EdTech developers can play an important role in collecting and curating datasets for the evaluation and benchmarking of LLM applications moving forward.<\/jats:p>\n                            <\/jats:list-item>\n                          <\/jats:list>\n                        <\/jats:p>\n                      <\/jats:boxed-text>\n                    <\/jats:p>\n                  <\/jats:sec>","DOI":"10.1111\/bjet.13505","type":"journal-article","created":{"date-parts":[[2024,7,12]],"date-time":"2024-07-12T08:25:39Z","timestamp":1720772739000},"page":"1982-2002","update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1002\/crossmark_policy","source":"Crossref","is-referenced-by-count":89,"title":["The life cycle of large language models in education: A framework for understanding sources of bias"],"prefix":"10.1111","volume":"55","author":[{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-9957-1342","authenticated-orcid":false,"given":"Jinsook","family":"Lee","sequence":"first","affiliation":[{"name":"Department of Information Science Cornell University  Ithaca New York USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-7234-7001","authenticated-orcid":false,"given":"Yann","family":"Hicke","sequence":"additional","affiliation":[{"name":"Department of Computer Science Cornell University  Ithaca New York USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0002-2375-3537","authenticated-orcid":false,"given":"Renzhe","family":"Yu","sequence":"additional","affiliation":[{"name":"Teachers College and Data Science Institute Columbia University  New York New York USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0003-0875-0204","authenticated-orcid":false,"given":"Christopher","family":"Brooks","sequence":"additional","affiliation":[{"name":"School of Information University of Michigan  Ann Arbor Michigan USA"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"ORCID":"https:\/\/2.zoppoz.workers.dev:443\/https\/orcid.org\/0000-0001-6283-5546","authenticated-orcid":false,"given":"Ren\u00e9 F.","family":"Kizilcec","sequence":"additional","affiliation":[{"name":"Department of Information Science Cornell University  Ithaca New York USA"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"311","published-online":{"date-parts":[[2024,7,12]]},"reference":[{"key":"e_1_2_11_2_1","unstructured":"Almazrouei E. Alobeidli H. Alshamsi A. Cappelli A. Cojocaru R. Debbah M. Goffinet \u00c9. Hesslow D. Launay J. Malartic Q. Mazzotta D. Noune B. Pannier B. &Penedo G.(2023).The falcon series of open language models.arXiv preprint arXiv:2311.16867."},{"key":"e_1_2_11_3_1","unstructured":"Anil R. Dai A. M. Firat O. Johnson M. Lepikhin D. Passos A. Shakeri S. Taropa E. Bailey P. Chen Z. Chu E. Clark J. H. El Shafey L. Huang Y. Meier\u2010Hellstern K. Mishra G. Moreira E. Omernick M. Robinson K. \u2026Wu Y.(2023).PaLM 2 technical report.arXiv preprint arXiv:2305.10403."},{"key":"e_1_2_11_4_1","unstructured":"Anthis J. R. Lum K. Ekstrand M. Feller A. D'Amour A. &Tan C.(2024).The impossibility of fair LLMs.arXiv e\u2010prints arXiv\u20132406."},{"key":"e_1_2_11_5_1","unstructured":"Attri. (2023).A comprehensive guide: Everything you need to know about LLMs' guardrails.https:\/\/2.zoppoz.workers.dev:443\/https\/attri.ai\/blog\/a\u2010comprehensive\u2010guide\u2010everything\u2010you\u2010need\u2010to\u2010know\u2010about\u2010llms\u2010guardrails"},{"key":"e_1_2_11_6_1","unstructured":"Bai J. Bai S. Chu Y. Cui Z. Dang K. Deng X. Fan Y. Ge W. Han Y. Huang F. Hui B. Ji L. Li M. Lin J. Lin R. Liu D. Liu G. Lu C. Lu K. \u2026Zhu T.(2023).Qwen technical report.https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2309.16609"},{"key":"e_1_2_11_7_1","unstructured":"Bai Y. Kadavath S. Kundu S. Askell A. Kernion J. Jones A. Chen A. Goldie A. Mirhoseini A. McKinnon C. Chen C. Olsson C. Olah C. Hernandez D. Drain D. Ganguli D. Li D. Tran\u2010Johnson E. Perez E. \u2026Kaplan J.(2022).Constitutional AI: Harmlessness from AI feedback.arXiv preprint arXiv:2212.08073."},{"key":"e_1_2_11_8_1","doi-asserted-by":"publisher","DOI":"10.1007\/s40593-021-00285-9"},{"key":"e_1_2_11_9_1","first-page":"671","article-title":"Big data's disparate impact","volume":"104","author":"Barocas S.","year":"2016","journal-title":"California Law Review"},{"key":"e_1_2_11_10_1","doi-asserted-by":"publisher","DOI":"10.1145\/3593013.3594007"},{"key":"e_1_2_11_11_1","doi-asserted-by":"publisher","DOI":"10.1038\/d41586-023-02361-7"},{"key":"e_1_2_11_12_1","unstructured":"BigScience Workshop Scao T. L. Fan A. Akiki C. Pavlick E. Ili\u0107 S. Hesslow D. Castagn\u00e9 R. Luccioni A. S. Yvon F. Gall\u00e9 M. Tow J. Rush A. M. Biderman S. Webson A. Ammanamanchi P. S. Wang T. Sagot B. Muennighoff N. \u2026Wolf T.(2023).Bloom: A 176b\u2010parameter open\u2010access multilingual language model."},{"key":"e_1_2_11_13_1","first-page":"21268","volume-title":"Advances in neural information processing systems","author":"Birhane A.","year":"2023"},{"key":"e_1_2_11_14_1","unstructured":"Birhane A. Prabhu V. U. &Kahembwe E.(2021).Multimodal datasets: Misogyny pornography and malignant stereotypes.arXiv preprint arXiv:2110.01963."},{"key":"e_1_2_11_15_1","doi-asserted-by":"publisher","DOI":"10.3102\/0013189X013006004"},{"key":"e_1_2_11_16_1","first-page":"4349","article-title":"Man is to computer programmer as woman is to homemaker? Debiasing word embeddings","volume":"29","author":"Bolukbasi T.","year":"2016","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_17_1","doi-asserted-by":"publisher","DOI":"10.1111\/nyas.15007"},{"key":"e_1_2_11_18_1","doi-asserted-by":"crossref","unstructured":"Bordia S. &Bowman S. R.(2019).Identifying and reducing gender bias in word\u2010level language models. InNorth American chapter of the association for computational linguistics.https:\/\/2.zoppoz.workers.dev:443\/https\/api.semanticscholar.org\/CorpusID:102352788","DOI":"10.18653\/v1\/N19-3002"},{"key":"e_1_2_11_19_1","first-page":"1877","volume-title":"Advances in neural information processing systems","author":"Brown T.","year":"2020"},{"key":"e_1_2_11_20_1","volume-title":"Proceedings of the 36th International Conference on Machine Learning","author":"Brunet M.\u2010E.","year":"2019"},{"key":"e_1_2_11_21_1","unstructured":"Bubeck S. Chandrasekaran V. Eldan R. Gehrke J. Horvitz E. Kamar E. Lee P. Lee Y. T. Li Y. Lundberg S. Nori H. Palangi H. Ribeiro M. T. &Zhang Y.(2023).Sparks of artificial general intelligence: Early experiments with GPT\u20104.arXiv preprint arXiv:2303.12712."},{"key":"e_1_2_11_22_1","doi-asserted-by":"publisher","DOI":"10.1126\/science.aal4230"},{"key":"e_1_2_11_23_1","unstructured":"Casper S. Davies X. Shi C. Gilbert T. K. Scheurer J. Rando J. Freedman R. Korbak T. Lindner D. Freire P. Wang T. Marks S. Segerie C.\u2010R. Carroll M. Peng A. Christoffersen P. Damani M. Slocum S. Anwar U. \u2026Hadfield\u2010Menell D.(2023).Open problems and fundamental limitations of reinforcement learning from human feedback.arXiv preprint arXiv:2307.15217."},{"issue":"70","key":"e_1_2_11_24_1","first-page":"1","article-title":"Scaling instruction\u2010finetuned language models","volume":"25","author":"Chung H. W.","year":"2024","journal-title":"Journal of Machine Learning Research"},{"key":"e_1_2_11_25_1","unstructured":"Coursera. (2023).New products tools and features.https:\/\/2.zoppoz.workers.dev:443\/https\/blog.coursera.org\/new\u2010products\u2010tools\u2010and\u2010features\u20102023"},{"key":"e_1_2_11_26_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.naacl-main.122"},{"key":"e_1_2_11_27_1","doi-asserted-by":"publisher","DOI":"10.3102\/01623737231169270"},{"key":"e_1_2_11_28_1","unstructured":"Denny P. Gulwani S. Heffernan N. T. K\u00e4ser T. Moore S. Rafferty A. N. &Singla A.(2024).Generative AI for education (GAIED): Advances opportunities and challenges.arXiv preprint arXiv:2402.01580."},{"key":"e_1_2_11_29_1","unstructured":"Devlin J. Chang M.\u2010W. Lee K. &Toutanova K.(2019).Bert: Pre\u2010training of deep bidirectional transformers for language understanding. InNorth American Chapter of the Association for Computational Linguistics.https:\/\/2.zoppoz.workers.dev:443\/https\/api.semanticscholar.org\/CorpusID:52967399"},{"key":"e_1_2_11_30_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445924"},{"key":"e_1_2_11_31_1","doi-asserted-by":"publisher","DOI":"10.4324\/9780429433726-31"},{"key":"e_1_2_11_32_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41539\u2010023\u201000174\u2010x"},{"key":"e_1_2_11_33_1","doi-asserted-by":"publisher","DOI":"10.1111\/bjso.12560"},{"key":"e_1_2_11_34_1","unstructured":"edX Press. (n.d.).edX Debuts Two AI\u2010Powered Learning Assistants Built on ChatGPT.https:\/\/2.zoppoz.workers.dev:443\/https\/press.edx.org\/edx\u2010debuts\u2010two\u2010ai\u2010powered\u2010learning\u2010assistants\u2010built\u2010on\u2010chatgpt"},{"key":"e_1_2_11_35_1","doi-asserted-by":"publisher","DOI":"10.1016\/S1364-6613(99)01294-2"},{"key":"e_1_2_11_36_1","doi-asserted-by":"publisher","DOI":"10.1162\/coli_a_00524"},{"key":"e_1_2_11_37_1","unstructured":"Ganguli D. Lovitt L. Kernion J. Askell A. Bai Y. Kadavath S. Mann B. Perez E. Schiefer N. Ndousse K. Jones A. Bowman S. Chen A. Conerly T. DasSarma N. Drain D. Elhage N. El\u2010Showk S. Fort S. \u2026Clark J.(2022).Red teaming language models to reduce harms: Methods scaling behaviors and lessons learned.arXiv preprint arXiv: 2209.07858."},{"key":"e_1_2_11_38_1","doi-asserted-by":"publisher","DOI":"10.1073\/pnas.1720347115"},{"key":"e_1_2_11_39_1","unstructured":"Gemini Team Anil R. Borgeaud S. Alayrac J.\u2010B. Yu J. Soricut R. Schalkwyk J. Dai A. M. Hauth A. Millican K. Silver D. Johnson M. Antonoglou I. Schrittwieser J. Glaese A. Chen J. Pitler E. Lillicrap T. Lazaridou A. \u2026Vinyals O.(2024).Gemini: A family of highly capable multimodal models.arXiv preprint arXiv:2312.11805."},{"key":"e_1_2_11_40_1","unstructured":"Gemma Team Mesnard T. Hardin C. Dadashi R. Bhupatiraju S. Pathak S. Sifre L. Rivi\u00e8re M. Kale M. S. Love J. Tafti P. Hussenot L. Sessa P. G. Chowdhery A. Roberts A. Barua A. Botev A. Castro\u2010Ros A. Slone A. \u2026Kenealy K.(2024).Gemma: Open models based on Gemini research and technology. arXiv preprint arXiv:2403.08295."},{"key":"e_1_2_11_41_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2021.acl-long.150"},{"key":"e_1_2_11_42_1","doi-asserted-by":"crossref","unstructured":"Gonen H. &Goldberg Y.(2019).Lipstick on a pig: Debiasing methods cover up systematic gender biases in word embeddings but do not remove them.arXiv preprint arXiv:1903.03862.","DOI":"10.18653\/v1\/N19-1061"},{"key":"e_1_2_11_43_1","unstructured":"Google Jigsaw. (2024).Perspective API documentation.https:\/\/2.zoppoz.workers.dev:443\/https\/perspectiveapi.com\/"},{"key":"e_1_2_11_44_1","doi-asserted-by":"crossref","unstructured":"Henkel O. Horne\u2010Robinson H. Kozhakhmetova N. &Lee A.(2024).Effective and scalable math support: Evidence on the impact of an AI\u2010 tutor on math achievement in Ghana.arXiv preprint arXiv:2402.09809.","DOI":"10.1007\/978-3-031-64315-6_34"},{"key":"e_1_2_11_45_1","unstructured":"Hicke Y. Agarwal A. Ma Q. &Denny P.(2023).Chata: Towards an intelligent question\u2010answer teaching assistant using open\u2010source LLMs.arXiv preprint arXiv:2311.02775."},{"key":"e_1_2_11_46_1","unstructured":"Hofmann V. Kalluri P. R. Jurafsky D. &King S.(2024).Dialect prejudice predicts AI decisions about people's character employability and criminality.arXiv preprint arXiv: 2403.00742."},{"key":"e_1_2_11_47_1","doi-asserted-by":"publisher","DOI":"10.1038\/s41598-023-30938-9"},{"key":"e_1_2_11_48_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533089"},{"key":"e_1_2_11_49_1","unstructured":"Inan H. Upasani K. Chi J. Rungta R. Iyer K. Mao Y. Tontchev M. Hu Q. Fuller B. Testuggine D. &Khabsa M.(2023).Llama Guard: LLM\u2010based input\u2010output safeguard for human\u2010AI conversations.arXiv preprint arXiv:2312.06674."},{"key":"e_1_2_11_50_1","doi-asserted-by":"publisher","DOI":"10.1145\/3442188.3445901"},{"key":"e_1_2_11_51_1","doi-asserted-by":"publisher","DOI":"10.1080\/10705511.2014.856694"},{"key":"e_1_2_11_52_1","unstructured":"Jiang A. Q. Sablayrolles A. Mensch A. Bamford C. Chaplot D. S. de lasCasas D. Bressand F. Lengyel G. Lample G. Saulnier L. Renard Lavaud L. Lachaux M.\u2010A. Stock P. Le Scao T. Lavril T. Wang T. Lacroix T. &El Sayed W.(2023).Mistral 7B.arXiv preprint arXiv: 2310.06825."},{"key":"e_1_2_11_53_1","unstructured":"Jurenka I. Kunesch M. McKee K. Gillick D. Zhu S. Wiltberger S. Phal S. M. Hermann K. Kasenberg D. Bhoopchand A. Anand A. P\u00eeslar M. Chan S. Wang L. She J. Mahmoudieh P. Rysbek A. Ko W.\u2010J. Huber A. \u2026Ibrahim L.(2024).Towards responsible development of generative AI for education: An evaluation\u2010driven approach. Google Technical Report.https:\/\/2.zoppoz.workers.dev:443\/https\/storage.googleapis.com\/deepmind\u2010media\/LearnLM\/LearnLM_paper.pdf"},{"key":"e_1_2_11_54_1","volume-title":"Noise: A flaw in human judgment","author":"Kahneman D.","year":"2021"},{"key":"e_1_2_11_55_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.lindif.2023.102274"},{"key":"e_1_2_11_56_1","unstructured":"Khan Academy. (n.d.).Khan Academy Labs.https:\/\/2.zoppoz.workers.dev:443\/https\/www.khanacademy.org\/khan\u2010labs"},{"key":"e_1_2_11_57_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.findings-emnlp.35"},{"key":"e_1_2_11_58_1","doi-asserted-by":"publisher","DOI":"10.4324\/9780429329067-10"},{"key":"e_1_2_11_59_1","doi-asserted-by":"publisher","DOI":"10.1145\/3582269.3615599"},{"key":"e_1_2_11_60_1","doi-asserted-by":"crossref","unstructured":"Kurita K. Vyas N. Pareek A. Black A. W. &Tsvetkov Y.(2019).Measuring bias in contextualized word representations.arXiv preprint arXiv:1906.07337.","DOI":"10.18653\/v1\/W19-3823"},{"key":"e_1_2_11_61_1","first-page":"1","article-title":"Bridging large language model disparities: Skill tagging of multilingual educational content","author":"Kwak Y.","year":"2024","journal-title":"British Journal of Educational Technology"},{"key":"e_1_2_11_62_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.caeai.2024.100210"},{"key":"e_1_2_11_63_1","unstructured":"LDNOOBW. (2023).List of dirty naughty obscene and otherwise bad words.https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/LDNOOBW\/List\u2010of\u2010Dirty\u2010Naughty\u2010Obscene\u2010and\u2010Otherwise\u2010Bad\u2010Words"},{"key":"e_1_2_11_64_1","unstructured":"Leiker D. Finnigan S. Gyllen A. R. &Cukurova M.(2023).Prototyping the use of large language models (LLMs) for adult learning content creation at scale.arXiv preprint arXiv:2306.01815."},{"key":"e_1_2_11_65_1","volume-title":"Proceedings of the 15th International Conference on educational data mining, International Educational Data Mining Society","author":"Levin N.","year":"2022"},{"key":"e_1_2_11_66_1","first-page":"9459","volume-title":"Advances in neural information processing systems","author":"Lewis P.","year":"2020"},{"key":"e_1_2_11_67_1","doi-asserted-by":"publisher","DOI":"10.1145\/3636555.3636860"},{"key":"e_1_2_11_68_1","unstructured":"Li Y. Bubeck S. Eldan R. Del Giorno A. Gunasekar S. &Lee Y. T.(2023).Textbooks are all you need ii: phi\u20101.5 technical report.arXiv preprint arXiv:2309.05463."},{"key":"e_1_2_11_69_1","unstructured":"Liang P. Bommasani R. Lee T. Tsipras D. Soylu D. Yasunaga M. Zhang Y. Narayanan D. Wu Y. Kumar A. Newman B. Yuan B. Yan B. Zhang C. Cosgrove C. Manning C. D. R\u00e9 C. Acosta\u2010Navas D. Hudson D. A. \u2026Koreeda Y.(2023).Holistic evaluation of language models.arXiv preprint arXiv:2211.09110."},{"key":"e_1_2_11_70_1","unstructured":"Lin J. Thomas D. R. Han F. Gupta S. Tan W. Nguyen N. D. &Koedinger K. R.(2023).Using large language models to provide explanatory feedback to human tutors.arXiv preprint arXiv:2306.15498."},{"key":"e_1_2_11_71_1","doi-asserted-by":"publisher","DOI":"10.1126\/sciadv.adg9405"},{"key":"e_1_2_11_72_1","doi-asserted-by":"publisher","DOI":"10.1145\/3560815"},{"key":"e_1_2_11_73_1","doi-asserted-by":"crossref","unstructured":"Liu X. Ji K. Fu Y. Tam W. L. Du Z. Yang Z. &Tang J.(2021).P\u2010tuning v2: Prompt tuning can be comparable to fine\u2010tuning universally across scales and tasks.arXiv preprint arXiv:2110.07602.","DOI":"10.18653\/v1\/2022.acl-short.8"},{"key":"e_1_2_11_74_1","unstructured":"Liu Y. Ott M. Goyal N. Du J. Joshi M. Chen D. Levy O. Lewis M. Zettlemoyer L. &Stoyanov V.(2019).Roberta: A robustly optimized BERT pretraining approach.arXiv preprint arXiv:1907.11692."},{"key":"e_1_2_11_75_1","volume-title":"The effects of virtual tutoring on young readers: Results from a randomized controlled trial","author":"Loeb S.","year":"2023"},{"key":"e_1_2_11_76_1","unstructured":"Lozhkov A. Ben Allal L. vonWerra L. &Wolf T.(2024 May).Fineweb\u2010edu.https:\/\/2.zoppoz.workers.dev:443\/https\/huggingface.co\/datasets\/HuggingFaceFW\/fineweb\u2010edu"},{"key":"e_1_2_11_77_1","unstructured":"Luo Y. Yang Z. Meng F. Li Y. Zhou J. &Zhang Y.(2023).An empirical study of catastrophic forgetting in large language models during continual fine\u2010tuning.arXiv preprint arXiv:2308.08747."},{"key":"e_1_2_11_78_1","doi-asserted-by":"crossref","unstructured":"May C. Wang A. Bordia S. Bowman S. R. &Rudinger R.(2019).On measuring social biases in sentence encoders.arXiv preprint arXiv:1903.10561.","DOI":"10.18653\/v1\/N19-1063"},{"key":"e_1_2_11_79_1","doi-asserted-by":"publisher","DOI":"10.1145\/3457607"},{"key":"e_1_2_11_80_1","doi-asserted-by":"publisher","DOI":"10.1016\/j.caeai.2023.100199"},{"key":"e_1_2_11_81_1","unstructured":"Minaee S. Mikolov T. Nikzad N. Chenaghlu M. Socher R. Amatriain X. &Gao J.(2024).Large language models: A survey.arXiv preprint arXiv:2402.06196."},{"key":"e_1_2_11_82_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2022.ltedi-1.4"},{"key":"e_1_2_11_83_1","doi-asserted-by":"publisher","DOI":"10.1111\/bjet.12156"},{"key":"e_1_2_11_84_1","doi-asserted-by":"publisher","DOI":"10.1145\/2207676.2208597"},{"key":"e_1_2_11_85_1","unstructured":"OpenAI. (2023).Gpt\u20104 technical report.https:\/\/2.zoppoz.workers.dev:443\/https\/arxiv.org\/abs\/2303.08774"},{"key":"e_1_2_11_86_1","doi-asserted-by":"crossref","unstructured":"Pankiewicz M. &Baker R. S.(2024).Navigating compiler errors with AI assistance\u2014A study of GPT hints in an introductory programming course.arXiv preprint arXiv:2403.12737.","DOI":"10.1145\/3649217.3653608"},{"key":"e_1_2_11_87_1","volume-title":"Conversational agents and natural language interaction: Techniques and effective practices: Techniques and effective practices","author":"Perez\u2010Marin D.","year":"2011"},{"key":"e_1_2_11_88_1","unstructured":"Radford A. &Narasimhan K.(2018).Improving language understanding by generative pre\u2010training.https:\/\/2.zoppoz.workers.dev:443\/https\/api.semanticscholar.org\/CorpusID:49313245"},{"key":"e_1_2_11_89_1","volume-title":"Improving language understanding by generative pre\u2010training","author":"Radford A.","year":"2018"},{"key":"e_1_2_11_90_1","unstructured":"Radford A. Wu J. Child R. Luan D. Amodei D. &Sutskever I.(2019).Language models are unsupervised multitask learners.https:\/\/2.zoppoz.workers.dev:443\/https\/api.semanticscholar.org\/CorpusID:160025533"},{"key":"e_1_2_11_91_1","first-page":"53728","article-title":"Direct preference optimization: Your language model is secretly a reward model","volume":"36","author":"Rafailov R.","year":"2024","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_92_1","doi-asserted-by":"crossref","unstructured":"Rajpurkar P. Zhang J. Lopyrev K. &Liang P.(2016).Squad: 100 000+ questions for machine comprehension of text.arXiv preprint arXiv:1606.05250.","DOI":"10.18653\/v1\/D16-1264"},{"key":"e_1_2_11_93_1","doi-asserted-by":"publisher","DOI":"10.1007\/s10462-021-10068-2"},{"key":"e_1_2_11_94_1","doi-asserted-by":"publisher","DOI":"10.18653\/v1\/2023.acl-long.163"},{"key":"e_1_2_11_95_1","first-page":"5861","article-title":"Process for adapting language models to society (palms) with values\u2010targeted datasets","volume":"34","author":"Solaiman I.","year":"2021","journal-title":"Advances in Neural Information Processing Systems"},{"key":"e_1_2_11_96_1","unstructured":"spamscanner. (2023).Spam scanner: A node.js anti\u2010spam email filtering and phishing prevention tool and service.https:\/\/2.zoppoz.workers.dev:443\/https\/github.com\/spamscanner\/spamscanner"},{"key":"e_1_2_11_97_1","doi-asserted-by":"publisher","DOI":"10.1145\/3465416.3483305"},{"key":"e_1_2_11_98_1","doi-asserted-by":"publisher","DOI":"10.1145\/3306618.3314270"},{"key":"e_1_2_11_99_1","doi-asserted-by":"crossref","unstructured":"Tao Y. Viberg O. Baker R. S. &Kizilcec R. F.(2024).Cultural bias and cultural alignment of large language models.arXiv preprint arXiv:2311.14096.","DOI":"10.1093\/pnasnexus\/pgae346"},{"key":"e_1_2_11_100_1","unstructured":"Team G. Mesnard T. Hardin C. Dadashi R. Bhupatiraju S. Pathak S. Sifre L. Rivi\u00e8re M. Kale M. S. Love J. Tafti P. Hussenot L. Sessa P. G. Chowdhery A. Roberts A. Barua A. Botev A. Castro\u2010Ros A. Slone A. \u2026Kenealy K.(2024).Gemma: Open models based on Gemini research and technology.arXiv preprint arXiv:2403.08295."},{"key":"e_1_2_11_101_1","unstructured":"Touvron H. Lavril T. Izacard G. Martinet X. Lachaux M.\u2010A. Lacroix T. Rozi\u00e8re B. Goyal N. Hambro E. Azhar F. Rodriguez A. Joulin A. Grave E. &Lample G.(2023).LLaMA: Open and efficient foundation language models.arXiv preprint arXiv:2302.13971."},{"key":"e_1_2_11_102_1","unstructured":"Wang A. Morgenstern J. &Dickerson J. P.(2024).Large language models cannot replace human participants because they cannot portray identity groups.arXiv preprint arXiv:2402.01908."},{"key":"e_1_2_11_103_1","doi-asserted-by":"crossref","unstructured":"Wang R. E. &Demszky D.(2024).Edu\u2010ConvoKit: An open\u2010source library for education conversation data.arXiv preprint arXiv:2402.05111.","DOI":"10.18653\/v1\/2024.naacl-demo.6"},{"key":"e_1_2_11_104_1","unstructured":"Wang R. E. Zhang Q. Robinson C. Loeb S. &Demszky D.(2023).Step\u2010by\u2010step remediation of students' mathematical mistakes.arXiv preprint arXiv:2310.10648."},{"key":"e_1_2_11_105_1","unstructured":"Webster K. Wang X. Tenney I. Beutel A. Pitler E. Pavlick E. Chen J. Chi E. &Petrov S.(2020).Measuring and reducing gendered correlations in pre\u2010trained models.arXiv preprint arXiv:2010.06032."},{"key":"e_1_2_11_106_1","unstructured":"Weidinger L. Mellor J. F. J. Rauh M. Griffin C. Uesato J. Huang P.\u2010S. Cheng M. Glaese M. Balle B. Kasirzadeh A. Kenton Z. Brown S. Hawkins W. Stepleton T. Biles C. Birhane A. Haas J. Rimell L. Hendricks L. A. \u2026Gabriel I.(2021).Ethical and social risks of harm from language models.ArXiv abs\/2112.04359.https:\/\/2.zoppoz.workers.dev:443\/https\/api.semanticscholar.org\/CorpusID:244954639"},{"key":"e_1_2_11_107_1","doi-asserted-by":"publisher","DOI":"10.1145\/3531146.3533088"},{"key":"e_1_2_11_108_1","unstructured":"Weights & Biases. (2023).Processing data for large language models.https:\/\/2.zoppoz.workers.dev:443\/https\/wandb.ai\/wandb_gen\/llm\u2010data\u2010processing\/reports\/Processing\u2010Data\u2010for\u2010Large\u2010Language\u2010Models\u2010\u2010VmlldzozMDg4MTM2"},{"key":"e_1_2_11_109_1","doi-asserted-by":"publisher","DOI":"10.3389\/frai.2021.654924"},{"key":"e_1_2_11_110_1","doi-asserted-by":"publisher","DOI":"10.1111\/bjet.13370"},{"key":"e_1_2_11_111_1","doi-asserted-by":"publisher","DOI":"10.1177\/2378023117737206"},{"key":"e_1_2_11_112_1","unstructured":"Zhai Y. Tong S. Li X. Cai M. Qu Q. Lee Y. J. &Ma Y.(2023).Investigating the catastrophic forgetting in multimodal large language models.arXiv preprint arXiv:2309.10313."},{"key":"e_1_2_11_113_1","doi-asserted-by":"crossref","unstructured":"Zhao J. Wang T. Yatskar M. Cotterell R. Ordonez V. &Chang K.\u2010W.(2019).Gender bias in contextualized word embeddings.arXiv preprint arXiv:1904.03310.","DOI":"10.18653\/v1\/N19-1064"},{"key":"e_1_2_11_114_1","first-page":"12697","volume-title":"International conference on machine learning","author":"Zhao Z.","year":"2021"},{"key":"e_1_2_11_115_1","unstructured":"Zheng H. Shen L. Tang A. Luo Y. Hu H. Du B. &Tao D.(2023).Learn from model beyond fine\u2010tuning: A survey.arXiv preprint arXiv:2310.08184."},{"key":"e_1_2_11_116_1","doi-asserted-by":"publisher","DOI":"10.1111\/bjet.13156"},{"key":"e_1_2_11_117_1","unstructured":"Zhou Y. Zanette A. Pan J. Levine S. &Kumar A.(2024).Archer: Training language model agents via hierarchical multi\u2010turn RL.arXiv preprint arXiv:2402.19446."}],"container-title":["British Journal of Educational Technology"],"original-title":[],"language":"en","link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/bera-journals.onlinelibrary.wiley.com\/doi\/pdf\/10.1111\/bjet.13505","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2025,10,29]],"date-time":"2025-10-29T13:58:55Z","timestamp":1761746335000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/bera-journals.onlinelibrary.wiley.com\/doi\/10.1111\/bjet.13505"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,12]]},"references-count":116,"journal-issue":{"issue":"5","published-print":{"date-parts":[[2024,9]]}},"alternative-id":["10.1111\/bjet.13505"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.1111\/bjet.13505","archive":["Portico"],"relation":{},"ISSN":["0007-1013","1467-8535"],"issn-type":[{"value":"0007-1013","type":"print"},{"value":"1467-8535","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,12]]},"assertion":[{"value":"2024-06-03","order":0,"name":"received","label":"Received","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-06-26","order":2,"name":"accepted","label":"Accepted","group":{"name":"publication_history","label":"Publication History"}},{"value":"2024-07-12","order":3,"name":"published","label":"Published","group":{"name":"publication_history","label":"Publication History"}}]}}