{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,1,28]],"date-time":"2026-01-28T21:19:18Z","timestamp":1769635158440,"version":"3.49.0"},"reference-count":64,"publisher":"Frontiers Media SA","license":[{"start":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T00:00:00Z","timestamp":1720051200000},"content-version":"vor","delay-in-days":0,"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/creativecommons.org\/licenses\/by\/4.0\/"}],"funder":[{"DOI":"10.13039\/100000145","name":"Division of Information and Intelligent Systems","doi-asserted-by":"publisher","award":["IIS-2202395"],"award-info":[{"award-number":["IIS-2202395"]}],"id":[{"id":"10.13039\/100000145","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000145","name":"Division of Information and Intelligent Systems","doi-asserted-by":"publisher","award":["IIS-2311969"],"award-info":[{"award-number":["IIS-2311969"]}],"id":[{"id":"10.13039\/100000145","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100006754","name":"Army Research Laboratory","doi-asserted-by":"publisher","award":["W911NF2320179"],"award-info":[{"award-number":["W911NF2320179"]}],"id":[{"id":"10.13039\/100006754","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000183","name":"Army Research Office","doi-asserted-by":"publisher","award":["W911NF2110299"],"award-info":[{"award-number":["W911NF2110299"]}],"id":[{"id":"10.13039\/100000183","id-type":"DOI","asserted-by":"publisher"}]},{"DOI":"10.13039\/100000006","name":"Office of Naval Research","doi-asserted-by":"publisher","award":["N00014-23-1-2850"],"award-info":[{"award-number":["N00014-23-1-2850"]}],"id":[{"id":"10.13039\/100000006","id-type":"DOI","asserted-by":"publisher"}]}],"content-domain":{"domain":["frontiersin.org"],"crossmark-restriction":true},"short-container-title":["Front. Artif. Intell."],"abstract":"<jats:p>This study focuses on a rescue mission problem, particularly enabling agents\/robots to navigate efficiently in unknown environments. Technological advances, including manufacturing, sensing, and communication systems, have raised interest in using robots or drones for rescue operations. Effective rescue operations require quick identification of changes in the environment and\/or locating the victims\/injuries as soon as possible. Several techniques have been developed in recent years for autonomy in rescue missions, including motion planning, adaptive control, and more recently, reinforcement learning techniques. These techniques rely on full knowledge of the environment or the availability of simulators that can represent real environments during rescue operations. However, in practice, agents might have little or no information about the environment or the number or locations of injuries, preventing\/limiting the application of most existing techniques. This study provides a probabilistic\/Bayesian representation of the unknown environment, which jointly models the stochasticity in the agent's navigation and the environment uncertainty into a vector called the belief state. This belief state allows offline learning of the optimal Bayesian policy in an unknown environment without the need for any real data\/interactions, which guarantees taking actions that are optimal given all available information. To address the large size of belief space, deep reinforcement learning is developed for computing an approximate Bayesian planning policy. The numerical experiments using different maze problems demonstrate the high performance of the proposed policy.<\/jats:p>","DOI":"10.3389\/frai.2024.1308031","type":"journal-article","created":{"date-parts":[[2024,7,4]],"date-time":"2024-07-04T04:46:28Z","timestamp":1720068388000},"update-policy":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.3389\/crossmark-policy","source":"Crossref","is-referenced-by-count":5,"title":["Bayesian reinforcement learning for navigation planning in unknown environments"],"prefix":"10.3389","volume":"7","author":[{"given":"Mohammad","family":"Alali","sequence":"first","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]},{"given":"Mahdi","family":"Imani","sequence":"additional","affiliation":[],"role":[{"role":"author","vocabulary":"crossref"}]}],"member":"1965","published-online":{"date-parts":[[2024,7,4]]},"reference":[{"key":"B1","doi-asserted-by":"crossref","DOI":"10.2514\/6.2019-1979","article-title":"A simulation-based development and verification architecture for micro UAV teams and swarms","author":"Akcakoca","year":"2019","journal-title":"AIAA Scitech 2019 Forum"},{"key":"B2","doi-asserted-by":"crossref","first-page":"3957","DOI":"10.23919\/ACC55779.2023.10155867","article-title":"Reinforcement learning data-acquiring for causal inference of regulatory networks","volume-title":"2023 American Control Conference (ACC)","author":"Alali","year":"2023"},{"key":"B3","doi-asserted-by":"publisher","DOI":"10.1109\/TCBB.2024.3402220","article-title":"Bayesian lookahead perturbation policy for inference of regulatory networks","author":"Alali","year":"2024","journal-title":"IEEE\/ACM Trans. Comput. Biol. Bioinform"},{"key":"B4","doi-asserted-by":"publisher","first-page":"2329260","DOI":"10.1080\/21642583.2024.2329260","article-title":"Deep reinforcement learning sensor scheduling for effective monitoring of dynamical systems","volume":"12","author":"Alali","year":"2024","journal-title":"Syst. Sci. Cont. Eng"},{"key":"B5","doi-asserted-by":"crossref","DOI":"10.1061\/9780784485514.035","article-title":"Privacy-preserved federated reinforcement learning for autonomy in signalized intersections","volume-title":"ASCE International Conference on Transportation and Development (ICTD)","author":"Asadi","year":"2024"},{"key":"B6","doi-asserted-by":"crossref","first-page":"1758","DOI":"10.1109\/CDC40024.2019.9030133","article-title":"An efficient reachability-based framework for provably safe autonomous navigation in unknown environments","volume-title":"2019 IEEE 58th Conference on Decision and Control (CDC)","author":"Bajcsy","year":"2019"},{"key":"B7","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2009.09595","article-title":"Rl star platform: Reinforcement learning for simulation based training of robots","author":"Blum","year":"2020","journal-title":"arXiv [Preprint]"},{"key":"B8","doi-asserted-by":"crossref","first-page":"523","DOI":"10.1109\/ICUAS.2019.8798254","article-title":"Deep reinforcement learning attitude control of fixed-wing UAVs using proximal policy optimization","volume-title":"2019 International Conference on Unmanned Aircraft Systems (ICUAS)","author":"B\u00f8hn","year":"2019"},{"key":"B9","doi-asserted-by":"publisher","first-page":"103673","DOI":"10.1016\/j.robot.2020.103673","article-title":"A novel UAV path planning algorithm to search for floating objects on the ocean surface based on object's trajectory prediction by regression","volume":"135","author":"Boulares","year":"2021","journal-title":"Rob. Auton. Syst"},{"key":"B10","doi-asserted-by":"crossref","first-page":"2518","DOI":"10.1109\/IROS45743.2020.9341361","article-title":"Autonomous spot: Long-range autonomous exploration of extreme environments with legged locomotion","volume-title":"2020 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Bouman","year":"2020"},{"key":"B11","doi-asserted-by":"publisher","first-page":"4","DOI":"10.3390\/drones3010004","article-title":"Survey on coverage path planning with unmanned aerial vehicles","volume":"3","author":"Cabreira","year":"2019","journal-title":"Drones"},{"key":"B12","doi-asserted-by":"publisher","first-page":"51","DOI":"10.1007\/s10514-020-09947-4","article-title":"Reinforcement based mobile robot path planning with improved dynamic window approach in unknown environment","volume":"45","author":"Chang","year":"2021","journal-title":"Auton. Robots"},{"key":"B13","doi-asserted-by":"crossref","first-page":"33","DOI":"10.1007\/978-3-030-28619-4_5","article-title":"A bayesian active learning approach to adaptive motion planning","author":"Choudhury","year":"2020","journal-title":"Robotics Research"},{"key":"B14","doi-asserted-by":"publisher","first-page":"32","DOI":"10.1016\/j.robot.2018.11.005","article-title":"Bio-inspired on-line path planner for cooperative exploration of unknown environment by a multi-robot system","volume":"112","author":"de Almeida","year":"2019","journal-title":"Rob. Auton. Syst"},{"key":"B15","doi-asserted-by":"publisher","first-page":"1312","DOI":"10.1109\/TMC.2020.2966989","article-title":"Autonomous UAV trajectory for localizing ground objects: A reinforcement learning approach","volume":"20","author":"Ebrahimi","year":"2020","journal-title":"IEEE Trans. Mobile Comp"},{"key":"B16","doi-asserted-by":"publisher","first-page":"102517","DOI":"10.1016\/j.rcim.2022.102517","article-title":"A review on reinforcement learning for contact-rich robotic manipulation tasks","volume":"81","author":"Elguea-Aguinaco","year":"2023","journal-title":"Robot. Comput. Integr. Manuf"},{"key":"B17","doi-asserted-by":"crossref","DOI":"10.2514\/6.2022-2497","article-title":"Deep reinforcement learning for autonomous aerobraking maneuver planning","author":"Falcone","year":"2022","journal-title":"AIAA SCITECH 2022 Forum"},{"key":"B18","doi-asserted-by":"crossref","first-page":"10820","DOI":"10.1109\/IROS47612.2022.9982175","article-title":"Bayesian active learning for sim-to-real robotic perception","volume-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Feng","year":"2022"},{"key":"B19","doi-asserted-by":"publisher","first-page":"359","DOI":"10.1561\/2200000049","article-title":"Bayesian reinforcement learning: a survey","volume":"8","author":"Ghavamzadeh","year":"2015","journal-title":"Found. Trends Mach. Learn"},{"key":"B20","doi-asserted-by":"publisher","first-page":"1681","DOI":"10.1007\/s10514-019-09829-4","article-title":"Reinforcement learning and model predictive control for robust embedded quadrotor guidance and control","volume":"43","author":"Greatwood","year":"2019","journal-title":"Auton. Robots"},{"key":"B21","article-title":"Efficient Bayes-adaptive reinforcement learning using sample-based search","volume-title":"Advances in Neural Information Processing Systems 25","author":"Guez","year":"2012"},{"key":"B22","doi-asserted-by":"publisher","first-page":"645","DOI":"10.22581\/muet1982.2103.17","article-title":"Reinforcement learning based hierarchical multi-agent robotic search team in uncertain environment","volume":"40","author":"Hamid","year":"2021","journal-title":"Mehran Univer. Res. J. Eng. Technol"},{"key":"B23","doi-asserted-by":"publisher","first-page":"14413","DOI":"10.1109\/TVT.2020.3034800","article-title":"Voronoi-based multi-robot autonomous exploration in unknown environments via deep reinforcement learning","volume":"69","author":"Hu","year":"2020","journal-title":"IEEE Trans. Vehicul. Technol"},{"key":"B24","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/ICARCV.2016.7838739","article-title":"Autonomous navigation of UAV by using real-time model-based reinforcement learning","volume-title":"2016 14th International Conference on Control, Automation, Robotics and Vision (ICARCV)","author":"Imanberdiyev","year":"2016"},{"key":"B25","doi-asserted-by":"publisher","first-page":"4125","DOI":"10.1109\/TNNLS.2021.3051012","article-title":"Scalable inverse reinforcement learning through multifidelity bayesian optimization","volume":"33","author":"Imani","year":"2022","journal-title":"IEEE Trans. Neural Netw. Learn. Syst"},{"key":"B26","article-title":"Bayesian control of large MDPs with unknown dynamics in data-poor environments","volume-title":"Advances in Neural Information Processing Systems","author":"Imani","year":"2018"},{"key":"B27","doi-asserted-by":"crossref","first-page":"25","DOI":"10.1007\/978-3-030-77939-9_2","volume-title":"Deep Learning and Reinforcement Learning for Autonomous Unmanned Aerial Systems: Roadmap for Theory to Deployment","author":"Jagannath","year":"2021"},{"key":"B28","doi-asserted-by":"publisher","first-page":"379","DOI":"10.1109\/TFUZZ.2011.2104364","article-title":"Evolutionary-group-based particle-swarm-optimized fuzzy controller with application to mobile-robot navigation in unknown environments","volume":"19","author":"Juang","year":"2011","journal-title":"IEEE Transact. Fuzzy Syst"},{"key":"B29","first-page":"1701","article-title":"Data-efficient reinforcement learning with probabilistic model predictive control","volume-title":"International Conference on Artificial Intelligence and Statistics","author":"Kamthe","year":"2018"},{"key":"B30","doi-asserted-by":"publisher","first-page":"323","DOI":"10.1007\/s00422-018-0753-2","article-title":"Planning and navigation as active inference","volume":"112","author":"Kaplan","year":"2018","journal-title":"Biol. Cybern"},{"key":"B31","doi-asserted-by":"crossref","first-page":"652","DOI":"10.1609\/icaps.v31i1.16014","article-title":"Plgrim: Hierarchical value learning for large-scale exploration in unknown environments","author":"Kim","year":"2021","journal-title":"Proceedings of the International Conference on Automated Planning and Scheduling"},{"key":"B32","doi-asserted-by":"publisher","first-page":"2493","DOI":"10.1109\/LRA.2019.2903259","article-title":"Bi-directional value learning for risk-aware planning under uncertainty","volume":"4","author":"Kim","year":"2019","journal-title":"IEEE Robot. Automat. Lett"},{"key":"B33","article-title":"Adam: a method for stochastic optimization","author":"Kingma","year":"2015","journal-title":"CoRR"},{"key":"B34","doi-asserted-by":"publisher","first-page":"8","DOI":"10.2478\/jaiscr-2019-0008","article-title":"Collision-free autonomous robot navigation in unknown environments utilizing pso for path planning","volume":"9","author":"Krell","year":"2019","journal-title":"J. Artif. Intell. Soft Comput. Res"},{"key":"B35","doi-asserted-by":"publisher","first-page":"2370","DOI":"10.1109\/LRA.2019.2903850","article-title":"A hybrid approach of learning and model-based channel prediction for communication relay UAVs in dynamic urban environments","volume":"4","author":"Ladosz","year":"2019","journal-title":"IEEE Robot. Automat. Lett"},{"key":"B36","doi-asserted-by":"publisher","first-page":"2064","DOI":"10.1109\/TNNLS.2019.2927869","article-title":"Deep reinforcement learning-based automatic exploration for navigation in unknown environment","volume":"31","author":"Li","year":"2019","journal-title":"IEEE Trans. Neural Netw. Learn. Syst"},{"key":"B37","doi-asserted-by":"crossref","first-page":"2493","DOI":"10.1109\/ICMA.2019.8816208","article-title":"End-to-end decentralized multi-robot navigation in unknown complex environments via deep reinforcement learning","volume-title":"2019 IEEE International Conference on Mechatronics and Automation (ICMA)","author":"Lin","year":"2019"},{"key":"B38","doi-asserted-by":"crossref","DOI":"10.1109\/CCTA60707.2024.10666622","article-title":"High-level human intention learning for cooperative decision-making","volume-title":"IEEE Conference on Control Technology and Applications (CCTA)","author":"Lin","year":"2024"},{"key":"B39","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/SIS.2014.7011782","article-title":"Sensor-based autonomous robot navigation under unknown environments with grid map representation","volume-title":"2014 IEEE Symposium on Swarm Intelligence","author":"Luo","year":"2014"},{"key":"B40","doi-asserted-by":"publisher","first-page":"529","DOI":"10.1038\/nature14236","article-title":"Human-level control through deep reinforcement learning","volume":"518","author":"Mnih","year":"2015","journal-title":"Nature"},{"key":"B41","doi-asserted-by":"publisher","first-page":"610","DOI":"10.1109\/LRA.2019.2891991","article-title":"Deep reinforcement learning robot for search and rescue applications: exploration in unknown cluttered environments","volume":"4","author":"Niroui","year":"2019","journal-title":"IEEE Robot. Automat. Lett"},{"key":"B42","doi-asserted-by":"crossref","first-page":"1374","DOI":"10.1109\/COASE.2016.7743569","article-title":"Multi-robot 3d coverage path planning for first responders teams","volume-title":"2016 IEEE International Conference on Automation Science and Engineering (CASE)","author":"Perez-Imaz","year":"2016"},{"key":"B43","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/SSRR.2018.8468611","article-title":"Reinforcement learning for autonomous UAV navigation using function approximation","volume-title":"2018 IEEE International Symposium on Safety, Security, and Rescue Robotics (SSRR)","author":"Pham","year":"2018"},{"key":"B44","doi-asserted-by":"publisher","first-page":"1027","DOI":"10.1109\/LCSYS.2022.3229054","article-title":"Optimal recursive expert-enabled inference in regulatory networks","volume":"7","author":"Ravari","year":"2023","journal-title":"IEEE Cont. Syst. Lett"},{"key":"B45","doi-asserted-by":"crossref","DOI":"10.23919\/ACC60939.2024.10644975","article-title":"Implicit human perception learning in complex and unknown environments","volume-title":"American Control Conference (ACC)","author":"Ravari","year":""},{"key":"B46","doi-asserted-by":"crossref","DOI":"10.1109\/TAI.2024.3358261","article-title":"Optimal inference of hidden Markov models through expert-acquired data","volume-title":"IEEE Transactions on Artificial Intelligence","author":"Ravari","year":""},{"key":"B47","doi-asserted-by":"crossref","first-page":"11932","DOI":"10.1109\/IROS47612.2022.9981738","article-title":"Informative path planning for active learning in aerial semantic mapping","volume-title":"2022 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Rckin","year":"2022"},{"key":"B48","doi-asserted-by":"publisher","first-page":"4279","DOI":"10.1109\/TRO.2023.3313811","article-title":"An informative path planning framework for active learning in uav-based semantic mapping","volume":"39","author":"Rckin","year":"2023","journal-title":"IEEE Trans. Robot"},{"key":"B49","first-page":"387","article-title":"Learning to plan for visibility in navigation of unknown environments","volume-title":"International Symposium on Experimental Robotics","author":"Richter","year":"2017"},{"key":"B50","doi-asserted-by":"publisher","first-page":"1142","DOI":"10.48550\/arXiv.2102.05762","article-title":"Risk-averse Bayes-adaptive reinforcement learning","volume":"34","author":"Rigter","year":"2021","journal-title":"Adv. Neural Inf. Process. Syst"},{"key":"B51","doi-asserted-by":"crossref","DOI":"10.1109\/TAC.2022.3210871","article-title":"The mixed-observable constrained linear quadratic regulator problem: the exact solution and practical algorithms","volume-title":"IEEE Transactions on Automatic Control","author":"Rosolia","year":"2022"},{"key":"B52","doi-asserted-by":"publisher","DOI":"10.48550\/arXiv.2403.16644","article-title":"Bridging the sim-to-real gap with Bayesian inference","author":"Rothfuss","year":"2024","journal-title":"arXiv [Preprint]"},{"key":"B53","doi-asserted-by":"publisher","first-page":"601","DOI":"10.1007\/s10846-018-0898-1","article-title":"A fully-autonomous aerial robot for search and rescue applications in indoor environments using learning-based techniques","volume":"95","author":"Sampedro","year":"2019","journal-title":"J. Intellig. Robot. Syst"},{"key":"B54","doi-asserted-by":"crossref","first-page":"200","DOI":"10.1109\/ICRA.2012.6224757","article-title":"Active learning from demonstration for robust autonomous navigation","volume-title":"2012 IEEE International Conference on Robotics and Automation","author":"Silver","year":"2012"},{"key":"B55","doi-asserted-by":"publisher","first-page":"102576","DOI":"10.1016\/j.mechatronics.2021.102576","article-title":"Active learning in robotics: a review of control principles","volume":"77","author":"Taylor","year":"2021","journal-title":"Mechatronics"},{"key":"B56","doi-asserted-by":"crossref","first-page":"1934","DOI":"10.1109\/IROS40897.2019.8968021","article-title":"Faster: Fast and safe trajectory planner for flights in unknown environments","volume-title":"2019 IEEE\/RSJ International Conference on Intelligent Robots and Systems (IROS)","author":"Tordesillas","year":"2019"},{"key":"B57","doi-asserted-by":"publisher","first-page":"2124","DOI":"10.1109\/TVT.2018.2890773","article-title":"Autonomous navigation of UAVs in large-scale complex environments: a deep reinforcement learning approach","volume":"68","author":"Wang","year":"2019","journal-title":"IEEE Trans. Vehic. Technol"},{"key":"B58","doi-asserted-by":"publisher","first-page":"6807","DOI":"10.1109\/TITS.2021.3062500","article-title":"Reinforcement learning and particle swarm optimization supporting real-time rescue assignments for multiple autonomous underwater vehicles","volume":"23","author":"Wu","year":"2021","journal-title":"IEEE Trans. Intellig. Transp. Syst"},{"key":"B59","doi-asserted-by":"crossref","first-page":"538","DOI":"10.1109\/IAEAC47372.2019.8998066","article-title":"Autonomous decision-making method for combat mission of UAV based on deep reinforcement learning","volume-title":"2019 IEEE 4th Advanced Information Technology, Electronic and Automation Control Conference (IAEAC)","author":"Xu","year":"2019"},{"key":"B60","doi-asserted-by":"crossref","first-page":"1","DOI":"10.1109\/SSRR.2018.8468643","article-title":"Robot navigation of environments with unknown rough terrain using deep reinforcement learning","volume-title":"2018 IEEE International Symposium on Safety, Security, and Rescue Robotics (SSRR)","author":"Zhang","year":"2018"},{"key":"B61","doi-asserted-by":"publisher","first-page":"172","DOI":"10.1016\/j.neucom.2012.09.019","article-title":"Robot path planning in uncertain environment using multi-objective particle swarm optimization","volume":"103","author":"Zhang","year":"2013","journal-title":"Neurocomputing"},{"key":"B62","article-title":"Modeling other players with Bayesian beliefs for games with incomplete information","author":"Zhang","year":"","journal-title":"arXiv"},{"key":"B63","article-title":"Collaborative AI teaming in unknown environments via active goal deduction","author":"Zhang","year":"","journal-title":"arXiv"},{"key":"B64","doi-asserted-by":"publisher","first-page":"1","DOI":"10.1145\/3580305.3599254","article-title":"VariBAD: a very good method for Bayes-adaptive deep RL via meta-learning","volume":"22","author":"Zintgraf","year":"2021","journal-title":"J. Mach. Learn. Res"}],"container-title":["Frontiers in Artificial Intelligence"],"original-title":[],"link":[{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.frontiersin.org\/articles\/10.3389\/frai.2024.1308031\/full","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2024,11,23]],"date-time":"2024-11-23T11:52:54Z","timestamp":1732362774000},"score":1,"resource":{"primary":{"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/www.frontiersin.org\/articles\/10.3389\/frai.2024.1308031\/full"}},"subtitle":[],"short-title":[],"issued":{"date-parts":[[2024,7,4]]},"references-count":64,"alternative-id":["10.3389\/frai.2024.1308031"],"URL":"https:\/\/2.zoppoz.workers.dev:443\/https\/doi.org\/10.3389\/frai.2024.1308031","relation":{},"ISSN":["2624-8212"],"issn-type":[{"value":"2624-8212","type":"electronic"}],"subject":[],"published":{"date-parts":[[2024,7,4]]},"article-number":"1308031"}}