Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving

Samak, Chinmay Vilas; Samak, Tanmay Vilas; Li, Bing; Krovi, Venkat

doi:10.1109/LRA.2025.3632723

Computer Science > Robotics

arXiv:2507.00236 (cs)

[Submitted on 30 Jun 2025 (v1), last revised 31 Oct 2025 (this version, v3)]

Title:Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving

Authors:Chinmay Vilas Samak, Tanmay Vilas Samak, Bing Li, Venkat Krovi

View PDF HTML (experimental)

Abstract:Simulation-based design, optimization, and validation of autonomous vehicles have proven to be crucial for their improvement over the years. Nevertheless, the ultimate measure of effectiveness is their successful transition from simulation to reality (sim2real). However, existing sim2real transfer methods struggle to address the autonomy-oriented requirements of balancing: (i) conditioned domain adaptation, (ii) robust performance with limited examples, (iii) modularity in handling multiple domain representations, and (iv) real-time performance. To alleviate these pain points, we present a unified framework for learning cross-domain adaptive representations through conditional latent diffusion for sim2real transferable automated driving. Our framework offers options to leverage: (i) alternate foundation models, (ii) a few-shot fine-tuning pipeline, and (iii) textual as well as image prompts for mapping across given source and target domains. It is also capable of generating diverse high-quality samples when diffusing across parameter spaces such as times of day, weather conditions, seasons, and operational design domains. We systematically analyze the presented framework and report our findings in terms of performance benchmarks and ablation studies. Additionally, we demonstrate its serviceability for autonomous driving using behavioral cloning case studies. Our experiments indicate that the proposed framework is capable of bridging the perceptual sim2real gap by over 40%.

Comments:	Accepted in IEEE Robotics and Automation Letters (RA-L)
Subjects:	Robotics (cs.RO)
Cite as:	arXiv:2507.00236 [cs.RO]
	(or arXiv:2507.00236v3 [cs.RO] for this version)
	https://doi.org/10.48550/arXiv.2507.00236
Journal reference:	IEEE Robotics and Automation Letters, vol. 11, no. 1, pp. 177-184, Jan. 2026
Related DOI:	https://doi.org/10.1109/LRA.2025.3632723

Submission history

From: Chinmay Samak [view email]
[v1] Mon, 30 Jun 2025 20:07:35 UTC (6,011 KB)
[v2] Tue, 15 Jul 2025 04:11:30 UTC (6,011 KB)
[v3] Fri, 31 Oct 2025 03:13:58 UTC (4,347 KB)

Computer Science > Robotics

Title:Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Robotics

Title:Sim2Real Diffusion: Leveraging Foundation Vision Language Models for Adaptive Automated Driving

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators