aalto1 untyped-item.component.html
Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation
Loading...
Access rights
openAccess
CC BY
CC BY
Creative Commons license
Except where otherwised noted, this item's license is described as openAccess
publishedVersion
URL
Journal Title
Journal ISSN
Volume Title
A4 Artikkeli konferenssijulkaisussa
This publication is imported from Aalto University research portal.
View publication in the Research portal (opens in new window)
View/Open full text file from the Research portal (opens in new window)
Other link related to publication (opens in new window)
View publication in the Research portal (opens in new window)
View/Open full text file from the Research portal (opens in new window)
Other link related to publication (opens in new window)
Unless otherwise stated, all rights belong to the author. You may download, display and print this publication for Your own personal use. Commercial use is prohibited.
Date
Major/Subject
Mcode
Degree programme
Language
en
Pages
18
Series
Proceedings of Machine Learning Research, Volume 270, pp. 4314-4331
Abstract
In vision-based behaviour cloning (BC), conventional image augmentations like Random Crop and Colour Jitter often fall short when addressing substantial visual domain shifts, such as variations in shadow, distractors and backgrounds. Superimposition-based augmentations, which blend in-domain and out-of-domain images, have shown promise for improving model generalisation in the computer vision community, but their suitability for BC remains uncertain due to the need to preserve task-critical semantics, spatial-temporal relationships, and agent-target interactions. To address this, we introduce RoboSaGA-a Saliency-Guided Augmentation method within the superimposition family, tailored for vision-based BC. RoboSaGA dynamically adjusts augmentation intensity per pixel based on policy-driven saliency, enabling aggressive augmentation in task-trivial areas while preserving task-critical information. Moreover, it integrates seamlessly into existing architectures without requiring structural changes or additional learning objectives. Empirical evaluations in both simulated and real-world settings show that RoboSaGA maintains in-domain performance while significantly enhancing robustness to visual domain shifts, including distractors and background variations, as well as handling lighting and shadow variations. Code available at: https://github.com/Zheyu-Zhuang/RoboSaGA.
Description
Other note
Citation
Zhuang, Z, Wang, R, Ingelhag, N, Kyrki, V & Kragic, D 2025, 'Enhancing Visual Domain Robustness in Behaviour Cloning via Saliency-Guided Augmentation', Proceedings of Machine Learning Research, vol. 270, pp. 4314-4331. < https://proceedings.mlr.press/v270/zhuang25b.html >
