aalto1 untyped-item.component.html
Trust-region variational autoencoders for multi-agent reinforcement learning
Loading...
URL
Journal Title
Journal ISSN
Volume Title
School of Electrical Engineering |
Master's thesis
Unless otherwise stated, all rights belong to the author. You may download, display and print this publication for Your own personal use. Commercial use is prohibited.
Authors
Date
Department
Major/Subject
Mcode
Degree programme
Language
en
Pages
56
Series
Abstract
Learning effective policies from high-dimensional observations such as images remains a challenge in multi-agent reinforcement learning (MARL), particularly due to sample inefficiency and representational instability. While recent advances in single-agent settings have explored representation learning through generative and contrastive methods, their extension to multi-agent domains remains limited. This thesis introduces a novel framework that integrates a variational autoencoder (VAE) with a trust-region constraint, called multi-agent trust-region variational autoencoder (MA-TRVAE), into the MARL training pipeline. The proposed approach enables stable and efficient learning of latent representations from pixel-based observations, which are shared across agents for policy optimization. By constraining the encoder updates within a trust region, the method mitigates representational drift during training and improves overall policy stability. We evaluate the framework on vision-based cooperative tasks within a multi-agent quadcopter control (MAQC) environment and compare it with the baselines. Experimental results demonstrate that MA-TRVAE achieves improved performance, scalability, and robustness in representation learning, providing a promising direction for integrating trust-region-constrained VAEs in multi-agent systems.