aalto1 untyped-item.component.html

Trust-region variational autoencoders for multi-agent reinforcement learning

Loading...
Thumbnail Image

URL

Journal Title

Journal ISSN

Volume Title

School of Electrical Engineering | Master's thesis

Department

Mcode

Language

en

Pages

56

Series

Abstract

Learning effective policies from high-dimensional observations such as images remains a challenge in multi-agent reinforcement learning (MARL), particularly due to sample inefficiency and representational instability. While recent advances in single-agent settings have explored representation learning through generative and contrastive methods, their extension to multi-agent domains remains limited. This thesis introduces a novel framework that integrates a variational autoencoder (VAE) with a trust-region constraint, called multi-agent trust-region variational autoencoder (MA-TRVAE), into the MARL training pipeline. The proposed approach enables stable and efficient learning of latent representations from pixel-based observations, which are shared across agents for policy optimization. By constraining the encoder updates within a trust region, the method mitigates representational drift during training and improves overall policy stability. We evaluate the framework on vision-based cooperative tasks within a multi-agent quadcopter control (MAQC) environment and compare it with the baselines. Experimental results demonstrate that MA-TRVAE achieves improved performance, scalability, and robustness in representation learning, providing a promising direction for integrating trust-region-constrained VAEs in multi-agent systems.

Description

Supervisor

Baumann, Dominik

Thesis advisor

Deng, Mingwei

Other note

Citation

Endorsement

Review

Supplemented By

Referenced By