Multi-agent trajectory planning in safety-critical systems needs to ensure safety while scaling to many agents. Sampling and optimization methods often adapt slowly and scale poorly. Reinforcement learning can improve adaptability, but it often violates safety constraints and suffers sample inefficiency. This work proposes IDDPG-MAF, which integrates Independent Deep Deterministic Policy Gradient (IDDPG) with a pre-trained Multi-head Action Filter Network (MAF-Net). We first cast the problem as a constrained mixed-integer nonlinear program and then reformulate it as a constrained decentralized Markov decision process for real-time adaptability and coordination. IDDPG enables scalable learning, while MAF-Net acts as a differentiable safety filter that masks unsafe actions and penalizes suboptimal behaviors. The IDDPG-MAF method is adapted to a complex multi-aircraft trajectory planning task under dynamic thunderstorm cells. Experimental results show that IDDPG-MAF achieves over 99% safe separation (vs. 82% for the state-of-the-art baseline), 95.5% task success even under moderate uncertainty, and scales safely to 45 aircraft in a compact spatiotemporal window, effectively doubling the maximum capacity of current operations.

Pang, B., Zhang, M., Hu, X., Pham, D., Alam, S., Lulli, G. (2026). Constrained Multi-Agent Reinforcement Learning with MAF-Net for Safe Trajectory Planning. In AAMAS '26: Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems (pp.1928-1937). Association for Computing Machinery, Inc [10.65109/MQDV9851].

Constrained Multi-Agent Reinforcement Learning with MAF-Net for Safe Trajectory Planning

Lulli, G
2026

Abstract

Multi-agent trajectory planning in safety-critical systems needs to ensure safety while scaling to many agents. Sampling and optimization methods often adapt slowly and scale poorly. Reinforcement learning can improve adaptability, but it often violates safety constraints and suffers sample inefficiency. This work proposes IDDPG-MAF, which integrates Independent Deep Deterministic Policy Gradient (IDDPG) with a pre-trained Multi-head Action Filter Network (MAF-Net). We first cast the problem as a constrained mixed-integer nonlinear program and then reformulate it as a constrained decentralized Markov decision process for real-time adaptability and coordination. IDDPG enables scalable learning, while MAF-Net acts as a differentiable safety filter that masks unsafe actions and penalizes suboptimal behaviors. The IDDPG-MAF method is adapted to a complex multi-aircraft trajectory planning task under dynamic thunderstorm cells. Experimental results show that IDDPG-MAF achieves over 99% safe separation (vs. 82% for the state-of-the-art baseline), 95.5% task success even under moderate uncertainty, and scales safely to 45 aircraft in a compact spatiotemporal window, effectively doubling the maximum capacity of current operations.
paper
action masking; decentralized decision making; deep reinforcement learning; Multi-agent system; planning under uncertainty;
English
25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026) - 25-29 May 2026
2026
AAMAS '26: Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems
9798400723179
2026
1928
1937
open
Pang, B., Zhang, M., Hu, X., Pham, D., Alam, S., Lulli, G. (2026). Constrained Multi-Agent Reinforcement Learning with MAF-Net for Safe Trajectory Planning. In AAMAS '26: Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems (pp.1928-1937). Association for Computing Machinery, Inc [10.65109/MQDV9851].
File in questo prodotto:
File Dimensione Formato  
Pang et al-2026-AAMAS-VoR.pdf

accesso aperto

Tipologia di allegato: Publisher’s Version (Version of Record, VoR)
Licenza: Creative Commons
Dimensione 1.53 MB
Formato Adobe PDF
1.53 MB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10281/611324
Citazioni
  • Scopus 0
  • ???jsp.display-item.citation.isi??? ND
Social impact