Nowadays, microservice-based applications are often deployed across a computing continuum that spans from edge devices (offering low-latency access but limited resources) to powerful cloud data centers. Balancing Quality of Service (QoS) constraints with cost efficiency in such a dynamic, heterogeneous environment remains a major challenge, especially when workloads fluctuate unpredictably. We introduce FIGARO (reinForcement learnInG mAnagement acRoss the computing cOntinuum), a hierarchical Reinforcement Learning (RL) framework for the dynamic resource management in the computing continuum. By observing component demands, resource characteristics, and QoS constraints, FIGARO guides layer-specific RL agents to autonomously adjust the number of active resources and handle unpredictable variations in the system state. Our approach includes an initial offline training phase, where agents gain insights from a heuristic-based expert, followed by an online deployment that adapts to evolving environment conditions.The experimental validation shows that FIGARO maintains QoS violations below 7% even when the distributions of components service times deviates significantly from the one considered during training, while reducing energy costs related to the resource usage by around 5% compared to baseline agents. The framework can therefore address large-scale, distributed computing environments and complex workflows, providing an effective solution to runtime microservice management in the computing continuum.
Cavadini, R., Da Mommio, M., Filippini, F., Sedghani, H., Lancellotti, R., Ardagna, D. (2026). FIGARO: Hierarchical reinforcement learning for scalable microservice management in the computing continuum. JOURNAL OF PARALLEL AND DISTRIBUTED COMPUTING, 218 [10.1016/j.jpdc.2026.105336].
FIGARO: Hierarchical reinforcement learning for scalable microservice management in the computing continuum
Filippini F.;
2026
Abstract
Nowadays, microservice-based applications are often deployed across a computing continuum that spans from edge devices (offering low-latency access but limited resources) to powerful cloud data centers. Balancing Quality of Service (QoS) constraints with cost efficiency in such a dynamic, heterogeneous environment remains a major challenge, especially when workloads fluctuate unpredictably. We introduce FIGARO (reinForcement learnInG mAnagement acRoss the computing cOntinuum), a hierarchical Reinforcement Learning (RL) framework for the dynamic resource management in the computing continuum. By observing component demands, resource characteristics, and QoS constraints, FIGARO guides layer-specific RL agents to autonomously adjust the number of active resources and handle unpredictable variations in the system state. Our approach includes an initial offline training phase, where agents gain insights from a heuristic-based expert, followed by an online deployment that adapts to evolving environment conditions.The experimental validation shows that FIGARO maintains QoS violations below 7% even when the distributions of components service times deviates significantly from the one considered during training, while reducing energy costs related to the resource usage by around 5% compared to baseline agents. The framework can therefore address large-scale, distributed computing environments and complex workflows, providing an effective solution to runtime microservice management in the computing continuum.| File | Dimensione | Formato | |
|---|---|---|---|
|
Cavadini et al-2026-Journal of Parallel and Distributed Computing-VoR.pdf
Solo gestori archivio
Tipologia di allegato:
Publisher’s Version (Version of Record, VoR)
Licenza:
Tutti i diritti riservati
Dimensione
1.9 MB
Formato
Adobe PDF
|
1.9 MB | Adobe PDF | Visualizza/Apri Richiedi una copia |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.


