The emergence of large-scale Low Earth Orbit (LEO) satellite constellations has renewed attention to satellite networks, with a vision to deliver global connectivity and performance levels comparable to terrestrial infrastructures. Despite this progress, such constellations are often treated as standalone systems rather than being fully integrated into next-generation communication architectures. To bridge this gap, extensive research has been devoted to the development of Non-Terrestrial Networks (NTN), aiming to enable unified operation across terrestrial and space segments within cellular infrastructures. Nonetheless, realizing this vision remains challenging due to persistent issues in latency, dynamic scheduling, and resource management. To tackle these challenges, we present the development of a Space Cloud, where Multi-access Edge Computing (MEC) services are deployed across satellite nodes and interconnected through Inter-Satellite Links (ISL), forming a distributed space-based data center. By enabling in-orbit computation, the Space Cloud reduces dependence on terrestrial infrastructure, as processing no longer needs to be offloaded to ground-based data centers. This architectural shift is key to meeting the lowlatency requirements of modern, latency-sensitive applications. This paper proposes a distributed Reinforcement Learning (RL) approach to manage computational resources in space-based MEC environments. The system leverages a scalable actorcritic framework, where neural network-based actors and critics are deployed on individual satellites. Each node operates autonomously and in a distributed manner, enabling the network to optimize actions locally while considering both instant rewards and long-term impact. The RL model leverages historical data and processing patterns to control the activation of on-orbit servers, aiming for efficient resource utilization. We evaluate the proposed strategy using a synthetic constellation designed with our in-house satellite emulation and MEC computation framework. Moreover, through Pareto-efficient analysis across key performance indicators (KPI), we benchmark our approach against conventional baselines and assess its applicability under the constraints of specific space missions. The results demonstrate that the RL controller achieves comparable task failure and latency rates to the baselines, while significantly reducing resource consumption.

Distributed Reinforcement Learning for Resource Management in Satellite Edge Computing Systems

Badini N.;Patrone F.;Marchese M.
2026-01-01

Abstract

The emergence of large-scale Low Earth Orbit (LEO) satellite constellations has renewed attention to satellite networks, with a vision to deliver global connectivity and performance levels comparable to terrestrial infrastructures. Despite this progress, such constellations are often treated as standalone systems rather than being fully integrated into next-generation communication architectures. To bridge this gap, extensive research has been devoted to the development of Non-Terrestrial Networks (NTN), aiming to enable unified operation across terrestrial and space segments within cellular infrastructures. Nonetheless, realizing this vision remains challenging due to persistent issues in latency, dynamic scheduling, and resource management. To tackle these challenges, we present the development of a Space Cloud, where Multi-access Edge Computing (MEC) services are deployed across satellite nodes and interconnected through Inter-Satellite Links (ISL), forming a distributed space-based data center. By enabling in-orbit computation, the Space Cloud reduces dependence on terrestrial infrastructure, as processing no longer needs to be offloaded to ground-based data centers. This architectural shift is key to meeting the lowlatency requirements of modern, latency-sensitive applications. This paper proposes a distributed Reinforcement Learning (RL) approach to manage computational resources in space-based MEC environments. The system leverages a scalable actorcritic framework, where neural network-based actors and critics are deployed on individual satellites. Each node operates autonomously and in a distributed manner, enabling the network to optimize actions locally while considering both instant rewards and long-term impact. The RL model leverages historical data and processing patterns to control the activation of on-orbit servers, aiming for efficient resource utilization. We evaluate the proposed strategy using a synthetic constellation designed with our in-house satellite emulation and MEC computation framework. Moreover, through Pareto-efficient analysis across key performance indicators (KPI), we benchmark our approach against conventional baselines and assess its applicability under the constraints of specific space missions. The results demonstrate that the RL controller achieves comparable task failure and latency rates to the baselines, while significantly reducing resource consumption.
2026
979-8-3315-7360-7
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11567/1313956
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact