Nested Multi-Agent Reinforcement Learning for Adaptive Resource Management in 6G Network Slicing: A Multi-Timescale Framework with Convergence Guarantees
Faculty / School
Faculty of Computer Sciences (FCS)
Department
Department of Computer Science
Was this content written or created while at IBA?
Yes
Document Type
Article
Source Publication
IEEE Open Journal of the Communications Society
ISSN
2644-125X
Disciplines
Artificial Intelligence and Robotics | Data Science | Digital Communications and Networking
Abstract
Sixth-generation (6G) networks are expected to rely on agentic artificial intelligence for zerotouch, self-managed orchestration of heterogeneous network slices serving enhanced mobile broadband (eMBB), ultra-reliable low-latency communication URLLC), and massive machine-type communication (mMTC). A central and under-studied challenge for adaptive multi-agent resource management (AMRM) in such settings is multi-timescale non-stationarity: channel fading evolves per time-slot, user demand shifts at the window scale, and service-level agreement (SLA) regimes change at an operational scale. Single-timescale multi-agent reinforcement learning (MARL) algorithms cannot track all three signals cleanly—a learning rate fast enough for the per-slot channel destabilises the coordination structure that governs longer-timescale policies. This paper proposes Nested-MARL, an independent-learner actor-critic algorithm in which each agent’s parameters are partitioned into three groups updated at separated rates α0 α1 α2, with a continuum-memory exponential moving average (EMA) anchoring the slowest group. The design is grounded in the Nested Learning paradigm of Behrouz et al. (2025) and is extended here from single-model continual learning to decentralised multi-agent coordination. We establish a finite-time convergence result in the two-timescale stochastic approximation framework showing that under standard regularity and timescale-separation conditions, Nested-MARL achieves O(T −1/2 ) fast-group convergence vs. an Ω(T −1/3 ) lower bound for any single-timescale algorithm. An empirical study on a three-agent 6G slicing simulator with continuous multi-timescale drift shows Nested-MARL outperforms independent PPO (IPPO) in mean reward at every drift severity we test (κ ∈ {0.5, 1.0, 1.5, 2.0}) and by + 8.6% in sample efficiency over the first 40 episodes at κ=1.5 (n=10 seeds, p < 0.05). A controlled ablation establishes that stripping timescale separation reduces performance below the IPPO baseline, isolating timescale separation as the causal mechanism. Nested-MARL also reduces policy switching cost by 16.6%, an operationally meaningful benefit for zero-touch orchestration. The complete simulator, agents, and 60 + per-seed training runs are released as open source.
Indexing Information
HJRS - W Category, Web of Science - Emerging Sources Citation Index (ESCI)
Recommended Citation
Rashid, A., Iradat, S., Iqbal, W., Syed, I., Khan, K., & Yahya, K. (2026). Nested Multi-Agent Reinforcement Learning for Adaptive Resource Management in 6G Network Slicing: A Multi-Timescale Framework with Convergence Guarantees. IEEE Open Journal of the Communications Society, 7, 7624-7640. Retrieved from https://ir.iba.edu.pk/faculty-research-articles/282
Publication Status
Published
COinS
