US2025350538A1PendingUtilityA1

System and method for reconfigurable intelligent surface (ris)-assisted energy-efficient (ee) radio access network (ran) using hierarchical reinforcement learning

Assignee: ERICSSON TELEFON AB L MPriority: May 6, 2022Filed: May 5, 2023Published: Nov 13, 2025
Est. expiryMay 6, 2042(~15.8 yrs left)· nominal 20-yr term from priority
Y02D30/70H04L 41/0833H04B 7/04013G06N 3/092H04W 40/10H04W 88/12H04W 28/0284H04W 52/0206H04L 41/16G06N 20/00
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A method, system and apparatus for reconfigurable intelligent surface-assisted energy-efficient radio access networks using hierarchical reinforcement learning are disclosed. A method in a network node operating as a meta-controller and configured to communicate with a wireless device and a plurality of sub-controllers is provided. The method includes determining a state of the meta-controller including traffic load ratios of the plurality of sub-controllers. The method also includes receiving an extrinsic reward that is based on an energy efficiency of a cell including the plurality of sub-controllers. The method further includes selecting a goal according to a policy, the selected goal being an on/off state of each of the plurality of sub-controllers, the policy being selected to increase the extrinsic reward. The method includes configuring the plurality of controllers with the selected goal and an indication of the policy for selecting the goal.

Claims

exact text as granted — not AI-modified
1 . A method implemented in a network node operating as a meta-controller and configured to communicate with a wireless device, WD, and a plurality of sub-controllers, the method comprising:
 determining a state of the meta-controller, the state of the meta-controller including traffic load ratios of the plurality of sub-controllers;   receiving an extrinsic reward, the extrinsic reward being based at least in part on an energy efficiency of a cell including the plurality of sub-controllers;   selecting a goal according to a policy, the selected goal being an on/off state of each of the plurality of sub-controllers, the policy being selected to increase the extrinsic reward; and   configuring the plurality of sub-controllers with the selected goal and an indication of the policy for selecting the goal.   
     
     
         2 . The method of  claim 1 , wherein the extrinsic reward is based at least in part a ratio of a sum of throughputs of network nodes in the cell to a sum of power consumptions of the network nodes in the cell. 
     
     
         3 . The method of  claim 2 , wherein the extrinsic reward is further based at least in part on a penalty factor to avoid overloading and based at least in part on a number of network nodes in the cell that are overloaded. 
     
     
         4 . The method of  claim 1 , wherein maximizing the extrinsic reward includes maximizing a throughput of a link between the network node and a plurality of WDs via a reconfigurable intelligent surface, RIS. 
     
     
         5 . The method of  claim 1 , wherein the policy is one of a greedy policy and an e-greedy policy. 
     
     
         6 . The method of  claim 5 , wherein the greedy policy provides a selected goal that increases the extrinsic reward when a random number exceeds a threshold and provides a randomly selected goal when the random number does not exceed the threshold. 
     
     
         7 . A network node operating as a meta-controller and configured to communicate with a wireless device, WD, and a plurality of sub-controllers, the network node comprising processing circuitry configured to:
 determine a state of the meta-controller, the state of the meta-controller including traffic load ratios of the plurality of sub-controllers;   receive an extrinsic reward, the extrinsic reward being based at least in part on an energy efficiency of a cell including the plurality of sub-controllers;   select a goal according to a policy, the selected goal being an on/off state of each of the plurality of sub-controllers, the policy being selected to increase the extrinsic reward; and   configure the plurality of sub-controllers with the selected goal and an indication of the policy for selecting the goal.   
     
     
         8 . The network node of  claim 7 , wherein the extrinsic reward is based at least in part a ratio of a sum of throughputs of network nodes in the cell to a sum of power consumptions of the network nodes in the cell. 
     
     
         9 . The network node of  claim 8 , wherein the extrinsic reward is further based at least in part on a penalty factor to avoid overloading and based at least in part on a number of network nodes in the cell that are overloaded. 
     
     
         10 . The network node of  claim 7 , wherein maximizing the extrinsic reward includes maximizing a throughput of a link between the network node and a plurality of WDs via a reconfigurable intelligent surface, RIS. 
     
     
         11 . The network node of  claim 7 , wherein the policy is one of a greedy policy and an e-greedy policy. 
     
     
         12 . The network node of  claim 11 , wherein the greedy policy provides a selected goal that increases the extrinsic reward when a random number exceeds a threshold and provides a randomly selected goal when the random number does not exceed the threshold. 
     
     
         13 . A method implemented in a network node operating as a sub-controller and configured to communicate with a wireless device (WD) and at least one network node operating as a meta-controller, the method comprising:
 receiving a goal and an indication of a policy from the meta-controller;   receiving an intrinsic reward, the intrinsic reward being based at least in part on a ratio of a throughput of the sub-controller to a transmission power of the sub-controller; and   selecting an action based at least in part on the goal and according to the policy, the action including adjusting the transmission power of the sub-controller to increase the intrinsic reward.   
     
     
         14 . The method of  claim 13 , wherein the intrinsic reward is further based at least in part on a penalty factor to avoid overloading and based at least in part on a number of network nodes in a cell that are overloaded. 
     
     
         15 . The method of  claim 14 , wherein maximizing the intrinsic reward includes maximizing a throughput of a link between the network node and a plurality of WDs via a reconfigurable intelligent surface, RIS. 
     
     
         16 . The method of  claim 13 , wherein the policy is one of a greedy policy and an e-greedy policy. 
     
     
         17 . The method of  claim 16 , wherein the greedy policy provides a selected action that increases the intrinsic reward when a random number exceeds a threshold and provides a randomly selected action when the random number does not exceed the threshold. 
     
     
         18 . The method of  claims 13-17   claim 13 , wherein the goal is an on/off state of the sub-controller. 
     
     
         19 .- 24 . (canceled) 
     
     
         25 . The method of  claim 2 , wherein maximizing the extrinsic reward includes maximizing a throughput of a link between the network node and a plurality of WDs via a reconfigurable intelligent surface, RIS. 
     
     
         26 . The method of  claim 2 , wherein the policy is one of a greedy policy and an ε-greedy policy.

Join the waitlist — get patent alerts

Track US2025350538A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.