US2022253896A1PendingUtilityA1

Methods and systems for using reinforcement learning for promotions

Assignee: THE BOSTON CONSULTING GROUP INCPriority: Oct 11, 2018Filed: Apr 26, 2022Published: Aug 11, 2022
Est. expiryOct 11, 2038(~12.2 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 3/006G06Q 30/0244G06N 20/00
62
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and systems of using reinforcement learning for promotions. A first promotion is offered to a customer for a product and/or service. A first reward or penalty is determined, via a reinforcement machine learning model, based on the customer's reaction to the first promotion, wherein the reinforcement machine learning model is at a first state. Feedback to the reinforcement machine learning model is provided based on the first reward or penalty. A state of the reinforcement machine learning model is changed, based on the feedback, from the first state to a second state.

Claims

exact text as granted — not AI-modified
1 . A method of using reinforcement learning for optimizing promotions, comprising:
 offering a first promotion to a customer for a product and/or a service;   determining, via a reinforcement machine learning model, a first reward or penalty based on the customer's reaction to the first promotion, wherein the reinforcement machine learning model is at a first state;   providing feedback to the reinforcement machine learning model based on the first reward or penalty;   automatically changing, based on the feedback, a state of the reinforcement machine learning model from the first state to a second state;   generating a second promotion based on the reinforcement machine learning model at the second state;   offering the second promotion to the customer for the product and/or the service;   determining a second reward or penalty based on the customer's reaction to the second promotion;   comparing the first reward or penalty with the second reward or penalty;   determining a profit and cost of the promotion caused by moving from first state to second state the reinforcement machine learning model based on the comparing; and   presenting a promotion for the product and/or the service based on the determining of the profit and cost.   
     
     
         2 . The method of  claim 1 , wherein the promotion is a discount. 
     
     
         3 . The method of  claim 1 , wherein the promotion is a non-discount promotion. 
     
     
         4 . The method of  claim 1 , wherein the promotion is a coupon, advertisement, or recommendation, or any combination thereof. 
     
     
         5 . The method of  claim 1 , wherein a time dimension is utilized in the reinforcement learning so that the promotion and/or the timing can be optimized. 
     
     
         6 . The method of  claim 1 , wherein a Q-learning algorithm is used in the reinforcement learning model. 
     
     
         7 . The method of  claim 1 , wherein the promotion relates to a consumption pattern of regularly consumed products. 
     
     
         8 . The method of  claim 1 , wherein the product is a regularly consumed product. 
     
     
         9 . The method of  claim 1 , wherein a deep Q-learning algorithm is used in the reinforcement learning model. 
     
     
         10 . The method of  claim 1 , wherein a double Q-learning algorithm is used in the reinforcement learning model. 
     
     
         11 . A system using reinforcement learning to optimizing promotions, comprising a processor and associated memory, the processor configured for:
 offering a first promotion to a customer for a product and/or a service;   determining, via a reinforcement machine learning model, a first reward or penalty based on the customer's reaction to the first promotion, wherein the reinforcement machine learning model is at a first state;   providing feedback to the reinforcement machine learning model based on the first reward or penalty;   automatically changing, based on the feedback, a state of the reinforcement machine learning model from the first state to a second state;   generating a second promotion based on the reinforcement machine learning model at the second state;   offering the second promotion to the customer for the product and/or the service;   determining a second reward or penalty based on the customer's reaction to the second promotion;   comparing the first reward or penalty with the second reward or penalty;   determining a profit and cost of the promotion caused by moving from first state to second state the reinforcement machine learning model based on the comparing; and   presenting a promotion for the product and/or the service based on the determining of the profit and cost.   
     
     
         12 . The system of  claim 11 , wherein the promotion is a discount. 
     
     
         13 . The system of  claim 11 , wherein the promotion is a non-discount promotion. 
     
     
         14 . The system of  claim 11 , wherein the promotion is a coupon, advertisement, or recommendation, or any combination thereof. 
     
     
         15 . The system of  claim 11 , wherein a time dimension is utilized in the reinforcement learning so that the promotion and/or the timing can be optimized. 
     
     
         16 . The system of  claim 11 , wherein a Q-learning algorithm is used in the reinforcement learning model. 
     
     
         17 . The system of  claim 11 , wherein the promotion relates to a consumption pattern of regularly consumed products. 
     
     
         18 . The system of  claim 11 , wherein the product is a regularly consumed product. 
     
     
         19 . The system of  claim 11 , wherein a deep Q-learning algorithm is used in the reinforcement learning model. 
     
     
         20 . The system of  claim 11 , wherein a double Q-learning algorithm is used in the reinforcement learning model.

Join the waitlist — get patent alerts

Track US2022253896A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.