US2022253896A1PendingUtilityA1
Methods and systems for using reinforcement learning for promotions
Assignee: THE BOSTON CONSULTING GROUP INCPriority: Oct 11, 2018Filed: Apr 26, 2022Published: Aug 11, 2022
Est. expiryOct 11, 2038(~12.2 yrs left)· nominal 20-yr term from priority
G06N 7/01G06N 3/006G06Q 30/0244G06N 20/00
62
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Methods and systems of using reinforcement learning for promotions. A first promotion is offered to a customer for a product and/or service. A first reward or penalty is determined, via a reinforcement machine learning model, based on the customer's reaction to the first promotion, wherein the reinforcement machine learning model is at a first state. Feedback to the reinforcement machine learning model is provided based on the first reward or penalty. A state of the reinforcement machine learning model is changed, based on the feedback, from the first state to a second state.
Claims
exact text as granted — not AI-modified1 . A method of using reinforcement learning for optimizing promotions, comprising:
offering a first promotion to a customer for a product and/or a service; determining, via a reinforcement machine learning model, a first reward or penalty based on the customer's reaction to the first promotion, wherein the reinforcement machine learning model is at a first state; providing feedback to the reinforcement machine learning model based on the first reward or penalty; automatically changing, based on the feedback, a state of the reinforcement machine learning model from the first state to a second state; generating a second promotion based on the reinforcement machine learning model at the second state; offering the second promotion to the customer for the product and/or the service; determining a second reward or penalty based on the customer's reaction to the second promotion; comparing the first reward or penalty with the second reward or penalty; determining a profit and cost of the promotion caused by moving from first state to second state the reinforcement machine learning model based on the comparing; and presenting a promotion for the product and/or the service based on the determining of the profit and cost.
2 . The method of claim 1 , wherein the promotion is a discount.
3 . The method of claim 1 , wherein the promotion is a non-discount promotion.
4 . The method of claim 1 , wherein the promotion is a coupon, advertisement, or recommendation, or any combination thereof.
5 . The method of claim 1 , wherein a time dimension is utilized in the reinforcement learning so that the promotion and/or the timing can be optimized.
6 . The method of claim 1 , wherein a Q-learning algorithm is used in the reinforcement learning model.
7 . The method of claim 1 , wherein the promotion relates to a consumption pattern of regularly consumed products.
8 . The method of claim 1 , wherein the product is a regularly consumed product.
9 . The method of claim 1 , wherein a deep Q-learning algorithm is used in the reinforcement learning model.
10 . The method of claim 1 , wherein a double Q-learning algorithm is used in the reinforcement learning model.
11 . A system using reinforcement learning to optimizing promotions, comprising a processor and associated memory, the processor configured for:
offering a first promotion to a customer for a product and/or a service; determining, via a reinforcement machine learning model, a first reward or penalty based on the customer's reaction to the first promotion, wherein the reinforcement machine learning model is at a first state; providing feedback to the reinforcement machine learning model based on the first reward or penalty; automatically changing, based on the feedback, a state of the reinforcement machine learning model from the first state to a second state; generating a second promotion based on the reinforcement machine learning model at the second state; offering the second promotion to the customer for the product and/or the service; determining a second reward or penalty based on the customer's reaction to the second promotion; comparing the first reward or penalty with the second reward or penalty; determining a profit and cost of the promotion caused by moving from first state to second state the reinforcement machine learning model based on the comparing; and presenting a promotion for the product and/or the service based on the determining of the profit and cost.
12 . The system of claim 11 , wherein the promotion is a discount.
13 . The system of claim 11 , wherein the promotion is a non-discount promotion.
14 . The system of claim 11 , wherein the promotion is a coupon, advertisement, or recommendation, or any combination thereof.
15 . The system of claim 11 , wherein a time dimension is utilized in the reinforcement learning so that the promotion and/or the timing can be optimized.
16 . The system of claim 11 , wherein a Q-learning algorithm is used in the reinforcement learning model.
17 . The system of claim 11 , wherein the promotion relates to a consumption pattern of regularly consumed products.
18 . The system of claim 11 , wherein the product is a regularly consumed product.
19 . The system of claim 11 , wherein a deep Q-learning algorithm is used in the reinforcement learning model.
20 . The system of claim 11 , wherein a double Q-learning algorithm is used in the reinforcement learning model.Join the waitlist — get patent alerts
Track US2022253896A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.