Selecting content items using reinforcement learning
Abstract
Methods, systems, and apparatus, including computer programs encoded on computer storage media, for using a machine learning model that has been trained through reinforcement learning to select a content item. One of the methods includes receiving first data characterizing a first context in which a first content item may be presented to a first user in a presentation environment; and providing the first data as input to a long-term engagement machine learning model, the model having been trained through reinforcement learning to: receive a plurality of inputs, and process each of the plurality of inputs to generate a respective engagement score for each input that represents a predicted, time-adjusted total number of selections by the respective user of future content items presented to the respective user in the presentation environment if the respective content item is presented in the respective context.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . (canceled)
2 . A system comprising one or more computers and one or more storage devices storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
receiving first data characterizing a first content item and a first context in which the first content item may be presented to a first user in a presentation environment; providing the first data as input to a long-term engagement machine learning model to obtain a first engagement score that represents a predicted, time-adjusted sum of rewards received at time windows that are subsequent to a current time window if the first content item is presented in the first context at the current time window; and determining, from at least the first engagement score, whether or not to present the first content item to the first user in the first context.
3 . The system of claim 2 , the operations further comprising:
generating, based at least on the first context, an estimate of a second engagement score that represents a predicted, time-adjusted sum of rewards received at time windows that are subsequent to a current time window if the first content item is not presented in the first context at the current time window; and determining, from at least the first engagement score and the second engagement score, whether or not to present the first content item to the first user in the first context.
4 . The system of claim 3 , wherein:
providing the first data as input to the long-term engagement machine learning model to obtain the first engagement score further comprises providing as input to the long-term engagement machine learning model data specifying a first action to present the first content item to the first user in the first context; and generating the estimate of the second engagement score further comprises providing as input to the long-term engagement machine learning model data specifying a second action to refrain from presenting the first content item to the first user in the first context.
5 . The system of claim 3 , wherein determining whether or not to present the first content item comprises determining to present the first content item to the first user in the first context only when the first engagement score is greater than the second engagement score.
6 . The system of claim 2 , the operations further comprising:
in response to determining to present the first content item, providing the first content item for presentation to the first user in the presentation environment or providing an indication to an external system that causes the external system to provide the first content item for presentation to the first user in the presentation environment.
7 . The system of claim 2 , wherein the long-term engagement machine learning model has been trained through reinforcement learning to determine trained values of parameters of the long-term engagement machine learning model.
8 . The system of claim 2 , wherein the data characterizing the first context comprises data characterizing content items previously presented to the first user in the presentation environment.
9 . The system of claim 2 , wherein the presentation environment is a response to a search query submitted by the user, and wherein the data characterizing the context comprises the search query.
10 . The system of claim 2 , wherein the first content item is a recommendation of content that may be of interest to the first user.
11 . The system of claim 9 , wherein the data characterizing the first context comprises data characterizing content items previously presented to the first user in response to one or more search queries previously submitted by the first user.
12 . The system of claim 2 , wherein the data characterizing the first context comprises data characterizing a quality of the first content item.
13 . The system of claim 2 , wherein the data characterizing the first context comprises a predicted likelihood that the first user will select the first content item if the first content item is presented to the first user in the first context.
14 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:
receiving first data characterizing a first content item and a first context in which the first content item may be presented to a first user in a presentation environment; providing the first data as input to a long-term engagement machine learning model to obtain a first engagement score that represents a predicted, time-adjusted sum of rewards received at time windows that are subsequent to a current time window if the first content item is presented in the first context at the current time window; and determining, from at least the first engagement score, whether or not to present the first content item to the first user in the first context.
15 . A method performed by one or more computers, the method comprising:
receiving first data characterizing a first content item and a first context in which the first content item may be presented to a first user in a presentation environment; providing the first data as input to a long-term engagement machine learning model to obtain a first engagement score that represents a predicted, time-adjusted sum of rewards received at time windows that are subsequent to a current time window if the first content item is presented in the first context at the current time window; and determining, from at least the first engagement score, whether or not to present the first content item to the first user in the first context.
16 . The method of claim 15 , further comprising:
generating, based at least on the first context, an estimate of a second engagement score that represents a predicted, time-adjusted sum of rewards received at time windows that are subsequent to a current time window if the first content item is not presented in the first context at the current time window; and determining, from at least the first engagement score and the second engagement score, whether or not to present the first content item to the first user in the first context.
17 . The method of claim 16 , wherein:
providing the first data as input to the long-term engagement machine learning model to obtain the first engagement score further comprises providing as input to the long-term engagement machine learning model data specifying a first action to present the first content item to the first user in the first context; and generating the estimate of the second engagement score further comprises providing as input to the long-term engagement machine learning model data specifying a second action to refrain from presenting the first content item to the first user in the first context.
18 . The method of claim 16 , wherein determining whether or not to present the first content item comprises determining to present the first content item to the first user in the first context only when the first engagement score is greater than the second engagement score.
19 . The method of claim 16 , wherein determining whether or not to present the first content item comprises determining to present the first content item to the first user in the first context only when the first engagement score is greater than a threshold value.
20 . The method of claim 15 , further comprising:
in response to determining to present the first content item, providing the first content item for presentation to the first user in the presentation environment or providing an indication to an external system that causes the external system to provide the first content item for presentation to the first user in the presentation environment.
21 . The method of claim 15 , wherein the data characterizing the first context comprises data characterizing content items previously presented to the first user in the presentation environment.
22 . The method of claim 15 , wherein the data characterizing the first context comprises data characterizing a quality of the first content item.
23 . The method of claim 15 , wherein the data characterizing the first context comprises a predicted likelihood that the first user will select the first content item if the first content item is presented to the first user in the first context.Join the waitlist — get patent alerts
Track US2021073638A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.