System and method for determining semantically related term
Abstract
Methods and systems for determining semantically related terms are disclosed. Generally, seed terms are received from a user. One or more potential terms semantically related to the seed terms are determined based on vectors comprising entries regarding a plurality of terms, a plurality of universal resource locators (“URLs”) associated with each term of the plurality of terms, and for each URL in a search log, a number of times that one or more users searched for each of the terms in the search log and clicked on the URL. At least a portion of the potential terms is then suggested to the user.
Claims
exact text as granted — not AI-modified1 . A method for determining semantically related terms, comprising:
receiving one or more seed terms; determining one or more potential terms semantically related to the seed terms based on one or more vectors comprising entries regarding a plurality of terms, a plurality of universal resource locators (“URLs”) associated with each term of the plurality of terms, and for each URL in a search log, a number of times that one or more users searched for each of the terms of the plurality of terms in the search log and clicked on the URL; and suggesting at least a portion of the determined potential terms.
2 . The method of claim 1 , wherein determining one or more potential terms semantically related to the seed terms based on one or more vectors comprises:
creating a set of URL vectors representing for each URL in the search log, a number of times that one or more users searched for each of the terms in the search log and clicked on the URL; creating a set of query vectors representing for each URL in the search log, whether a user searched for each of the terms in the search log and clicked on the URL; and determining one or more potential terms semantically related to the seed terms based on the set of URL vectors, set of query vectors, and at least one seed term.
3 . The method of claim 2 , wherein determining one or more potential terms semantically related to the seed terms based on one or more vectors further comprises:
weighing each entry in the set of URL vectors and the set of query vectors based on a number of times a URL and a term appears in the search log; and normalizing the set of URL vectors and the set of query vectors.
4 . The method of claim 2 , wherein determining one or more potential terms semantically related to the seed terms based on one or more vectors further comprises:
determining a weighted average over each URL in the search logs as a function of the set of URL vectors, the set of query vectors, and the at least one seed term.
5 . The method of claim 4 , wherein the weighted average is calculated using the formula:
T =Sum of ( V 1*cosine( V 2, S)),
wherein V 1 *cosine(V 2 ,S) is calculated for each URL in the search log; V 1 is the relevant query vector indicating for each term in the search log, whether a searcher searched for the term and clicked on the relevant URL; V 2 is the relevant URL vector indicating for each term in the search log, a number of times a user searched for the term and clicked on the relevant URL; S is a seed term vector indicating for each term in the search log, whether the term is one of the seed terms; and T is a vector indicating for each term in the search log, how relevant the term is to the seed terms.
6 . The method of claim 1 , wherein receiving one or more search terms comprises:
receiving a location of a webpage; and pulling one or more terms from the content of the webpage.
7 . The method of claim 6 , further comprising:
suggesting at least one term of the terms pulled from the content of the webpage.
8 . The method of claim 6 , wherein pulling one or more terms from the content of the webpage comprises:
retrieving the content of the webpage from the location of the webpage; stripping code from the content of the webpage; pulling one or more terms from the content of the webpage; and weighing each term of the terms pulled from the content of the webpage based on a location of where the term was located on the webpage.
9 . The method of claim 1 , wherein the seed terms comprise at least one term sent to an internet search engine.
10 . A computer-readable storage medium comprising a set of instructions for determining semantically related terms, the set of instructions to direct a computer system to perform the acts of:
receiving one or more seed terms; determining one or more potential terms semantically related to the seed terms based on one or more vectors comprising entries regarding a plurality of terms, a plurality of universal resource locators (“URLs”) associated with each term of the plurality of terms, and for each URL in a search log, a number of times that one or more users searched for each of the terms of the plurality of terms in the search log and clicked on the URL; and suggesting at least a portion of the determined potential terms.
11 . The computer-readable storage medium of claim 10 , wherein determining one or more potential terms semantically related to the seed terms based on one or more vectors comprises:
creating a set of URL vectors representing for each URL in the search log, a number of times that one or more users searched for each of the terms in the search log and clicked on the URL; creating a set of query vectors representing for each URL in the search log, whether a user searched for each of the terms in the search log and clicked on the URL; and determining one or more potential terms semantically related-to the seed terms based on the set of URL vectors, set of query vectors, and at least one seed term.
12 . The computer-readable storage medium of claim 11 , wherein determining one or more potential terms semantically related to the seed terms based on one or more vectors further comprises:
weighing each entry in the set of URL vectors and the set of query vectors based on a number of times a URL and a term appears in the search log; and normalizing the set of URL vectors and the set of query vectors.
13 . The computer-readable storage medium of claim 12 , wherein the weighted average is calculated using the formula:
T =Sum of ( V 1*cosine( V 2, S)),
wherein V 1 *cosine(V 2 ,S) is calculated for each URL in the search log; V 1 is the relevant query vector indicating for each term in the search log, whether a searcher searched for the term and clicked on the relevant URL; V 2 is the relevant URL vector indicating for each term in the search log, a number of times a user searched for the term and clicked on the relevant URL; S is a seed term vector indicating for each term in the search log, whether the term is one of the seed terms; and T is a vector indicating for each term in the search log, how relevant the term is to the seed terms.
14 . The computer-readable storage medium of claim 10 , wherein receiving one or more search terms comprises:
receiving a location of a webpage; retrieving the content of the webpage from the location of the webpage; stripping code from the content of the webpage; pulling one or more terms from the content of the webpage; and weighing each term of the terms pulled from the content of the webpage based on a location of where the term was located on the webpage.
15 . The computer-readable storage medium of claim 10 , wherein the seed terms comprises at least one term sent to an internet search engine.
16 . A system for determining semantically related words comprising:
at least one database comprising:
a set of universal resource locators (“URL”) vectors representing for each URL in a search log, a number of times that one or more users searched for each of the terms in the search log and clicked on the URL; and
a set of query vectors representing for each URL in the search log, whether a user searched for each of the terms in the search log and clicked on the URL; and
at least one server operative to access the set of URL vectors and set of query vectors of the at least one database, the at least one server configured to:
receive one or more seed terms;
determine a plurality of potential terms semantically related to the seed terms based on the set of URL vectors, set of query vectors, and the seed terms; and
suggest at least a portion of the plurality of potential terms.
17 . The system of claim 16 , wherein to receive one or more seed terms, the at least one server is further configured to:
receive a location of a webpage; retrieve the content of the webpage from the location of the webpage; strip code from the content of the webpage; pull one or more terms from the content of the webpage; and weigh each term of the terms pulled from the content of the webpage based on a location of where the term was located on the webpage.
18 . The system of claim 16 , wherein the seed terms comprise at least one term sent to an internet search engine.
19 . A method for determining semantically related terms, comprising:
receiving one or more seed terms; searching an index to determine a plurality of potential terms semantically related to the seed terms, the index comprising associations between terms regarding which terms in a search log resulted in one or more searchers clicking on the same universal resource locator (“URL”); and suggesting at least a first term of the plurality of potential terms.
20 . The method of claim 19 , further comprising:
receiving an indication of relevance of the first term to a user; and modifying the terms which comprise the seed terms based at least in part on the received indication of relevance.
21 . The method of claim 20 , wherein the indication of relevance to the user is an indication of relevance to the user on a scale.
22 . The method of claim 20 , wherein modifying the terms which comprise the seed terms comprises:
receiving an indication that a first term is relevant to the user; and modifying the seed terms to comprise the first term as a positive seed term.
23 . The method of claim 22 , wherein modifying the terms which comprise the seed terms comprises:
receiving an indication that a second term is not relevant to the user; and modifying the seed terms to comprise the second term as a negative seed term.
24 . The method of claim 20 , further comprising:
searching the index to determine a second plurality of potential terms semantically related to the modified seed terms; and suggesting at least one term of the second plurality of potential terms.
25 . The method of claim 20 , further comprising:
examining, with a function learning algorithm, the associations in the index between terms regarding at least which terms in the search log resulted in the same or different searcher clicking on the same URL; and predicting a degree of relevance between two terms based on the algorithm learned by the function learning algorithm in examining the index.
26 . The method of claim 20 , wherein receiving one or more seed terms comprises:
receiving a location of a webpage; retrieving the content of the webpage from the location of the webpage; stripping code from the content of the webpage; pulling one or more terms from the content of the webpage; and weighing each term of the terms pulled from the content of the webpage based on a location of where the term was located on the webpage.
27 . The method of claim 20 , wherein the seed terms comprise at least one term sent to an internet search engine.
28 . A computer-readable storage medium storing a set of instructions for determine semantically related terms, the set of instructions to direct a computer system to perform acts of:
receiving one or more seed terms; searching an index to determine a plurality of potential terms semantically related to the seed terms, the index comprising associations between terms regarding which terms in a search log resulted in one or more searchers clicking on the same universal resource locator (“URL”); and suggesting at least a first term of the plurality of potential terms.
29 . The computer-readable storage medium of claim 28 , further comprising a set of instructions to direct the computer system to perform acts of:
receiving an indication of relevance of the first term to a user; and modifying the terms which comprise the seed terms based at least in part on the received indication of relevance.
30 . The computer-readable storage medium of claim 29 , wherein the indication of relevance to the seed terms is an indication of relevance to the seed terms on a scale.
31 . The computer-readable storage medium of claim 28 , further comprising a set of instructions to direct the computer system to perform acts of:
examining, with a function learning algorithm, the associations in the index between terms regarding at least which terms in the search log resulted in the same or different searcher clicking on the same URL; and predicting a degree of relevance between two terms based on the algorithm learned by the function learning algorithm in examining the index.
32 . The computer-readable storage medium of claim 28 , wherein receiving at least one search term comprises:
receiving a location of a webpage; retrieving the content of the webpage from the location of the webpage; stripping code from the content of the webpage; pulling one or more terms from the content of the webpage; and weighing each term of the terms pulled from the content of the webpage based on a location of where the term was located on the webpage.
33 . The computer-readable storage medium of claim 28 , wherein the seed terms comprise at least one search term sent to an internet search engine.
34 . A system for determining semantically related terms comprising:
a database comprising an index comprising associations between terms regarding which terms in a search log resulted in one or more searchers clicking on the same universal resource locator (“URL”); at least one server operative to access the index, the at least one server configured to:
receive one or more seed terms;
search the index to determine a plurality of potential terms semantically related to the seed terms; and
suggest at least a first term of the plurality of potential terms.
35 . The system of claim 34 , wherein the at least one server is further configured to:
receiving an indication of relevance of the first term to a user; and modifying the terms which comprise the seed terms based at least in part on the received indication of relevance.
36 . The system of claim 34 , wherein the indication of relevance to the user is an indication of relevance to the seed terms on a scale.
37 . The system of claim 34 , wherein the at least one server is further configured to:
examine, with a function learning algorithm, the associations in the index between terms regarding which terms in the search log resulted in one or more searchers clicking on the same URL; and predict a degree of relevance between two terms based on the algorithm learned by the function learning algorithm in examining the index.
38 . The system of claim 34 , wherein to receive the one or more seed terms, the at least one server is further configured to:
receive a location of a webpage; retrieve the content of the webpage from the location of the webpage; strip code from the content of the webpage; pull one or more terms from the content of the webpage; and weigh each term of the terms pulled from the content of the webpage based on a location of where the term was located on the webpage.
39 . The system of claim 34 , wherein the seed terms comprise at least one term sent to an internet search engine.
40 . The system of claim 34 , wherein the first term of the plurality of potential terms is suggested via a user interface of an advertisement campaign management system.
41 . The system of claim 34 , wherein the first term of the plurality of potential terms is suggested via an application program interface of an advertisement campaign management system.Join the waitlist — get patent alerts
Track US2007027865A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.