US2008312911A1PendingUtilityA1
Dictionary word and phrase determination
Est. expiryJun 14, 2027(~0.9 yrs left)· nominal 20-yr term from priority
Inventors:Po Zhang
G06F 40/53G06F 40/242G06F 16/3338
44
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Context signals in documents are identified, characters bounded by the context signals are identified, one or more candidate words defined by the characters bounded by the context signals are identified, and one or more of the candidate words are added to an input method editor dictionary.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method, comprising:
identifying context signals in documents; identifying characters bounded by the context signals; identifying one or more candidate words defined by the characters bounded by the context signals; and adding one or more of the candidate words to an input method editor dictionary.
2 . The method of claim 1 wherein identifying context signals in documents comprises identifying Chinese book title marks.
3 . The method of claim 2 wherein the Chinese book title marks comprise single book title marks or double book title marks.
4 . The method of claim 1 wherein identifying characters bounded by the context signals comprises identifying Hanzi characters bounded by the context signals.
5 . The method of claim 1 wherein the candidate words comprise Chinese words.
6 . The method of claim 1 , wherein identifying context signals in documents comprises identifying hypertext markup language tags in electronic documents.
7 . The method of claim 1 wherein the input method editor dictionary comprises a Chinese input method editor dictionary.
8 . The method of claim 1 , comprising determining a count of each candidate word.
9 . The method of claim 8 wherein adding one or more of the candidate words to the input method editor dictionary comprises adding candidate words having a count that exceeds a threshold to the input method editor dictionary.
10 . The method of claim 8 , wherein identifying context signals in documents comprises identifying non-duplicative documents.
11 . The method of claim 10 , wherein determining a count of each candidate word comprises determining the count of each candidate word based on only the non-duplicative documents.
12 . The method of claim 1 wherein the documents comprise web documents obtained from the Internet.
13 . The method of claim 1 , comprising identifying candidate words in search queries and adding one or more of the candidate words to the input method editor dictionary.
14 . The method of claim 13 wherein identifying candidate words in search queries comprises:
for each candidate word,
determining a first count representing a number of times that the candidate word is the only word in the search queries, and
determining a second count representing a number of times that the candidate word and one or more other words are included in each of the search queries, and
adding one or more of the candidate words to the input method editor dictionary based on a relationship between the first count and the second count.
15 . A method, comprising:
establishing a dictionary that includes words that are identified based on characters bounded by context signals; and providing an input method editor configured to select words from the dictionary.
16 . The method of claim 15 wherein establishing the dictionary comprises identifying words based on characters bounded by Chinese book title marks.
17 . The method of claim 15 , comprising identifying candidate words in search queries and adding one or more of the candidate words to the dictionary.
18 . An apparatus, comprising:
a dictionary that includes words that are identified based on candidate words that are associated with characters found in documents, in which each candidate word is associated with one or more characters bounded by the context signals; and an input method editor configured to select words from the dictionary
19 . The apparatus of claim 18 wherein the candidate words comprise Hanzi characters.
20 . The apparatus of claim 18 wherein the context signals comprise Chinese book title marks.
21 . The apparatus of claim 20 wherein the Chinese book title marks comprise at least one of single book title marks or double book title marks.
22 . The apparatus of claim 18 wherein the dictionary comprises words identified based on a first count representing a number of times that the word is the only word in search queries and a second count representing a number of times that the word and one or more other words are in each of the search queries.
23 . The apparatus of claim 18 wherein the input method editor dictionary comprises a Chinese input method editor dictionary.
24 . A system, comprising:
a data store to store a document corpus; and a processing engine stored in computer readable medium and comprising instructions executable by a processing device that upon such execution cause the processing device to:
identify candidate words by finding characters in documents of the document corpus in which the characters are enclosed in pairs of Chinese book title marks, and
add one or more of the candidate words to an input method editor dictionary.
25 . A system, comprising:
means for identifying context signals in documents; means for identifying characters bounded by the context signals; means for identifying one or more candidate words defined by the characters bounded by the context signals; and means for adding one or more of the candidate words to an input method editor dictionary.Join the waitlist — get patent alerts
Track US2008312911A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.