US2008270115A1PendingUtilityA1
System and method for diacritization of text
Individually held — no corporate assignee on recordPriority: Mar 22, 2006Filed: Jun 3, 2008Published: Oct 30, 2008
Est. expiryMar 22, 2026(expired)· nominal 20-yr term from priority
G06F 40/53G06F 40/232
53
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A system and method for restoration of diacritics includes making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration. A best diacritic representation is determined for graphemes in the utterance based upon a best match with the diacritization model. A diacritically restored representation of the utterance is output.
Claims
exact text as granted — not AI-modified1 . A method for restoration of diacritics, comprising:
making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration; determining a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model; and outputting a diacritically restored representation of the utterance.
2 . The method as recited in claim 1 , wherein the diacritization model includes information for a Semitic language.
3 . The method as recited in claim 1 , further comprising integrating a plurality of sources of information to train the diacritization model.
4 . The method as recited in claim 1 , wherein the diacritization model is trained using a Maximum Entropy technique.
5 . The method as recited in claim 1 , wherein the diacritization model is trained using one or more of support vector machines, boosting, statistical decision trees, and/or their combinations to restore diacritics or to generate a diacritic lattice.
6 . The method as recited in claim 1 , wherein determining a best diacritic representation includes performing a Viterbi search to find the best diacritic representation assigned to the utterance.
7 . The method as recited in claim 1 , wherein the diacritization model includes one or more of previously predicted diacritics, graphemes, morphological units, words, part-of-speech (POS) tags, syntactic information, semantic analysis, and linguistic knowledge.
8 . The method as recited in claim 1 , wherein the diacritization model restores a complete set of diacritics.
9 . The method as recited in claim 1 , further comprising building a Diacritization Parse Tree (DPT) to extract information from in the form of features to train the diacritization model.
10 . The method as recited in claim 9 , wherein the diacritization model includes a statistical model able to learn the DPT, such that each node in the DPT can be predicted.
11 . The method as recited in claim 9 , wherein the diacritization model is constrained to learn leaves of the DPT that correspond to diacritics only.
12 . The method as recited in claim 1 , wherein making classification decisions further comprises building a Diacritization Parse Tree (DPT) for an input utterance to generate a diacritic lattice for each grapheme in the input utterance to determine a best diacritic for each grapheme.
13 . The method as recited in claim 1 , wherein the plurality of information sources include non-overlapping and overlapping sources of information.
14 . The method as recited in claim 1 , further comprising generating a lattice of diacritics for each utterance, and rescoring the diacritics with alternative models to improve the accuracy.
15 . The method as recited in claim 1 , wherein outputting a diacritically restored representation includes outputting one of text and speech.
16 . The method as recited in claim 1 , wherein the utterance includes one or text and speech.
17 . A computer program product for restoration of diacritics comprising a computer useable medium including a computer readable program, wherein the computer readable program when executed on a computer causes the computer to perform:
making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration; determining a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model; and outputting a diacritically restored representation of the utterance.
18 . The program product as recited in claim 17 , wherein the diacritization model restores a complete set of diacritics.
19 . The program product as recited in claim 17 , wherein making classification decisions further comprises building a Diacritization Parse Tree (DPT) for an input utterance to generate a diacritic lattice for each grapheme in the input utterance to determine a best diacritic for each grapheme.
20 . A system for restoration of diacritics, comprising:
a diacritization model configured to make classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources for diacritic restoration; a processing module configured to determine a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model to output a diacritically restored representation of the utterance.Join the waitlist — get patent alerts
Track US2008270115A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.