US2008270115A1PendingUtilityA1

System and method for diacritization of text

Individually held — no corporate assignee on recordPriority: Mar 22, 2006Filed: Jun 3, 2008Published: Oct 30, 2008
Est. expiryMar 22, 2026(expired)· nominal 20-yr term from priority
G06F 40/53G06F 40/232
53
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for restoration of diacritics includes making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration. A best diacritic representation is determined for graphemes in the utterance based upon a best match with the diacritization model. A diacritically restored representation of the utterance is output.

Claims

exact text as granted — not AI-modified
1 . A method for restoration of diacritics, comprising:
 making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration;   determining a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model; and   outputting a diacritically restored representation of the utterance.   
     
     
         2 . The method as recited in  claim 1 , wherein the diacritization model includes information for a Semitic language. 
     
     
         3 . The method as recited in  claim 1 , further comprising integrating a plurality of sources of information to train the diacritization model. 
     
     
         4 . The method as recited in  claim 1 , wherein the diacritization model is trained using a Maximum Entropy technique. 
     
     
         5 . The method as recited in  claim 1 , wherein the diacritization model is trained using one or more of support vector machines, boosting, statistical decision trees, and/or their combinations to restore diacritics or to generate a diacritic lattice. 
     
     
         6 . The method as recited in  claim 1 , wherein determining a best diacritic representation includes performing a Viterbi search to find the best diacritic representation assigned to the utterance. 
     
     
         7 . The method as recited in  claim 1 , wherein the diacritization model includes one or more of previously predicted diacritics, graphemes, morphological units, words, part-of-speech (POS) tags, syntactic information, semantic analysis, and linguistic knowledge. 
     
     
         8 . The method as recited in  claim 1 , wherein the diacritization model restores a complete set of diacritics. 
     
     
         9 . The method as recited in  claim 1 , further comprising building a Diacritization Parse Tree (DPT) to extract information from in the form of features to train the diacritization model. 
     
     
         10 . The method as recited in  claim 9 , wherein the diacritization model includes a statistical model able to learn the DPT, such that each node in the DPT can be predicted. 
     
     
         11 . The method as recited in  claim 9 , wherein the diacritization model is constrained to learn leaves of the DPT that correspond to diacritics only. 
     
     
         12 . The method as recited in  claim 1 , wherein making classification decisions further comprises building a Diacritization Parse Tree (DPT) for an input utterance to generate a diacritic lattice for each grapheme in the input utterance to determine a best diacritic for each grapheme. 
     
     
         13 . The method as recited in  claim 1 , wherein the plurality of information sources include non-overlapping and overlapping sources of information. 
     
     
         14 . The method as recited in  claim 1 , further comprising generating a lattice of diacritics for each utterance, and rescoring the diacritics with alternative models to improve the accuracy. 
     
     
         15 . The method as recited in  claim 1 , wherein outputting a diacritically restored representation includes outputting one of text and speech. 
     
     
         16 . The method as recited in  claim 1 , wherein the utterance includes one or text and speech. 
     
     
         17 . A computer program product for restoration of diacritics comprising a computer useable medium including a computer readable program, wherein the computer readable program when executed on a computer causes the computer to perform:
 making classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources in a diacritization model for diacritic restoration;   determining a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model; and   outputting a diacritically restored representation of the utterance.   
     
     
         18 . The program product as recited in  claim 17 , wherein the diacritization model restores a complete set of diacritics. 
     
     
         19 . The program product as recited in  claim 17 , wherein making classification decisions further comprises building a Diacritization Parse Tree (DPT) for an input utterance to generate a diacritic lattice for each grapheme in the input utterance to determine a best diacritic for each grapheme. 
     
     
         20 . A system for restoration of diacritics, comprising:
 a diacritization model configured to make classification decisions regarding an utterance in accordance with an aggregate of a plurality of information sources for diacritic restoration;   a processing module configured to determine a best diacritic representation for graphemes in the utterance based upon a best match with the diacritization model to output a diacritically restored representation of the utterance.

Join the waitlist — get patent alerts

Track US2008270115A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.