US2009055158A1PendingUtilityA1

Speech translation apparatus and method

Assignee: TOSHIBA KKPriority: Aug 21, 2007Filed: Aug 21, 2008Published: Feb 26, 2009
Est. expiryAug 21, 2027(~1.1 yrs left)· nominal 20-yr term from priority
G10L 13/04G06F 40/58G10L 15/26G10L 19/09G10L 19/0018
46
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A speech translation apparatus includes a speech recognition unit configured to recognize input speech of a first language to generate a first text of the first language, an extraction unit configured to compare original prosody information of the input speech with first synthesized prosody information based on the first text to extract paralinguistic information about each of first words of the first text, a machine translation unit configured to translate the first text to a second text of a second language, a mapping unit configured to allocate the paralinguistic information about each of the first words to each of second words of the second text in accordance with synonymity, a generating unit configured to generate second synthesized prosody information based on the paralinguistic information allocated to each of the second words, and a speech synthesis unit configured to synthesize output speech based on the second synthesized prosody information.

Claims

exact text as granted — not AI-modified
1 . A speech translation apparatus comprising:
 a speech recognition unit configured to recognize input speech of a first language to generate a first text of the first language;   a prosody analysis unit configured to analyze a prosody of the input speech to obtain original prosody information;   a first language-analysis unit configured to split the first text into first words to obtain first linguistic information;   a first generating unit configured to generate first synthesized prosody information based on the first linguistic information;   an extraction unit configured to compare the original prosody information with the first synthesized prosody information to extract paralinguistic information about each of the first words;   a machine translation unit configured to translate the first text to a second text of a second language;   a second language-analysis unit configured to split the second text into second words to obtain second linguistic information;   a mapping unit configured to allocate the paralinguistic information about each of the first words to each of the second words in accordance with synonymity;   a second generating unit configured to generate second synthesized prosody information based on the second linguistic information and the paralinguistic information allocated to each of the second words; and   a speech synthesis unit configured to synthesize output speech based on the second linguistic information and the second synthesized prosody information.   
   
   
       2 . The apparatus according to  claim 1 , wherein the extraction unit normalizes the original prosody information to calculate a first characteristic quantity for each of the first words, and normalizes the first synthesized prosody information to calculate a second characteristic quantity for each of the first words, and compares the first characteristic quantity with the second characteristic quantity to extract the paralinguistic information about each of the first words. 
   
   
       3 . The apparatus according to  claim 1 , wherein the extraction unit normalizes the original prosody information to calculate a first characteristic quantity for each of the first words, and normalizes the first synthesized prosody information to calculate a second characteristic quantity about each of the first words, and compares the first characteristic quantity with the second characteristic quantity to extract the paralinguistic information about each of the first words; and the second generating unit generates third synthesized prosody information based on the second linguistic information, normalizes the third synthesized prosody information to calculate a third characteristic quantity for each of the second words, corrects the third characteristic quantity based on the paralinguistic information to calculate a fourth characteristic quantity, and uses the fourth characteristic quantity to generate the second synthesized prosody information. 
   
   
       4 . The apparatus according to  claim 3 , wherein the paralinguistic information is a value obtained by subtracting the second characteristic quantity from the first characteristic quantity, and the fourth characteristic quantity is a value obtained by adding the paralinguistic information to the third characteristic quantity. 
   
   
       5 . The apparatus according to  claim 4 , wherein the mapping unit allocates the paralinguistic information to each of the second words only when the paralinguistic information is a positive value. 
   
   
       6 . The apparatus according to  claim 3 , wherein the first characteristic quantity is a ratio of a peak value to a linear regression value of a basic frequency of the original prosody information for each of the first words; the second characteristic quantity is a ratio of a peak value to a linear regression value of a basic frequency of the first synthesized prosody information for each of the first words; and the third characteristic quantity is a ratio of a peak value to a linear regression value of a basic frequency of the third synthesized prosody information for each of the second words. 
   
   
       7 . The apparatus according to  claim 3 , wherein the first characteristic quantity is a ratio of a peak value to a linear regression value of an average power of the original prosody information for each of the first words; the second characteristic quantity is a ratio of a peak value to a linear regression value of an average power of the first synthesized prosody information for each of the first words; and the third characteristic quantity is a ratio of a peak value to a linear regression value of an average power of the third synthesized prosody information for each of the second words. 
   
   
       8 . The apparatus according to  claim 3 , wherein the first characteristic quantity is determined by a ratio of a duration of each of first phonetic units obtained by splitting each of the first words, to an average duration of the first phonetic units about the original prosody information; the second characteristic quantity is determined by a ratio of the duration of each the first phonetic units to an average duration of the first phonetic units about the first synthesized prosody information; and the third characteristic quantity is determined by a ratio of the duration of each of second phonetic units obtained by splitting each of the second word, to an average duration of the second phonetic units about the third synthesized prosody information. 
   
   
       9 . A speech translation method comprising:
 recognizing input speech of a first language to generate a first text of the first language;   analyzing a prosody of the input speech to obtain original prosody information;   splitting the first text into first words to obtain first linguistic information;   generating first synthesized prosody information based on the first linguistic information;   comparing the original prosody information with the first synthesized prosody information to extract paralinguistic information about each of the first words;   translating the first text to a second text of a second language;   splitting the second text into second words to obtain second linguistic information;   allocating the paralinguistic information about each of the first words to each of the second words in accordance with synonymity;   generating second synthesized prosody information based on the second linguistic information and the paralinguistic information allocated to each of the second words; and   synthesizing output speech based on the second linguistic information and the second synthesized prosody information.   
   
   
       10 . A computer readable storage medium storing instructions of a computer program which when executed by a computer results in performance of steps comprising:
 recognizing input speech of a first language to generate a first text of the first language;   analyzing a prosody of the input speech to obtain original prosody information;   splitting the first text into first words to obtain first linguistic information;   generating first synthesized prosody information based on the first linguistic information;   comparing the original prosody information with the first synthesized prosody information to extract paralinguistic information about each of the first words;   translating the first text to a second text of a second language;   splitting the second text into second words to obtain second linguistic information;   allocating the paralinguistic information about each of the first words to each of the second words in accordance with synonymity;   generating second synthesized prosody information based on the second linguistic information and the paralinguistic information allocated to each of the second words; and   synthesizing output speech based on the second linguistic information and the second synthesized prosody information.

Join the waitlist — get patent alerts

Track US2009055158A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.