US2024420682A1PendingUtilityA1

Method and apparatus for improving performance of artificial intelligence model using speech recognition results as text input

Assignee: ELECTRONICS & TELECOMMUNICATIONS RES INSTPriority: Jun 15, 2023Filed: Feb 23, 2024Published: Dec 19, 2024
Est. expiryJun 15, 2043(~16.9 yrs left)· nominal 20-yr term from priority
G10L 2015/0631G10L 2015/221G06N 20/00G10L 15/04G10L 15/22G10L 15/063G10L 15/26G10L 15/16
47
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The present disclosure relates to a method and device for improving the performance of an AI model that uses voice recognition results as text input. A method of training an AI model according to an embodiment of the present disclosure may include: generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; generating second time information by adding a pre-configured delay time to the first time information; generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and performing training of the AI model based on a second training sample including the voice and the modified transcription.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of training an artificial intelligence (AI) model, the method comprising:
 generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;   generating second time information by adding a pre-configured delay time to the first time information;   generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and   performing training of the AI model based on a second training sample including the voice and the modified transcription,   wherein the pre-configured delay time is variably adjusted depending on a degree of the training.   
     
     
         2 . The method of  claim 1 ,
 wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.   
     
     
         3 . The method of  claim 1 ,
 wherein the first time information includes information on an end time for each word for the plurality of words.   
     
     
         4 . The method of  claim 3 ,
 wherein the second time information is generated by adding the pre-configured delay time to an end time of each word for the plurality of words.   
     
     
         5 . The method of  claim 1 ,
 wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.   
     
     
         6 . The method of  claim 1 ,
 wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.   
     
     
         7 . The method of  claim 1 ,
 wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.   
     
     
         8 . The method of  claim 1 ,
 wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.   
     
     
         9 . The method of  claim 1 ,
 wherein the pre-configured delay time is set to gradually increase up to a delay time of the voice recognizer as a stage of training increases.   
     
     
         10 . An apparatus for training an artificial intelligence (AI) model, the apparatus comprising:
 a processor and a memory,   wherein the processor is configured to:
 generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; 
 generate second time information by adding a pre-configured delay time to the first time information; 
 generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and 
 perform training of the AI model based on a second training sample including the voice and the modified transcription, 
   wherein the pre-configured delay time is variably adjusted depending on a degree of the training.   
     
     
         11 . The apparatus of  claim 10 ,
 wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.   
     
     
         12 . The apparatus of  claim 10 ,
 wherein the first time information includes information on an end time for each word for the plurality of words.   
     
     
         13 . The apparatus of  claim 12 ,
 wherein the second time information is generated by adding the pre-configured delay time to an end time of each word for the plurality of words.   
     
     
         14 . The apparatus of  claim 10 ,
 wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.   
     
     
         15 . The apparatus of  claim 10 ,
 wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.   
     
     
         16 . The apparatus of  claim 10 ,
 wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.   
     
     
         17 . The apparatus of  claim 10 ,
 wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.   
     
     
         18 . The apparatus of  claim 10 ,
 wherein the pre-configured delay time is set to gradually increase up to a delay time of the voice recognizer as a stage of training increases.   
     
     
         19 . One or more non-transitory computer readable medium storing one or more instructions,
 wherein the one or more instructions are executed by one or more processors and control an apparatus for training an artificial intelligence (AI) model to:
 generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; 
 generate second time information by adding a pre-configured delay time to the first time information; 
 generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and 
 perform training of the AI model based on a second training sample including the voice and the modified transcription, 
   wherein the pre-configured delay time is variably adjusted depending on a degree of the training.   
     
     
         20 . The computer readable medium of  claim 19 ,
 wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.

Join the waitlist — get patent alerts

Track US2024420682A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.