Method and apparatus for improving performance of artificial intelligence model using speech recognition results as text input
Abstract
The present disclosure relates to a method and device for improving the performance of an AI model that uses voice recognition results as text input. A method of training an AI model according to an embodiment of the present disclosure may include: generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; generating second time information by adding a pre-configured delay time to the first time information; generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and performing training of the AI model based on a second training sample including the voice and the modified transcription.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method of training an artificial intelligence (AI) model, the method comprising:
generating first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription; generating second time information by adding a pre-configured delay time to the first time information; generating a modified transcription based on an end time of a last word among the plurality of words and the second time information; and performing training of the AI model based on a second training sample including the voice and the modified transcription, wherein the pre-configured delay time is variably adjusted depending on a degree of the training.
2 . The method of claim 1 ,
wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.
3 . The method of claim 1 ,
wherein the first time information includes information on an end time for each word for the plurality of words.
4 . The method of claim 3 ,
wherein the second time information is generated by adding the pre-configured delay time to an end time of each word for the plurality of words.
5 . The method of claim 1 ,
wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.
6 . The method of claim 1 ,
wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.
7 . The method of claim 1 ,
wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
8 . The method of claim 1 ,
wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
9 . The method of claim 1 ,
wherein the pre-configured delay time is set to gradually increase up to a delay time of the voice recognizer as a stage of training increases.
10 . An apparatus for training an artificial intelligence (AI) model, the apparatus comprising:
a processor and a memory, wherein the processor is configured to:
generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;
generate second time information by adding a pre-configured delay time to the first time information;
generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and
perform training of the AI model based on a second training sample including the voice and the modified transcription,
wherein the pre-configured delay time is variably adjusted depending on a degree of the training.
11 . The apparatus of claim 10 ,
wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.
12 . The apparatus of claim 10 ,
wherein the first time information includes information on an end time for each word for the plurality of words.
13 . The apparatus of claim 12 ,
wherein the second time information is generated by adding the pre-configured delay time to an end time of each word for the plurality of words.
14 . The apparatus of claim 10 ,
wherein the modified transcription is generated by removing one or more words from among the plurality of words whose end time for each word is greater than or equal to the end time based on the second time information.
15 . The apparatus of claim 10 ,
wherein the pre-configured delay time is set to a value of 0 in an initial section of the training.
16 . The apparatus of claim 10 ,
wherein the pre-configured delay time is set between a 0 value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
17 . The apparatus of claim 10 ,
wherein the pre-configured delay time is set between a minimum delay time value and a maximum delay time value in a section where the training exceeds a pre-determined training stage.
18 . The apparatus of claim 10 ,
wherein the pre-configured delay time is set to gradually increase up to a delay time of the voice recognizer as a stage of training increases.
19 . One or more non-transitory computer readable medium storing one or more instructions,
wherein the one or more instructions are executed by one or more processors and control an apparatus for training an artificial intelligence (AI) model to:
generate first time information on a plurality of words included in a voice and transcription, using a first learning sample including the voice and the transcription;
generate second time information by adding a pre-configured delay time to the first time information;
generate a modified transcription based on an end time of a last word among the plurality of words and the second time information; and
perform training of the AI model based on a second training sample including the voice and the modified transcription,
wherein the pre-configured delay time is variably adjusted depending on a degree of the training.
20 . The computer readable medium of claim 19 ,
wherein the pre-configured delay time is related to a delay time for text output of a voice recognizer.Join the waitlist — get patent alerts
Track US2024420682A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.