Pre-training method, pre-training device, and pre-training program
Abstract
A pre-training method executed by a training apparatus includes converting an input acoustic feature amount sequence into a corresponding intermediate acoustic feature amount sequence having a first length using a first conversion model to which a conversion model parameter is provided, converting a correct answer symbol sequence to generate a first frame unit symbol sequence having the first length and generating a second frame unit symbol sequence having the first length by delaying the first frame unit symbol sequence by one frame, converting the second frame unit symbol sequence into an intermediate character feature amount sequence having the first length using a second conversion model to which a character feature amount estimation model parameter is provided, and performing label estimation using an estimation model to which an estimation model parameter is provided based on the intermediate acoustic feature amount sequence and the intermediate character feature amount sequence.
Claims
exact text as granted — not AI-modified1 . A pre-training method executed by a training apparatus, the pre-training method comprising:
converting an input acoustic feature amount sequence into a corresponding intermediate acoustic feature amount sequence having a first length using a first conversion model to which a conversion model parameter is provided; converting a correct answer symbol sequence to generate a first frame unit symbol sequence having the first length and generating a second frame unit symbol sequence having the first length by delaying the first frame unit symbol sequence by one frame; converting the second frame unit symbol sequence into an intermediate character feature amount sequence having the first length using a second conversion model to which a character feature amount estimation model parameter is provided; performing label estimation using an estimation model to which an estimation model parameter is provided based on the intermediate acoustic feature amount sequence and the intermediate character feature amount sequence and outputting an output probability distribution of a two-dimensional matrix; and calculating a cross entropy (CE) loss of the output probability distribution with respect to the first frame unit symbol sequence based on the first frame unit symbol sequence and the output probability distribution.
2 . The pre-training method according to claim 1 , further including updating the conversion model parameter, the character feature amount estimation model parameter, and the estimation model parameter based on the CE loss and repeating the first conversion process, the second conversion process, the third conversion process, the estimation process, and the calculation process until a termination condition is satisfied.
3 . The pre-training method according to claim 2 , wherein the updating includes inputting the second frame unit symbol sequence having the first length to the third conversion process such that the first conversion model, the second conversion model, and the estimation model are pre-trained as an autoregressive model for predicting a next label.
4 . A pre-training apparatus comprising:
processing circuitry configured to: convert an input acoustic feature amount sequence into a corresponding intermediate acoustic feature amount sequence having a first length using a first conversion model to which a conversion model parameter is provided; convert a correct answer symbol sequence to generate a first frame unit symbol sequence having the first length and to generate a second frame unit symbol sequence having the first length by delaying the first frame unit symbol sequence by one frame; convert the second frame unit symbol sequence into an intermediate character feature amount sequence having the first length using a second conversion model to which a character feature amount estimation model parameter is provided; perform label estimation using an estimation model to which an estimation model parameter is provided based on the intermediate acoustic feature amount sequence and the intermediate character feature amount sequence and to output an output probability distribution of a two-dimensional matrix; and calculate a cross entropy (CE) loss of the output probability distribution with respect to the first frame unit symbol sequence based on the first frame unit symbol sequence and the output probability distribution.
5 . A non-transitory computer-readable recording medium storing therein a pre-training program that causes a computer to execute a process comprising:
converting an input acoustic feature amount sequence into a corresponding intermediate acoustic feature amount sequence having a first length using a first conversion model to which a conversion model parameter is provided; converting a correct answer symbol sequence to generate a first frame unit symbol sequence having the first length and generating a second frame unit symbol sequence having the first length by delaying the first frame unit symbol sequence by one frame; converting the second frame unit symbol sequence into an intermediate character feature amount sequence having the first length using a second conversion model to which a character feature amount estimation model parameter is provided; performing label estimation using an estimation model to which an estimation model parameter is provided based on the intermediate acoustic feature amount sequence and the intermediate character feature amount sequence and outputting an output probability distribution of a two-dimensional matrix; and calculating a cross entropy (CE) loss of the output probability distribution with respect to the first frame unit symbol sequence based on the first frame unit symbol sequence and the output probability distribution.Join the waitlist — get patent alerts
Track US2024071369A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.