US2023402046A1PendingUtilityA1

Audio encoder and decoder using a frequency domain processor with full-band gap filling and a time domain processor

Assignee: FRAUNHOFER GES FORSCHUNGPriority: Jul 28, 2014Filed: Aug 25, 2023Published: Dec 14, 2023
Est. expiryJul 28, 2034(~8 yrs left)· nominal 20-yr term from priority
G10L 21/038G10L 19/18G10L 19/028G10L 19/032G10L 19/06G10L 19/265G10L 19/20G10L 19/02G10L 19/04G10L 19/24
78
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

An audio encoder for encoding an audio signal has: a first encoding processor for encoding a first audio signal portion in a frequency domain, having: a time frequency converter for converting the first audio signal portion into a frequency domain representation; an analyzer for analyzing the frequency domain representation to determine first spectral portions to be encoded with a first spectral resolution and second regions to be encoded with a second resolution; and a spectral encoder for encoding the first spectral portions with the first spectral resolution and encoding the second portions with the second resolution; a second encoding processor for encoding a second different audio signal portion in the time domain; a controller for analyzing and determining, which portion of the audio signal is the first audio signal portion encoded in the frequency domain and which portion is the second audio signal portion encoded in the time domain; and an encoded signal former for forming an encoded audio signal having a first encoded signal portion for the first audio signal portion and a second encoded signal portion for the second portion.

Claims

exact text as granted — not AI-modified
1 . An audio encoder for encoding an audio signal, the audio signal comprising a first audio signal portion and a timely subsequent second audio signal portion having an audio sampling rate, to generate an encoded audio signal, comprising:
 a first encoding processor configured for encoding the first audio signal portion in a frequency domain to obtain a first encoded signal portion;   a second encoding processor configured for encoding the second audio signal portion in a time domain to obtain a second encoded signal portion, the second audio signal portion comprising a low band and a high band; and   a controller configured for analyzing a portion of the audio signal and for determining, that the portion of the audio signal is either the first audio signal portion encoded in the frequency domain or the second audio signal portion encoded in the time domain; and   an encoded signal former configured for forming the encoded audio signal comprising the first encoded signal portion for the first audio signal portion and the second encoded signal portion for the second audio signal portion.   
     
     
         2 . The audio encoder of  claim 1 , further comprising:
 a preprocessor configured for preprocessing the first audio signal portion and the second audio signal portion,   wherein the preprocessor comprises:
 a prediction analyzer configured for determining prediction coefficients; and 
   wherein the second encoding processor comprises:
 a prediction coefficient quantizer configured for generating a quantized version of the prediction coefficients; and 
 an entropy coder configured for generating an encoded version of the quantized prediction coefficients, 
   wherein the encoded signal former is configured for introducing the encoded version of the quantized prediction coefficients into the encoded audio signal.   
     
     
         3 . The audio encoder of  claim 1 , comprising a preprocessor,
 wherein the preprocessor comprises a resampler configured for resampling the audio signal to a lower sampling rate of the second encoding processor to obtain a resampled audio signal; and   wherein a prediction analyzer is configured to determine prediction coefficients using the resampled audio signal.   
     
     
         4 . The audio encoder of  claim 1 , comprising a preprocessor, wherein the preprocessor comprises a long term prediction analysis stage configured for determining one or more long term prediction parameters for the first audio signal portion. 
     
     
         5 . The audio encoder of  claim 1 , further comprising a cross-processor configured for calculating, from an encoded spectral representation of the first audio signal portion, initialization data of the second encoding processor, so that the second encoding processor is initialized to encode the second audio signal portion immediately following the first audio signal portion in time in the audio signal. 
     
     
         6 . The audio encoder of  claim 5 , wherein the cross-processor comprises:
 a spectral decoder configured for calculating a decoded version of the first encoded signal portion;   a weighted prediction coefficient analysis filtering block configured for filtering the decoded version of the first encoded signal portion to obtain a filtered decoded version and for feeding the filtered version into a codebook determinator of the second encoding processor for an initialization.   
     
     
         7 . The audio encoder of  claim 5 , wherein the cross-processor comprises:
 a spectral decoder for calculating a decoded version of the first encoded signal portion; and   an analysis filtering stage configured for filtering the decoded version of the first encoded signal portion or a pre-emphasized decoded version of the first encoded signal portion to obtain a filter residual and configured for feeding the filter residual into an adaptive codebook determinator of the second encoding processor for an initialization.   
     
     
         8 . The audio encoder of  claim 5 , wherein the cross-processor comprises:
 a spectral decoder configured for calculating a decoded version of the first encoded signal portion;   a pre-emphasis filter configured for filtering the decoded version of the first encoded signal portion to obtain a pre-emphasized version and for feeding the pre-emphasized version or a delayed pre-emphasized version to a synthesis filtering stage of the second encoding processor for an initialization.   
     
     
         9 . The audio encoder of  claim 1 ,
 wherein an analyzer is configured to perform a temporal tile shaping or temporal noise shaping analysis or an operation of setting to zero spectral values in the second spectral portions,   wherein the first encoding processor is configured to perform a shaping of spectral values of first spectral portions using prediction coefficients derived from the first audio signal portion, and wherein the first encoding processor is furthermore configured to perform a quantization and entropy coding operation of shaped spectral values of the first spectral portions, and   wherein spectral values of the second spectral portions are set to zero.   
     
     
         10 . The audio encoder of  claim 1 , wherein the second encoding processor comprises at least one block of the following group of blocks:
 a prediction analysis filter;   an ACELP/gain coding stage; and   a prediction synthesis filtering stage.   
     
     
         11 . The audio encoder of  claim 1 , wherein the second encoding processor comprises at least one block of the following group of blocks:
 an adaptive codebook stage;   an innovative codebook stage;   an estimator configured for estimating an innovative codebook entry;   
     
     
         12 . The audio encoder of  claim 1 , wherein the second encoding processor comprises at least one block of the following group of blocks:
 a de-emphasis stage; and   a bass post-filter analysis stage.   
     
     
         13 . The audio encoder of  claim 1 ,
 wherein the time frequency converter is configured for converting the first audio signal portion into the frequency domain representation comprising spectral lines up to a maximum frequency of the first audio signal portion, and   wherein the analyzer is configured for analyzing the frequency domain representation up to the maximum frequency.   
     
     
         14 . The audio encoder of  claim 1 , wherein the second encoding processor comprises:
 a sampling rate converter configured for converting the second audio signal portion to a lower sampling rate representation of the second audio signal portion, wherein the sampling rate converter is configured so that a lower sampling rate of the lower sampling rate representation is lower than the audio sampling rate of the second audio signal portion, and so that the lower sampling rate representation of the second audio signal portion comprises the low band of the second audio signal portion and does not comprise the high band of the second audio signal portion;   a time domain low band encoder configured for time domain encoding the lower sampling rate representation of the second audio signal portion; and   a time domain bandwidth extension encoder configured for parametrically encoding the high band of the second audio signal portion.   
     
     
         15 . An audio decoder for decoding an encoded audio signal comprising a first encoded audio signal portion and a second encoded audio signal portion to obtain a decoded audio signal, comprising:
 a first decoding processor configured for decoding the first encoded audio signal portion in a frequency domain to obtain a decoded time domain first audio signal portion;   a second decoding processor configured for decoding the second encoded audio signal portion in the time domain to acquire a decoded time domain second audio signal portion having a low band and a high band; and   a combiner configured for combining the decoded time domain first audio signal portion and the decoded time domain second audio signal portion to acquire the decoded audio signal.   
     
     
         16 . The audio decoder of  claim 15 ,
 wherein an upsampler of the second decoding processor comprises an analysis filterbank operating at a first sampling rate and a synthesis filterbank operating at a second sampling rate.   
     
     
         17 . The audio decoder of  claim 15 ,
 wherein the time domain low band decoder comprises a decoder and a synthesis filter configured for filtering a residual signal using synthesis filter coefficients,   wherein a time domain bandwidth extension decoder of the second decoding processor is configured to upsample the residual signal to obtain an upsampled residual signal and to process the upsampled residual signal using a non-linear operation to acquire a high band residual signal, and to spectrally shape the high band residual signal to acquire a high band of the decoded time domain second audio signal portion having a second sampling rate.   
     
     
         18 . The audio decoder of  claim 15 ,
 wherein the first decoding processor comprises an adaptive long term prediction post-filter configured for post-filtering the decoded first audio signal portion, wherein the adaptive long term prediction post-filter is controlled by one or more long term prediction parameters comprised in the encoded audio signal.   
     
     
         19 . The audio decoder of  claim 15 , wherein the second decoding processor comprises:
 a time domain low band decoder configured for decoding to obtain a low band time domain signal having a first sampling rate;   an upsampler configured for upsampling the low band time domain signal to obtain an upsampled low band time domain signal having a second sampling rate being higher than the first sampling rate, the upsampled low band time domain signal representing the low band of the decoded time domain second audio signal portion;   a time domain bandwidth extension decoder configured for synthesizing the high band of the decoded time domain second audio signal portion having the second sampling rate using the low band time domain signal; and   a mixer configured for mixing the high band of the decoded time domain second audio signal portion having the second sampling rate and the upsampled low band time domain signal having the second sampling rate to obtain the decoded time domain second audio signal portion.   
     
     
         20 . The audio decoder of  claim 15 , wherein the first decoding processor comprises:
 a spectral decoder configured for decoding first spectral portions with a high spectral resolution and for synthesizing second spectral portions using a parametric representation of the second spectral portions and at least a decoded first spectral portion to acquire a decoded spectral representation; and   a frequency-time converter configured for converting the decoded spectral representation into a time domain to acquire a decoded time domain first audio signal portion.   
     
     
         21 . The audio decoder of  claim 15 , wherein the second decoding processor comprises: an adaptive codebook synthesis stage. 
     
     
         22 . The audio decoder of  claim 15 , wherein the second decoding processor comprises an ACELP block configured for decoding gains and an innovative codebook. 
     
     
         23 . The audio decoder of  claim 15 , wherein the second decoding processor comprises at least one block of the group of blocks comprising:
 an ACELP post-processor; and   a de-emphasis stage.   
     
     
         24 . The audio decoder of  claim 15 , wherein the second decoding processor comprises a prediction synthesis filter. 
     
     
         25 . The audio decoder of  claim 15 , comprising: a cross-processor configured for calculating, from the decoded spectral representation of the first encoded audio signal portion, initialization data of the second decoding processor, so that the second decoding processor is initialized to decode the encoded second audio signal portion following in time the first audio signal portion in the encoded audio signal, wherein the cross-processor comprises:
 a frequency-time converter configured for converting the decoded spectral representation to obtain a further decoded first signal portion in the time domain;   a delay stage configured for delaying the further decoded first signal portion to obtain a delayed version and for feeding the delayed version into a de-emphasis stage of the second decoding processor for an initialization.   
     
     
         26 . The audio decoder of  claim 15 , comprising a cross-processor configured for calculating, from the decoded spectral representation of the first encoded audio signal portion, initialization data of the second decoding processor, so that the second decoding processor is initialized to decode the encoded second audio signal portion following in time the first audio signal portion in the encoded audio signal, wherein the cross-processor comprises:
 a frequency-time converter configured for converting the decoded spectral representation to obtain a further decoded first signal portion in the time domain;   a prediction analysis filter configured for generating a prediction residual signal from the further decoded first signal portion or a pre-emphasized further decoded first signal portion and for feeding the prediction residual signal into a codebook synthesizer of the second decoding processor.   
     
     
         27 . A method of encoding an audio signal, the audio signal comprising a first audio signal portion and a timely subsequent second audio signal portion having an audio sampling rate, to generate an encoded audio signal, comprising:
 first encoding the first audio signal portion in a frequency domain to obtain a first encoded signal portion;   second encoding the second audio signal portion in a time domain to obtain a second encoded signal portion, the second audio signal portion comprising a low band and a high band;   analyzing a portion of the audio signal and determining that the portion of the audio signal is either the first audio signal portion encoded in the frequency domain or is the second audio signal portion encoded in the time domain; and   forming the encoded audio signal comprising the first encoded signal portion for the first audio signal portion and the second encoded signal portion for the second audio signal portion.   
     
     
         28 . A method of decoding an encoded audio signal comprising a first encoded audio signal portion and a second encoded audio signal portion to obtain a decoded audio signal, comprising:
 first decoding the first encoded audio signal portion in a frequency domain to acquire a decoded time domain first audio signal portion;   second decoding the second encoded audio signal portion in the time domain to acquire a decoded second time domain audio signal portion having a low band and a high band; and   combining the decoded time domain first audio signal portion and the decoded time domain second audio signal portion to acquire the decoded audio signal.   
     
     
         29 . A non-transitory digital storage medium having stored thereon a computer program for performing, when running on a computer, a method of encoding an audio signal, the audio signal comprising a first audio signal portion and a timely subsequent second audio signal portion having an audio sampling rate, to generate an encoded audio signal, the method comprising:
 first encoding the first audio signal portion in a frequency domain to obtain a first encoded signal portion;   second encoding the second audio signal portion in a time domain to obtain a second encoded signal portion, the second audio signal portion comprising a low band and a high band;   analyzing a portion of the audio signal and determining that the portion of the audio signal is either the first audio signal portion encoded in the frequency domain or the second audio signal portion encoded in the time domain; and   forming the encoded audio signal comprising the first encoded signal portion for the first audio signal portion and the second encoded signal portion for the second audio signal portion.   
     
     
         30 . A non-transitory digital storage medium having stored thereon a computer program for performing, when running on a computer, a method of decoding an encoded audio signal comprising a first encoded audio signal portion and a second encoded audio signal portion to obtain a decoded audio signal, the method comprising:
 first decoding the first encoded audio signal portion in a frequency domain to acquire a decoded time domain first audio signal portion;   second decoding the second encoded audio signal portion in the time domain to acquire a decoded time domain second audio signal portion having a low band and a high band; and   combining the decoded time domain first audio signal portion and the decoded time domain second audio signal portion to acquire the decoded audio signal.

Join the waitlist — get patent alerts

Track US2023402046A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.