US2026011337A1PendingUtilityA1

Noise reduction in audio mixing systems including a beamformer

Assignee: SYNAPTICS INCPriority: Jul 2, 2024Filed: Jul 2, 2024Published: Jan 8, 2026
Est. expiryJul 2, 2044(~17.9 yrs left)· nominal 20-yr term from priority
G10L 21/034G10L 25/18G10L 21/0364G10L 2021/02166G10L 25/30G10L 21/0232H04R 3/005G06N 3/044G10L 19/008H04R 2227/003H04R 27/00H04R 2430/23H04R 1/406H04R 2201/401G10L 21/0208
59
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

This disclosure provides methods, devices, and systems for audio signal mixing. The present implementations more specifically relate to mixing audio signals from a microphone array by performing fixed beamforming to generate beams, reducing noise on the beams, and mixing the beams to generate a final audio signal for playback. In some aspects, an audio mixing system includes a fixed beamformer to generate beams from audio signals from a microphone array and noise reduction units (NRUs) to reduce a noise component of each audio beam. The system also includes logic to calculate a signal characteristic of each reduced noise audio beam to determine, based on the signal characteristics, the reduced noise audio beams that include a speech component. The logic also generates a gain for each audio beam based on the selection, with the gains used in beam mixing. In some aspects, the NRU includes a neural network noise reduction unit.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A method of audio mixing, comprising:
 receiving a plurality of audio beams, wherein the audio beams are generated from a plurality of audio signals from a microphone array;   for each audio beam of the plurality of audio beams:
 generating a reduced noise audio beam from the audio beam by reducing a noise component of the audio beam; and 
 calculating a signal characteristic of the reduced noise audio beam; 
   determining, based on the plurality of signal characteristics of the plurality of reduced noise audio beams, one or more reduced noise audio beams of the plurality of reduced noise audio beams that include a speech component; and   for each audio beam of the plurality of audio beams, generating a gain for the audio beam based on the determination.   
     
     
         2 . The method of  claim 1 , further comprising:
 for each reduced noise audio beam of the plurality of reduced noise audio beams, generating a time measurement based on when the reduced noise audio beam includes the speech component, wherein generating the gain for the audio beam corresponding to the reduced noise audio beam is further based on the time measurement.   
     
     
         3 . The method of  claim 2 , wherein:
 generating a time measurement for the reduced noise audio beam includes counting by a counter a number of frames of the reduced noise audio beam that includes the speech component, wherein the counting includes:
 incrementing the counter by one or more in response to determining that the reduced noise audio beam includes the speech component in a current frame of the reduced noise audio beam; and 
 decrementing the counter by one or more in response to determining that the reduced noise audio beam does not include the speech component in the current frame of the reduced noise audio beam; and 
   generating the gain for the audio beam corresponding to the reduced noise audio beam includes reducing the gain towards zero based on the counter being at zero.   
     
     
         4 . The method of  claim 1 , wherein for each reduced noise audio beam of the plurality of reduced noise audio beams, the signal characteristic of the reduced noise audio beam includes one of:
 a signal level, wherein the signal level indicates an instantaneous signal power of the reduced noise audio beam; or   a signal-to-noise ratio (SNR), wherein the SNR indicates a ratio between the signal level of the reduced noise audio beam and a noise level of the noise component of the audio beam corresponding to the reduced noise audio beam.   
     
     
         5 . The method of  claim 4 , wherein for each reduced noise audio beam of the plurality of reduced noise audio beams, calculating the SNR of the reduced noise audio beam includes:
 calculating a first signal level of the audio beam corresponding to the reduced noise audio beam;   calculating a second signal level of the reduced noise audio beam;   calculating the noise level as a difference between the first signal level and the second signal level; and   calculating a ratio of the second signal level to the noise level as the SNR.   
     
     
         6 . The method of  claim 1 , wherein for each audio beam of the plurality of audio beams, reducing the noise component of the audio beam includes:
 inputting the audio beam to a neural network noise reduction unit (NNNRU) dedicated to processing the audio beam, wherein the NNNRU includes a recurrent neural network configured to receive samples of the audio beam based on a frequency spectrum of the audio beam; and   denoising the audio beam to generate the reduced noise audio beam by the NNNRU.   
     
     
         7 . The method of  claim 1 , further comprising:
 calculating a direction of arrival (DOA) of audio to the microphone array based on the one or more reduced noise audio beams that include the speech component; and   generating a control signal to control one or more of an audio unit or a video unit based on the DOA.   
     
     
         8 . The method of  claim 1 , further comprising mixing the plurality of audio beams to generate a mixed audio signal, wherein mixing the plurality of audio beams includes:
 for each audio beam of the plurality of audio beams, multiplying the audio beam with the gain for the audio beam to generate a processed audio beam; and   combining the plurality of processed audio beams to generate the mixed audio signal.   
     
     
         9 . The method of  claim 8 , further comprising reducing a noise in the mixed audio signal by a neural network noise reduction unit (NNNRU) to generate an output audio signal. 
     
     
         10 . The method of  claim 8 , further comprising:
 for each audio beam of the plurality of audio beams:
 receiving audio at one or more microphones of the microphone array; 
 for each microphone of the one or more microphones, generating an audio signal from the audio received at the microphone, wherein the plurality of audio signals includes the audio signal; and 
 generating, by a fixed beamformer, the audio beam from the one or more audio signals. 
   
     
     
         11 . An audio mixing system comprising:
 a processing system; and   a memory storing instructions that, when executed by the processing system, causes the audio mixing system to perform operations comprising:
 receiving a plurality of audio beams, wherein the audio beams are generated from a plurality of audio signals from a microphone array; 
 for each audio beam of the plurality of audio beams:
 generating a reduced noise audio beam from the audio beam by reducing a noise component of the audio beam; and 
 calculating a signal characteristic of the reduced noise audio beam; 
 
 determining, based on the plurality of signal characteristics of the plurality of reduced noise audio beams, one or more noise reduced audio beams of the plurality of reduced noise audio beams that include a speech component; and 
 for each audio beam of the plurality of audio beams, generating a gain for the audio beam based on the determination. 
   
     
     
         12 . The audio mixing system of  claim 11 , wherein the operations further comprise:
 for each reduced noise audio beam of the plurality of reduced noise audio beams, generating a time measurement based on when the reduced noise audio beam includes the speech component, wherein generating the gain for the audio beam corresponding to the reduced noise audio beam is further based on the time measurement.   
     
     
         13 . The audio mixing system of  claim 12 , wherein:
 generating a time measurement for the reduced noise audio beam includes counting by a counter a number of frames of the reduced noise audio beam that includes the speech component, wherein the counting includes:
 incrementing the counter by one or more in response to determining that the reduced noise audio beam includes the speech component in a current frame of the reduced noise audio beam; and 
 decrementing the counter by one or more in response to determining that the reduced noise audio beam does not include the speech component in the current frame of the reduced noise audio beam; and 
   generating the gain for the audio beam corresponding to the reduced noise audio beam includes reducing the gain towards zero based on the counter being at zero.   
     
     
         14 . The audio mixing system of  claim 11 , wherein for each reduced noise audio beam of the plurality of reduced noise audio beams, the signal characteristic of the reduced noise audio beam includes one of:
 a signal level, wherein the signal level indicates an instantaneous signal power of the reduced noise audio beam; or   a signal-to-noise ratio (SNR), wherein the SNR indicates a ratio between the signal level of the reduced noise audio beam and a noise level of the noise component of the audio beam corresponding to the reduced noise audio beam.   
     
     
         15 . The audio mixing system of  claim 14 , wherein for each reduced noise audio beam of the plurality of reduced noise audio beams, calculating the SNR of the reduced noise audio beam includes:
 calculating a first signal level of the audio beam corresponding to the reduced noise audio beam;   calculating a second signal level of the reduced noise audio beam;   calculating the noise level as a difference between the first signal level and the second signal level; and   calculating a ratio of the second signal level to the noise level as the SNR.   
     
     
         16 . The audio mixing system of  claim 11 , wherein for each audio beam of the plurality of audio beams, reducing the noise component of the audio beam includes:
 inputting the audio beam to a neural network noise reduction unit (NNNRU) dedicated to processing the audio beam, wherein the NNNRU includes a recurrent neural network configured to receive samples of the audio beam based on a frequency spectrum of the audio beam; and   denoising the audio beam to generate the reduced noise audio beam by the NNNRU.   
     
     
         17 . The audio mixing system of  claim 11 , wherein the operations further comprise:
 calculating a direction of arrival (DOA) of audio to the microphone array based on the one or more reduced noise audio beams that include the speech component; and   generating a control signal to control one or more of an audio unit or a video unit based on the DOA.   
     
     
         18 . The audio mixing system of  claim 11 , wherein the operations further comprise mixing the plurality of audio beams to generate a mixed audio signal, wherein mixing the plurality of audio beams includes:
 for each audio beam of the plurality of audio beams, multiplying the audio beam with the gain for the audio beam to generate a processed audio beam; and   combining the plurality of processed audio beams to generate the mixed audio signal.   
     
     
         19 . The audio mixing system of  claim 18 , wherein the operations further comprise reducing a noise in the mixed audio signal by a neural network noise reduction unit (NNNRU) to generate an output audio signal. 
     
     
         20 . The audio mixing system of  claim 18 , further comprising the microphone array, wherein the operations further comprise:
 receiving audio at one or more microphones of the microphone array;   for each microphone of the one or more microphones, generating an audio signal from the audio received at the microphone, wherein the plurality of audio signals includes the audio signal; and   generating, by a fixed beamformer, the audio beam from the one or more audio signals.

Join the waitlist — get patent alerts

Track US2026011337A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.