US2026038479A1PendingUtilityA1

Systems and methods for real-time accent mimicking

Assignee: SANAS AI INCPriority: Aug 1, 2024Filed: Aug 19, 2025Published: Feb 5, 2026
Est. expiryAug 1, 2044(~18 yrs left)· nominal 20-yr term from priority
G10L 13/08G10L 13/027G10L 2021/0135G10L 21/003
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

The disclosed technology relates to methods, speech processing systems, and non-transitory computer readable media for real-time accent mimicking. In some examples, trained machine learning model(s) are applied to first input audio data to extract accent features of first input speech associated with a first accent of a first user. Obtained second input data associated with second input speech associated with a second accent of a second user is analyzed to generate characteristics specific to a natural voice of the second user. A modified version of the second input speech is synthesized based on the generated characteristics and the extracted accent features. The modified version of the second input speech advantageously preserves aspects of the natural voice of the second user and mimics the first accent. Output audio data generated based on the modified version of the second input speech is provided for output via an audio output device.

Claims

exact text as granted — not AI-modified
What is claimed is: 
     
         1 . A speech processing system, comprising memory having instructions stored thereon and one or more processors coupled to the memory and configured to execute the instructions to:
 apply one or more machine learning models to first input audio data to generate accent features associated with a first accent of a first user;   analyze second input audio data associated with a second accent of a second user to generate vocal characteristics that are distinct to the second user and comprise a voice quality or one or more phonetic patterns, prosodic features, articulation styles, or intonation patterns;   modify the second input audio data based on the vocal characteristics and the accent features; and   generate output audio data based on the modified second audio data and output the output audio data.   
     
     
         2 . The speech processing system of  claim 1 , wherein the one or more processors are further configured to execute the instructions to extract from the first input audio data one or more other prosodic features, linguistic features, or global speaker characteristics. 
     
     
         3 . The speech processing system of  claim 1 , wherein the accent features comprise one or more pitch contours, other intonation patterns, or phoneme pronunciations and the pitch contours comprise variations in pitch throughout the first input speech, the other intonation patterns comprise the rise and fall of pitch at the ends of phrases or sentences, or the phoneme pronunciations comprise a unique production of phonemes in the first accent. 
     
     
         4 . The speech processing system of  claim 1 , wherein the one or more processors are further configured to execute the instructions to apply a mel frequency cepstral coefficient (MFCC) analysis to extract a unique fingerprint of a voice of the second user, wherein the vocal characteristics comprise the unique fingerprint. 
     
     
         5 . The speech processing system of  claim 1 , wherein the one or more processors are further configured to execute the instructions to apply a speaker identity encoding technique to encode speaker-specific voice characteristics, wherein the vocal characteristics comprise the speaker-specific voice characteristics. 
     
     
         6 . The speech processing system of  claim 1 , wherein the one or more processors are further configured to execute the instructions to receive the second input audio data via one or more communication networks and from a user computing device that is remote from the speech processing system and the second input audio data is captured at the user computing device. 
     
     
         7 . A method implemented by a speech processing system and comprising:
 applying one or more machine learning models to first input audio data to generate accent features associated with a first accent of a first user;   analyzing obtained second input audio data associated with a second accent of a second user to generate vocal characteristics specific to the second user, wherein the vocal characteristics comprise a unique fingerprint of a voice of the second user;   modifying the second input audio data based on the vocal characteristics and the accent features; and   providing output audio data generated based on the modified second input audio data.   
     
     
         8 . The method of  claim 7 , wherein the modified second input audio data preserves aspects of a natural voice of the second user and mimics the first accent. 
     
     
         9 . The method of  claim 7 , further comprising extracting from the first input audio data one or more prosodic features, linguistic features, or global speaker characteristics. 
     
     
         10 . The method of  claim 7 , wherein the accent features comprise one or more pitch contours, intonation patterns, or phoneme pronunciations and the pitch contours comprise variations in pitch, the intonation patterns comprise the rise and fall of pitch at the ends of phrases or sentences, or the phoneme pronunciations comprise a unique production of phonemes in the first accent. 
     
     
         11 . The method of  claim 7 , further comprising applying a mel frequency cepstral coefficient (MFCC) analysis to extract the unique fingerprint of the voice of the second user. 
     
     
         12 . The method of  claim 7 , further comprising applying a speaker identity encoding technique to encode speaker-specific voice characteristics, wherein the vocal characteristics comprise the speaker-specific voice characteristics. 
     
     
         13 . The method of  claim 7 , further comprising receiving the second input audio data via one or more communication networks and from a user device that is remote from the speech processing system, wherein the second input audio data is captured at the user device. 
     
     
         14 . A non-transitory computer-readable medium comprising instructions that, when executed by at least one processor, cause the at least one processor to:
 apply one or more machine learning models to first input audio data to generate accent features associated with a first accent of a first user;   analyze second input audio data associated with a second accent of a second user to generate speaker-specific voice characteristics specific to the second user;   modify the second input audio data based on the speaker-specific voice characteristics and the accent features; and   provide output audio data generated based on the modified second input audio data.   
     
     
         15 . The non-transitory computer-readable medium of  claim 14 , wherein the modified second input audio data preserves aspects of a natural voice of the second user and mimics the first accent. 
     
     
         16 . The non-transitory computer-readable medium of  claim 14 , wherein the instructions, when executed by the at least one processor further cause the at least one processor to extract from the first input audio data one or more prosodic features, linguistic features, or global speaker characteristics. 
     
     
         17 . The non-transitory computer-readable medium of  claim 14 , wherein the accent features comprise one or more pitch contours, intonation patterns, or phoneme pronunciations and the pitch contours comprise variations in pitch, the intonation patterns comprise the rise and fall of pitch at the ends of phrases or sentences, or the phoneme pronunciations comprise a unique production of phonemes in the first accent. 
     
     
         18 . The non-transitory computer-readable medium of  claim 14 , wherein the instructions, when executed by the at least one processor further causes the at least one processor to apply a mel frequency cepstral coefficient (MFCC) analysis to extract a unique fingerprint of a voice of the second user, wherein the speaker-specific voice characteristics comprise the unique fingerprint. 
     
     
         19 . The non-transitory computer-readable medium of  claim 14 , wherein the instructions, when executed by the at least one processor further cause the at least one processor to apply a speaker identity encoding technique to encode the speaker-specific voice characteristics. 
     
     
         20 . The non-transitory computer-readable medium of  claim 14 , wherein the instructions, when executed by the at least one processor further cause the at least one processor to receive the second input audio data via one or more communication networks and from a user device that is remote from the speech processing system, wherein the second input audio data is captured at the user device.

Join the waitlist — get patent alerts

Track US2026038479A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.