US2006167698A1PendingUtilityA1

System and method for generating an identification signal for electronic devices

Assignee: NELLYMOSER INC A MASSACHUSETTSPriority: Dec 31, 2001Filed: Mar 22, 2006Published: Jul 27, 2006
Est. expiryDec 31, 2021(expired)· nominal 20-yr term from priority
G10H 2240/056G10H 2250/291G10H 2230/021G10H 3/125G10H 2250/285H04M 19/041G10H 2250/235G10H 2250/265
41
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

A system and method for creating a ring tone for an electronic device takes as input a phrase sung in a human voice and transforms it into a control signal controlling, for example, a ringer on a cellular telephone. Time-varying features of the input signal are analyzed to segment the signal into a set of discrete notes and assigning to each note a chromatic pitch value. The set of note start and stop times and pitches are then translated into a format suitable for controlling the device.

Claims

exact text as granted — not AI-modified
1 . A method for generating an identification signal, comprising: 
 accepting as input a monophonic audio signal of limited duration;    translating said monophonic audio signal to a representation of a series of discrete tones; and    producing a control signal from said representation of discrete tones, said control signal suitable for causing a transponder to generate a signal,    where said generated signal is a translation of said monophonic audio signal;    wherein translating said monophonic audio signal to the representation of the series of discrete tones includes segmenting the monophonic audio signal into a series of segments according to time varying features of the audio signal that include a feature associated with energy and a feature associated with spectral composition, wherein each tone in the series of discrete tones is associated with a different segment in the series of segments.    
   
   
       2 . A method for generating an identification signal, comprising: 
 accepting as input a voice signal of limited duration;    translating said voice signal to a representation of a series of discrete tones; and    producing a control signal from said representation of discrete tones, said control signal suitable for causing a transponder to generate a signal,    where said generated signal is a translation of said voice signal;    wherein translating said voice signal to the representation of the series of discrete tones includes segmenting the voice signal into a series of segments according to time varying features of the voice signal that include a feature associated with energy and a feature associated with spectral composition, wherein each tone in the series of discrete tones is associated with a different segment in the series of segments.    
   
   
       3 . The method of  claim 2  wherein said generated signal is melodically human-recognizable.  
   
   
       4 . The method of  claim 2  wherein said generated signal is rhythmically human-recognizable.  
   
   
       5 . The method of  claim 2  wherein accepting as input further comprises receiving said voice signal over a telephone connection.  
   
   
       6 . The method of  claim 5  wherein said telephone connection is wireless.  
   
   
       7 . The method of  claim 2  wherein said step of accepting as input further comprises receiving said voice signal over a microphone attached to a computer.  
   
   
       8 . The method of  claim 2  wherein said translating step further comprises translating said voice signal to a range of tones within the capability of a mobile telephone audio output synthesizer.  
   
   
       9 . The method of  claim 2  further comprising the step of transmitting said control signal to a tone-producing output device responsive to said control signal.  
   
   
       10 . The method of  claim 2  wherein said translating step further comprises: 
 generating a digital representation of said voice signal;    dividing said digitized signal into a plurality of frames;    extracting analysis data from each said frame; and    formatting said analysis data into a frame representation.    
   
   
       11 . The method of  claim 10  further comprising the step of segmenting said signal by counting instances of increased signal amplitude in said frames, and 
 for each instance of increased amplitude, determining a change in each of pitch, energy, and spectral composition in a region around said instance of increased amplitude,    whereby a segment is defined by a start frame having an instance of increased amplitude and an end frame is defined by changes in pitch, energy and spectral composition in relation to selected thresholds.    
   
   
       12 . The method of  claim 10  wherein said translating step further comprises grouping said frames into a plurality of regions.  
   
   
       13 . The method of  claim 12  wherein each said region is determined from a count of consecutive upward short-term average change in cepstral-domain energy followed by a count of consecutive downward short-term average change in cepstral-domain energy.  
   
   
       14 . The method of  claim 12  further comprising the step of determining the existence of a candidate note start frame in each said region.  
   
   
       15 . The method of  claim 13  further comprising the step of determining a candidate note start frame in each said region as the last frame within said region in which the count of consecutive upward short-term average change in cepstral-domain energy is not zero.  
   
   
       16 . The method of  claim 14  further comprising the step of determining which regions of said plurality have a valid note start frame.  
   
   
       17 . The method of  claim 14 , wherein determining a candidate note start frame further comprises the step of determining if the cepstral domain energy of a particular frame is greater than a cepstral domain energy threshold and a frame immediately before said particular frame was below said cepstral domain energy threshold.  
   
   
       18 . The method of  claim 14 , wherein determining a candidate note start frame further comprises the step of determining whether a fundamental frequency range of a particular frame is above a fundamental frequency range threshold and whether an energy range for said particular frame is above an energy range threshold.  
   
   
       19 . The method of  claim 14 , further comprising the step of determining a stop frame corresponding to each start frame.  
   
   
       20 . The method of  claim 15 , further comprising the step of determining a stop frame by locating the first frame after a start frame in which cepstral energy is below said cepstral domain energy threshold.  
   
   
       21 . The method of  claim 20 , further comprising the step of defining the stop frame as a frame between two and ten frames before a subsequent start frame if no frame having cepstral energy below said cepstral domain energy threshold is found.  
   
   
       22 . The method of  claim 19  further comprising the step of verifying each start and stop frame pair by determining whether 
 a) average voicing probability is above a voicing probability threshold,    b) average short-time energy is above an average short-time energy threshold, and    c) average fundamental frequency is above an average fundamental frequency threshold.    
   
   
       23 . The method of  claim 2  wherein the feature associated with energy includes a time-domain energy.  
   
   
       24 . The method of  claim 2  wherein the feature associated with energy includes a cepstral-domain energy.  
   
   
       25 . The method of  claim 2  wherein the time varying features according to which the voice signal is segmented include at least two features associated with energy.  
   
   
       26 . The method of clam  2  wherein the feature associated with spectral composition includes a cepstral coefficient.  
   
   
       27 . The method of  claim 2  wherein the time varying features according to which the voice signal is segmented further include a feature associated with periodicity.  
   
   
       28 . The method of  claim 27  wherein the feature associated with periodicity includes a fundamental frequency.  
   
   
       29 . The method of  claim 27  wherein feature associated with periodicity includes a voicing probability.  
   
   
       30 . Apparatus for generating an identification signal comprising: 
 a voice signal receiver;    a translator having as its input a voice signal received by said voice signal receiver and having as its output a representation of discrete tones where an audio presentation of said discrete tones would be human-recognizable as a translation of said voice signal;    wherein the translator includes an estimation module with outputs of a time varying feature associated with each of energy and spectral composition from the voice signal and a segmentation module responsive to the time varying features with an output of a segmentation of the voice signal into a series of segments according to the time varying features, such that each in the series of output discrete tones is associated with a different segment in the series of segments.    
   
   
       31 . The apparatus of  claim 30  wherein said voice signal receiver comprises an analog telephone receiver.  
   
   
       32 . The apparatus of  claim 30  wherein said voice signal receiver further comprises a voice-to-digital signal transducer.  
   
   
       33 . The apparatus of  claim 30  wherein said voice signal receiver further comprises a recording device.  
   
   
       34 . The apparatus of  claim 30  wherein said translator further comprises a feature estimation module to determine values for at least one time-varying feature of said input signal.  
   
   
       35 . The apparatus of  claim 34  wherein said translator further comprises a pitch assignment module responsive to signal energy in each segment output by said segmentation module.  
   
   
       36 . The apparatus of  claim 34  wherein said feature estimation module further comprises a primary feature module, a secondary feature module and a tertiary feature module.  
   
   
       37 . The apparatus of  claim 36  wherein said primary feature module determines a plurality of values for each of time-domain energy, fundamental frequency, cepstral-domain energy, and voicing probability.  
   
   
       38 . The apparatus of  claim 35  wherein said segmentation module further comprises a first-phase segmentation module and a second-phase segmentation module.  
   
   
       39 . The apparatus of  claim 38  wherein said first-phase segmentation module groups a plurality of successive frames of said input signal into at least one region in response to output of said feature estimation module.  
   
   
       40 . The apparatus of  claim 39  wherein said region is a plurality of frames in which a change in energy increases immediately followed by frames in which change in energy decreases.  
   
   
       41 . The apparatus of  claim 40  in which said region has a minimum rider of frames.  
   
   
       42 . The apparatus of  claim 39  wherein said second-phase segmentation module determines if said at least one region has a valid note start frame and if so, determines a stop frame.  
   
   
       43 . The apparatus of  claim 42  wherein said second-phase segmentation module determines said valid note start frame in response to cepstral domain energy by determining whether a frame has a cepstral domain energy greater than a cepstral domain energy threshold preceded by a frame having a cepstral domain energy less than said cepstral domain threshold.  
   
   
       44 . The apparatus of  claim 42  wherein said second-phase segmentation module determines a valid note start frame if the fundamental frequency exceeds a fundamental energy threshold and if the non-cepstral domain energy exceeds an energy threshold.  
   
   
       45 . The apparatus of  claim 39  further comprising a segmentation post-processor to verify said start and stop frame in response to average voicing probability, average short-time energy, and average fundamental frequency of said start and stop frame.  
   
   
       46 . The apparatus of  claim 35  wherein said pitch assignment module assigns an integer between 32 and 83, said integer corresponding to the MIDI note number for pitch.  
   
   
       47 . The apparatus of  claim 35  wherein said pitch assignment module comprises an intranote pitch assignment subsystem and an internote pitch assignment subsystem.  
   
   
       48 . The apparatus of  claim 47  wherein said internote pitch assignment subsystem corrects pitches determined by said intranote pitch assignment subsystem.  
   
   
       49 . The apparatus of  claim 48  wherein said internote pitch assignment subsystem further comprises a key finding stage to assign a scale to a note sequence output by said intranote pitch assignment subsystem.  
   
   
       50 . The apparatus of  claim 48  wherein said internote pitch assignment subsystem further comprises a pairwise correction stage to examine a pitch and its preceding pitch for conformity to voice-leading rules, 
 if a pair is determined to be dissonant according to said voice-leading rules, the internote pitch assignment subsystem corrects the pitches of said pair if the pitch adjustment does not cause dissonance in an adjacent pair.

Join the waitlist — get patent alerts

Track US2006167698A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.