System and method for adaptive multi-cultural searching and matching of personal names
Abstract
An automated name searching system incorporates an automatic name classifier and a multi-path architecture in which different algorithms are applied based on cultural identity of the query name. The name classifier operates with a preemptive list, analysis of morphological elements, length, and linguistic rules. A name regularizer produces a character based computational representation of the name. A pronunciation equivalent representation such as an IPA language representation, and language specific rules to generate name searching keys, are used in a first pass to eliminate database entries which are obviously not matches for the query name. The methods can also be implemented as a callable set of library routines including an intelligent preprocessor and a name evaluator that produces a score comparing a query name and database name, based on a variety of user-adjustable parameters. The user-controlled parameters permit tuning of the search methodologies for specific custom applications.
Claims
exact text as granted — not AI-modified1 .- 31 . (canceled)
32 . A method of providing an indication of whether an input name matches a known name, the method comprising:
accessing a text input name entered as an input name by one or more of a user or a system; determining multiple phonetic representations for a portion of the text input name, each of the multiple phonetic representations being for a different pronunciation of the text input name; comparing each of the multiple phonetic representations of the portion of the text input name to a phonetic representation of a portion of a text known name stored in a database; and providing an indication of whether the text input name matches the text known name based on the comparing.
33 . The method of claim 32 further comprising:
classifying the text input name as belonging to a particular culture; selecting a rule based on the classifying of the text input name; and applying the rule in determining the multiple phonetic representations for the portion of the text input name.
34 . The method of claim 32 further comprising:
classifying the text input name as belonging to a particular culture; selecting multiple rules based on the classifying of the text input name; and applying the multiple rules in determining the multiple phonetic representations for the portion of the text input name.
35 . The method of claim 32 wherein:
comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises determining articulatory similarity between at least one of the multiple phonetic representations of the portion of the text input name and the phonetic representation of the portion of the text known name, and providing the indication comprises providing an indication of articulatory similarity between the text input name and the text known name, the indication of articulatory similarity being based on the determining of articulatory similarity.
36 . The method of claim 35 further comprising:
identifying an articulatory variation between (i) one or more of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name; and classifying the articulatory variation as likely or unlikely, and wherein determining articulatory similarity comprises attributing less significance to the articulatory variation, so as to indicate greater articulatory similarity, if the articulatory variation is likely than if the articulatory variation is unlikely.
37 . The method of claim 35 wherein determining articulatory similarity comprises determining articulatory similarity based on a culture-specific rule.
38 . The method of claim 35 wherein:
determining articulatory similarity between at least one of the multiple phonetic representations of the portion of the text input name and the phonetic representation of the portion of the text known name comprises determining, for the at least one of the multiple phonetic representations of the portion of the text input name, how many phonetic features are in common between corresponding portions of the at least one phonetic representation of the portion of the text input name and the phonetic representation of the portion of the text known name, and providing the indication of articulatory similarity comprises providing an indication that is based on the determining of how many phonetic features are in common.
39 . The method of claim 38 wherein:
the at least one phonetic representation of the portion of the text input name comprises an International Phonetic Alphabet (“IPA”) representation of the text input name, the phonetic representation of the portion of the text known name comprises an IPA representation of the portion of the text known name, and determining how many phonetic features are in common between corresponding portions of the at least one phonetic representation of the portion of the text input name and the phonetic representation of the portion of the text known name comprises determining how many phonetic features are in common between corresponding symbols from the IPA representation of the portion of the text input name and the IPA representation of the portion of the text known name.
40 . The method of claim 39 wherein determining how many phonetic features are in common between corresponding symbols from the IPA representation of the portion of the text input name and the IPA representation of the portion of the text known name is based on a culture-specific rule.
41 . The method of claim 32 wherein determining multiple phonetic representations comprises determining multiple representations that are each based on an IPA.
42 . The method of claim 32 further comprising comparing each of the multiple phonetic representations of the portion of the text input name to a second phonetic representation of the portion of the text known name.
43 . The method of claim 32 wherein accessing the text input name comprises accessing a character representation of the text input name.
44 . The method of claim 43 wherein determining multiple phonetic representations comprises using a rule relating character representations to sounds.
45 . The method of claim 43 wherein: the character representation of the text input name reflects a spelling from a specific culture, and
determining multiple phonetic representations comprises using a rule for determining phonetic representations, the rule being based on the specific culture.
46 . The method of claim 43 wherein:
the character representation of the text input name reflects a spelling from a specific culture, the text input name belongs to another culture that is different from the specific culture, and determining multiple phonetic representations comprises using a rule for determining phonetic representations, the rule being based on the specific culture.
47 . The method of claim 43 wherein:
the character representation of the text input name reflects a spelling from a specific culture, the text input name belongs to another culture that is different from the specific culture, and determining multiple phonetic representations comprises using a rule for determining phonetic representations, the rule being based on the other culture.
48 . The method of claim 43 wherein:
the character representation of the text input name reflects a spelling from a specific culture, the text input name belongs to the specific culture, and determining multiple phonetic representations comprises using a rule for determining phonetic representations, the rule being based on the specific culture.
49 . The method of claim 32 wherein providing the indication comprises providing an indication that the text input name exactly matches the text known name.
50 . The method of claim 32 wherein providing the indication comprises providing an indication that the text input name does not exactly match the text known name.
51 . The method of claim 32 wherein comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises comparing, for at least one of the multiple phonetic representations of the portion of the text input name, corresponding parts of (i) the at least one phonetic representation of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name.
52 . The method of claim 51 wherein the corresponding parts include parts that correspond at a syntactic level
53 . The method of claim 51 wherein the corresponding parts include parts that correspond at a syllabic level.
54 . The method of claim 53 wherein the parts that correspond at the syllabic level include (i) a first part that relates to a left-most syllable of the portion of the text input name and (ii) a second part that relates to a left-most syllable of the portion of the text known name.
55 . The method of claim 54 wherein:
the first part further relates to both an initial phonologic element and a final phonologic element of the left-most syllable of the portion of the text input name, and the second part further relates to an initial phonologic element and a final phonologic element of the left-most syllable of the portion of the text known name.
56 . The method of claim 55 further comprising:
producing a result from the comparing of the first part and the second part; and determining, based on the result, whether to continue comparing the at least one of the multiple phonetic representations of the portion of the text input name and the phonetic representation of the portion of the text known name.
57 . The method of claim 51 wherein the corresponding parts include parts that correspond at a morphologic level.
58 . The method of claim 51 wherein the corresponding parts include parts that correspond at a phonologic level.
59 . The method of claim 58 wherein the parts that correspond at the phonologic level include (i) a first part that relates to a final phoneme of the portion of the text input name and (ii) a second part that relates to a final phoneme of the portion of the text known name.
60 . The method of claim 32 wherein comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises comparing, for at least one of the multiple phonetic representations of the portion of the text input name, sonority level between at least part of (i) the at least one of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name.
61 . The method of claim 32 wherein providing the indication of whether the text input name matches the text known name comprises providing a rank-ordered list of names, with rank-order indicating a likelihood of matching the text input name.
62 . The method of claim 61 wherein providing the rank-ordered list of names comprises ranking names on the rank-ordered list based on a degree of articulatory similarity between names on the rank-ordered list and the text input name.
63 . The method of claim 61 wherein the rank-ordered list of names includes the text known name.
64 . The method of claim 63 wherein providing the rank-ordered list comprises:
comparing, for at least one of the multiple phonetic representations of the portion of the text input name, sonority level between at least part of (i) the at least one of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name, and basing rank-order of the text known name on the comparing of sonority level.
65 . The method of claim 63 wherein providing the rank-ordered list comprises:
determining whether the text known name includes a morphological element, and basing rank-order of the text known name on whether the text known name includes a morphological element.
66 . The method of claim 63 wherein providing the rank-ordered list comprises:
comparing, for at least one of the multiple phonetic representations of the portion of the text input name, (i) an initial sound of the at least one of the multiple phonetic representations of the portion of the text input name and (ii) an initial sound of the phonetic representation of the portion of the text known name, and basing rank-order of the text known name on the comparing of initial sounds.
67 . The method of claim 63 wherein providing the rank-ordered list comprises:
comparing, for at least one of the multiple phonetic representations of the portion of the text input name, syllabic structure of (i) the at least one of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name, and basing rank-order of the text known name on the comparing of syllabic structure.
68 . The method of claim 67 wherein comparing syllabic structure comprises comparing syllabic similarity.
69 . The method of claim 63 wherein providing the rank-ordered list comprises:
comparing, for at least one of the multiple phonetic representations of the portion of the text input name, location of stress in (i) the at least one of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name, and basing rank-order of the text known name on the comparing of location of stress.
70 . The method of claim 63 wherein providing the rank-ordered list comprises:
comparing, for at least one of the multiple phonetic representations of the portion of the text input name, orthographic similarity between (i) the at least one of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name, and basing rank-order of the text known name on the comparing of orthographic similarity.
71 . The method of claim 32 wherein comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises discounting, for at least one of the multiple phonetic representations of the portion of the text input name, an occurrence of a likely articulatory variation between the at least one of the multiple phonetic representations of the portion of the text input name and the phonetic representation of the portion of the text known name.
72 . The method of claim 32 further comprising:
identifying a particle in the text input name; and attributing less significance to the particle, than to another part of the text input name, in providing the indication of whether the text input name matches the text known name.
73 . The method of claim 72 wherein attributing less significance to the particle comprises deciding not to determine a phonetic representation of the particle.
74 . The method of claim 72 wherein attributing less significance to the particle comprises deciding not to compare a phonetic representation of the particle to a phonetic representation of a part of the text known name.
75 . The method of claim 72 wherein identifying a particle comprises identifying a title, affix, or qualifier as the particle.
76 . The method of claim 32 wherein accessing an the text input name comprises accessing a portion of a complete name.
77 . The method of claim 32 wherein the portion of the text input name comprises the entire text input name.
78 . An apparatus comprising a computer readable storage medium having instructions stored thereon that when executed by a machine result in at least the following:
accessing a text input name entered as an input name by one or more of a user or a system; determining multiple phonetic representations for a portion of the text input name, each of the multiple phonetic representations being for a different pronunciation of the text input name; comparing each of the multiple phonetic representations of the portion of the text input name to a phonetic representation of a portion of a text known name stored in a database; and providing an indication of whether the text input name matches the text known name based on the comparing.
79 . The apparatus of claim 78 wherein the instructions, when executed by a machine, further result in at least the following:
classifying the text input name as belonging to a particular culture; selecting a rule based on the classifying of the text input name; and applying the rule in determining the multiple phonetic representations for the portion of the text input name.
80 . The apparatus of claim 78 wherein the instructions, when executed by a machine, further result in at least the following:
classifying the text input name as belonging to a particular culture; selecting multiple rules based on the classifying of the text input name; and applying the multiple rules in determining the multiple phonetic representations for the portion of the text input name.
81 . The apparatus of claim 78 wherein:
comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises determining articulatory similarity between at least one of the multiple phonetic representations of the portion of the text input name and the phonetic representation of the portion of the text known name, and providing the indication comprises providing an indication of articulatory similarity between the text input name and the text known name, the indication of articulatory similarity being based on the determining of articulatory similarity.
82 . The apparatus of claim 81 wherein the instructions, when executed by a machine, further result in at least the following:
identifying an articulatory variation between (i) one or more of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name; and classifying the articulatory variation as likely or unlikely, and wherein determining articulatory similarity comprises attributing less significance to the articulatory variation, so as to indicate greater articulatory similarity, if the articulatory variation is likely than if the articulatory variation is unlikely.
83 . The apparatus of claim 81 wherein determining articulatory similarity comprises determining articulatory similarity based on a culture-specific rule.
84 . The apparatus of claim 81 wherein:
determining articulatory similarity between at least one of the multiple phonetic representations of the portion of the text input name and the phonetic representation of the portion of the text known name comprises determining, for the at least one of the multiple phonetic representations of the portion of the text input name, how many phonetic features are in common between corresponding portions of the at least one phonetic representation of the portion of the text input name and the phonetic representation of the portion of the text known name, and providing the indication of articulatory similarity comprises providing an indication that is based on the determining of how many phonetic features are in common.
85 . The apparatus of claim 84 wherein:
the at least one phonetic representation of the portion of the text input name comprises an International Phonetic Alphabet (“IPA”) representation of the text input name, the phonetic representation of the portion of the text known name comprises an IPA representation of the portion of the text known name, and determining how many phonetic features are in common between corresponding portions of the at least one phonetic representation of the portion of the text input name and the phonetic representation of the portion of the text known name comprises determining how many phonetic features are in common between corresponding symbols from the IPA representation of the portion of the text input name and the IPA representation of the portion of the text known name.
86 . The apparatus of claim 85 wherein determining how many phonetic features are in common between corresponding symbols from the IPA representation of the portion of the text input name and the IPA representation of the portion of the text known name is based on a culture-specific rule.
87 . The apparatus of claim 78 wherein determining multiple phonetic representations comprises determining multiple representations that are each based on an IPA.
88 . The apparatus of claim 78 wherein the instructions, when executed by a machine, further result in comparing each of the multiple phonetic representations of the portion of the text input name to a second phonetic representation of the portion of the text known name.
89 . The apparatus of claim 78 wherein comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises comparing, for at least one of the multiple phonetic representations of the portion of the text input name, corresponding parts of (i) the at least one phonetic representation of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name.
90 . The apparatus of claim 89 wherein the corresponding parts include parts that correspond at a syllabic level.
91 . The apparatus of claim 90 wherein the parts that correspond at the syllabic level include (i) a first part that relates to a left-most syllable of the portion of the text input name and (ii) a second part that relates to a left-most syllable of the portion of the text known name.
92 . The apparatus of claim 78 wherein comparing each of the multiple phonetic representations of the portion of the text input name to the phonetic representation of the portion of the text known name comprises comparing, for at least one of the multiple phonetic representations of the portion of the text input name, sonority level between at least part of (i) the at least one of the multiple phonetic representations of the portion of the text input name and (ii) the phonetic representation of the portion of the text known name.
93 . The apparatus of claim 78 wherein providing the indication of whether the text input name matches the text known name comprises providing a rank-ordered list of names, with rank-order indicating a likelihood of matching the text input name.
94 . The apparatus of claim 93 wherein providing the rank-ordered list of names comprises ranking names on the rank-ordered list based on a degree of articulatory similarity between names on the rank-ordered list and the text input name.Join the waitlist — get patent alerts
Track US2007005567A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.