Speech and Textual Analysis Device and Corresponding Method
Abstract
A speech and textual analysis device and method for forming a search and/or classification catalog. The device is based on a linguistic database and includes a taxonomy table containing variable taxon nodes. The speech and textual analysis device includes a weighting module, a weighting parameter being additionally assigned to each taxon node to register the recurrence frequency of terms in the linguistic and/or textual data that is to be classified and/or sorted. The speech and/or textual analysis device includes an integration module for determining a predefinable number of agglomerates based on the weighting parameters of the taxon nodes in the taxonomy table and at least one neuronal network module for classifying and/or sorting the speech and/or textual data based on the agglomerates in the taxonomy table.
Claims
exact text as granted — not AI-modified1 - 22 . (canceled)
23 . A language and text analysis apparatus for formation of a search and/or classification catalog, comprising:
at least one linguistic databank for association of linguistic terms with data records, in which case the language and text analysis apparatus can be used to classify and/or to sort language and/or text data corresponding to the data records, and in which the linguistic terms comprise at least keywords and/or search terms, wherein: the language and text analysis apparatus includes a taxonomy table with variable taxon nodes on the basis of the linguistic databank, in which case one or more data records can be associated with one taxon node in the taxonomy table, and in which case each data record includes a variable significance factor for weighting of terms on the basis of at least filling words and/or linking words and/or keywords, the language and text analysis apparatus includes a weighting module, in which a weighting parameter for recording of frequencies of occurrence of terms within the language and/or text data to be sorted and/or to be classified is additionally stored associated with each taxon node, the language and/or text analysis apparatus includes an integration module for determination of a predefinable number of agglomerates on the basis of the weighting parameters of the taxon nodes in the taxonomy table, with one agglomerate including at least one taxon node, and the language and/or text analysis apparatus includes at least one neural network module for classification and/or for sorting of the language and/or text data on the basis of the agglomerates in the taxonomy table.
24 . The language and text analysis apparatus as claimed in claim 23 , wherein the neural network module includes at least one self-organizing Kohonen map.
25 . The language and text analysis apparatus as claimed in claim 23 , wherein the language and text analysis apparatus includes an entropy module for determination of an entropy parameter, which can be stored in a memory module, on the basis of distribution of a data record in the language and/or text data.
26 . The language and text analysis apparatus as claimed in claim 23 , wherein the linguistic databank includes multilingual data records.
27 . The language and text analysis apparatus as claimed in claim 23 , wherein the language and text analysis apparatus includes a hash table that is associated with the linguistic databank in which case a hash value can be used to identify linguistically linked data records in the hash table.
28 . The language and text analysis apparatus as claimed in claim 23 , wherein a language parameter can be used to associate the data records with a language and can be identified as a synonym in the taxonomy table.
29 . The language and text analysis apparatus as claimed in claim 23 , wherein the entropy parameter is given by:
Entropy DR =In(freqsum DR )−Σ F DR In( F DR )/freqsum DR in which freqsum isyn =Σ F isyn ( ) and F isyn is the frequency for each synonym group and each item of language and/or text data.
30 . The language and text analysis apparatus as claimed in claim 23 , wherein the agglomerates form an n-dimensional content space.
31 . The language and text analysis apparatus as claimed in claim 30 , wherein n is equal to 100.
32 . The language and text analysis apparatus as claimed in claim 23 , wherein the language and text analysis apparatus includes descriptors by which constraints that correspond to definable descriptors can be determined for a subject group.
33 . An automated language and text analysis method for formation of a search and/or classification catalog, with a linguistic databank being used to record data records and to classify and/or sort language and/or text data on the basis of the data records, wherein:
the data records in the linguistic databank are stored associated with a taxon node in a taxonomy table, with each data record including a variable significance factor for weighting of terms based at least on filling words and/or linking words and/or keywords, the language and/or text data is recorded on the basis of the taxonomy table, with the frequency of individual data records in the language and/or text data being determined by a weighting module and being associated with a weighting parameter for the taxon node, an integration module is used to determine a determinable number of agglomerates in the taxonomy table on the basis of the weighting parameters of one or more taxon nodes, a neural network module is used to classify and/or sort the language and/or text data on the basis of the agglomerates in the taxonomy table.
34 . The automated language and text analysis method as claimed in claim 33 , wherein the neural network module includes at least one self-organizing Kohonen map.
35 . The automated language and text analysis method as claimed in claim 33 , wherein an entropy module is used to determine an entropy factor on the basis of distribution of a data record in the language and/or text data.
36 . The automated language and text analysis method as claimed in claim 33 , wherein the linguistic databank includes multilingual data records.
37 . The automated language and text analysis method as claimed in claim 33 , wherein a hash table is stored associated with the linguistic databank, with the hash table including an identification of linked data records by a hash value.
38 . The automated language and text analysis method as claimed in claim 33 , wherein the data records can be associated with a language and can be weighted synonymously in the taxonomy table by a language parameter.
39 . The automated language and text analysis method as claimed in claim 33 , wherein the entropy factor is given by the term
Entropy DR =In(freqsum DR )−Σ F DR In( F DR )/freqsum DR in which freqsum isyn =Σ F DR ( ) and F isyn is the frequency for each synonym group and each item of language and/or text data.
40 . The automated language and text analysis method as claimed in claim 33 , wherein the agglomerates form an n-dimensional content space.
41 . The automated language and text analysis method as claimed in claim 40 , wherein n is equal to 100.
42 . The automated language and text analysis method as claimed in claim 33 , wherein definable descriptors can be used to determine corresponding constraints for a subject group.
43 . A computer program product on a computer-readable medium with computer program code means contained therein to control one or more processors in a computer-based system for automated language and text analysis by formation of a search and/or classification catalog, with data records being recorded on the basis of a linguistic databank, and with language and/or text data being classified and/or sorted on the basis of the data records,
wherein: the computer program product can be used to store the data records in the linguistic databank associated with a taxon node in a taxonomy table, with each data record including a variable significance factor for weighting of terms on the basis at least of filling words and/or linking words and/or keywords, the computer program product can be used to record the language and/or text data on the basis of the taxonomy table, with frequency of individual data records in the language and/or text data determining a weighting parameter for the taxon nodes, the computer program product can be used to determine a determinable number of agglomerates in the taxonomy table on the basis of the weighting parameter of one or more taxon nodes, the computer program product can be used to generate a neural network, which can be used to classify and/or sort the language and/or text data on the basis of the agglomerates in the taxonomy table, the language, and/or text data.
44 . A computer program product which can be loaded in an internal memory of a digital computer and includes software code sections by which the operation as claimed in claim 33 can be carried out when the product is run on a computer, in which case the neural networks can be generated with software and/or hardware.Join the waitlist — get patent alerts
Track US2008215313A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.