Methods and systems for language-agnostic machine learning in natural language processing using feature extraction
Abstract
Methods, apparatuses, and systems are presented for generating natural language models using a novel system architecture for feature extraction. A method for extracting features for natural language processing comprises: accessing one or more tokens generated from a document to be processed; receiving one or more feature types defined by user; receiving selection of one or more feature types from a plurality of system-defined and user-defined feature types, wherein each feature type comprises one or more rules for generating features; receiving one or more parameters for the selected feature types, wherein the one or more rules for generating features are defined at least in part by the parameters; generating features associated with the document to be processed based on the selected feature types and the received parameters; and outputting the generated features in a format common among all feature types.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for extracting features for natural language processing, the method comprising:
accessing, by one or more processors in a natural language processing platform, one or more tokens generated from a document to be processed; receiving, by the one or more processors, one or more feature types defined by a user; receiving, by the one or more processors, a selection of one or more feature types from a plurality of system-defined and user-defined feature types, wherein each feature type comprises one or more rules for generating features; receiving, by the one or more processors, one or more parameters of the selected feature types, wherein the one or more rules for generating features are defined at least in part by the parameters; generating, by the one or more processors, features associated with the document to be processed based on the selected feature types and the received parameters; and outputting, by the one or more processors, the generated features in a format common among all feature types.
2 . The method of claim 1 , wherein the plurality of feature types comprises one or more feature types that generate features each comprising at least one combination of the accessed tokens.
3 . The method of claim 1 , further comprising accessing one or more tags attached to the one or more tokens, wherein the plurality of feature types comprises one or more feature types that generate features containing information in the tags.
4 . The method of claim 1 , further comprising accessing metadata associated with the document to be processed, wherein the plurality of feature types comprises one or more feature types that generate features containing information in the metadata.
5 . The method of claim 4 , wherein the plurality of feature types comprises a feature type that generates features each comprising information in the metadata and a combination of tokens.
6 . The method of claim 1 , further comprising generating statistics across a pool of documents, wherein the plurality of feature types comprises one or more feature types that generate features based on the statistics.
7 . The method of claim 6 , wherein the statistics comprise an average or median document length in the pool, and the plurality of feature types comprises a feature type that generates a feature indicating whether the document to be processed is longer than, is shorter than, or equals to the average or median document length.
8 . The method of claim 1 , further comprising accessing a list of entries, wherein the plurality of feature types comprises a feature type that generates a feature indicating whether the document to be processed contains one or more tokens that match one or more of the list of entries.
9 . The method of claim 1 , further comprising accessing a list of word vectors, wherein the plurality of feature types comprises a feature type that generates features containing one or more word vectors each corresponding to a combination of the accessed tokens.
10 . The method of claim 1 , further comprising:
calculating frequencies of occurrence of one or more generated features within a pool of documents; and storing the frequencies of occurrence in a format accessible by a module for submitting documents for human annotation.
11 . The method of claim 1 , further comprising presenting, in a user interface, one or more features associated with a document.
12 . The method of claim 1 , wherein:
the document to be processed is in one or more languages; the one or more tokens are accessed in a language agnostic format; and the generated features are outputted in a language agnostic format.
13 . An apparatus for extracting features for natural language processing, the apparatus comprising one or more processors configured to:
access one or more tokens generated from a document to be processed; receive one or more feature types defined by user; receive selection of one or more feature types from a plurality of system-defined and user-defined feature types, wherein each feature type comprises one or more rules for generating features; receive one or more parameters for the selected feature types, wherein the one or more rules for generating features are defined at least in part by the parameters; generate features associated with the document to be processed based on the selected feature types and the received parameters; and output the generated features in a format common among all feature types.
14 . The apparatus of claim 13 , wherein the plurality of feature types comprises one or more feature types that generate features each comprising at least one combination of the accessed tokens.
15 . The apparatus of claim 13 , wherein the one or more processors are further configured to access one or more tags attached to the one or more tokens, and the plurality of feature types comprises one or more feature types that generate features containing information in the tags.
16 . The apparatus of claim 13 , wherein the one or more processors are further configured to access metadata associated with the document to be processed, and the plurality of feature types comprises one or more feature types that generate features containing information in the metadata.
17 . The apparatus of claim 13 , wherein the one or more processors are further configured to generate statistics across a pool of documents, and the plurality of feature types comprises one or more feature types that generate features based on the statistics.
18 . The apparatus of claim 13 , wherein
the document to be processed is in one or more languages; the one or more tokens are accessed in a language agnostic format; and the generated features are outputted in a language agnostic format.
19 . A non-transitory computer readable medium comprising instructions that, when executed by a processor, cause the processor to:
access one or more tokens generated from a document to be processed; receive one or more feature types defined by user; receive selection of one or more feature types from a plurality of system-defined and user-defined feature types, wherein each feature type comprises one or more rules for generating features; receive one or more parameters for the selected feature types, wherein the one or more rules for generating features are defined at least in part by the parameters; generate features associated with the document to be processed based on the selected feature types and the received parameters; and output the generated features in a format common among all feature types.
20 . The non-transitory computer readable medium of claim 19 , wherein the plurality of feature types comprises one or more feature types that generate features each comprising at least one combination of the accessed tokens.Join the waitlist — get patent alerts
Track US2018157636A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.