US2019384809A1PendingUtilityA1
Methods and systems for providing universal portability in machine learning
Est. expiryDec 9, 2034(~8.4 yrs left)· nominal 20-yr term from priority
Inventors:Schuyler D. ErleRobert J. MunroBrendan D. CallahanGary C. KingJason BrenierJames B. Robinson
G06F 40/30G06F 40/284G06F 40/169G06F 17/2785G06F 17/241G06F 17/277
59
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
Systems, methods, and apparatuses are presented for a trained language model to be stored in an efficient manner such that the trained language model may be utilized in virtually any computing device to conduct natural language processing. Unlike other natural language processing engines that may be computationally intensive to the point of being capable of running only on high performance machines, the organization of the natural language models according to the present disclosures allows for natural language processing to be performed even on smaller devices, such as mobile devices.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for classifying a document in natural language processing using a natural language model stored in one or more data files, the method comprising:
accessing one or more feature types from the data file, the one or more feature types each defining a data structure configured to access a tokenized sequence of the document and generate linguistic features from content within the tokenized sequence; performing a tokenizing operation of the document, the tokenizing operation configured to generate one or more tokenized sequences from the content within the document; generating a plurality of features for the document from the one or more tokenized sequences based on parameters defined by the one or more feature types; accessing a plurality of probabilities stored in the data file, each probability among the plurality of probabilities associated with a feature among the plurality of features and defining a pre-computed likelihood that said feature predicts a presence of absence of a label that the document is to be classified into; computing a prediction score indicating how likely the document is to be classified into said label, based on the plurality of probabilities; and classifying the document into said label based on comparing the prediction score to a threshold.
2 . The method of claim 1 , wherein generating the plurality of features is further based on parameters defined in task configuration data in the data file, the task configuration data associated with a type of task analysis that the natural language model is configured to classify the document into.
3 . The method of claim 1 , wherein the plurality of probabilities are pre-computed during a model training process configured to train the natural language model to classify documents according to at least said label and said task.
4 . The method of claim 1 , wherein the plurality of probabilities comprise a first probability that said feature predicts the presence of said label and a second probability that said feature predicts the absence of said label.
5 . The method of claim 1 , wherein the plurality of probabilities comprise a first probability that said feature appears at a beginning of a subset of the document, a second probability that said feature appears at an inside of the subset of the document, and a third probability that said feature appears on an outside of the subset of the document.
6 . The method of claim 1 , wherein each feature among the generated plurality of features comprises a first array storing integer indices of every label in the natural language model having a non-zero probability.
7 . The method of claim 6 , wherein each feature among the generated plurality of features further comprises a second array storing a quantized 16-bit fixed-point value of each probability converted as a logarithm.
8 . The method of claim 1 , wherein generating the plurality of features is based further on task configuration data stored in the data file, the task configuration data including training feature types used to train the natural language model.
9 . The method of claim 8 , wherein the task configuration data stored in the data file includes executable code configured perform a user-specified data transformation to generate custom feature types.
10 . The method of claim 1 , wherein the task configuration data further comprises analyst-defined tuning rules.
11 . A natural language platform configured to classify a document in natural language processing using a natural language model stored in one or more data files, the natural language platform comprising:
a memory configured to store the data file; and a processor coupled to the memory and configured to: access one or more feature types from the data file, the one or more feature types each defining a data structure configured to access a tokenized sequence of the document and generate linguistic features from content within the tokenized sequence; perform a tokenizing operation of the document, the tokenizing operation configured to generate one or more tokenized sequences from the content within the document; generate a plurality of features for the document from the one or more tokenized sequences based on parameters defined by the one or more feature types; access a plurality of probabilities stored in the data file, each probability among the plurality of probabilities associated with a feature among the plurality of features and defining a pre-computed likelihood that said feature predicts a presence of absence of a label that the document is to be classified into; compute a prediction score indicating how likely the document is to be classified into said label, based on the plurality of probabilities; and classify the document into said label based on comparing the prediction score to a threshold.
12 . The natural language platform of claim 11 , wherein generating the plurality of features is further based on parameters defined in task configuration data in the data file, the task configuration data associated with a type of task analysis that the natural language model is configured to classify the document into.
13 . The natural language platform of claim 11 , wherein the plurality of probabilities are pre-computed during a model training process configured to train the natural language model to classify documents according to at least said label and said task.
14 . The natural language platform of claim 11 , wherein the plurality of probabilities comprise a first probability that said feature predicts the presence of said label and a second probability that said feature predicts the absence of said label.
15 . The natural language platform of claim 11 , wherein the plurality of probabilities comprise a first probability that said feature appears at a beginning of a subset of the document, a second probability that said feature appears at an inside of the subset of the document, and a third probability that said feature appears on an outside of the subset of the document.
16 . The natural language platform of claim 11 , wherein each feature among the generated plurality of features comprises a first array storing integer indices of every label in the natural language model having a non-zero probability.
17 . The natural language platform of claim 16 , wherein each feature among the generated plurality of features further comprises a second array storing a quantized 16-bit fixed-point value of each probability converted as a logarithm.
18 . A non-transitory computer-readable medium embodying instructions that, when executed by a processor, perform operations comprising:
accessing one or more feature types from the data file, the one or more feature types each defining a data structure configured to access a tokenized sequence of the document and generate linguistic features from content within the tokenized sequence; performing a tokenizing operation of the document, the tokenizing operation configured to generate one or more tokenized sequences from the content within the document; generating a plurality of features for the document from the one or more tokenized sequences based on parameters defined by the one or more feature types; accessing a plurality of probabilities stored in the data file, each probability among the plurality of probabilities associated with a feature among the plurality of features and defining a pre-computed likelihood that said feature predicts a presence of absence of a label that the document is to be classified into; computing a prediction score indicating how likely the document is to be classified into said label, based on the plurality of probabilities; and classifying the document into said label based on comparing the prediction score to a threshold.
19 . The computer readable medium of claim 18 , wherein generating the plurality of features is further based on parameters defined in task configuration data in the data file, the task configuration data associated with a type of task analysis that the natural language model is configured to classify the document into.
20 . The computer readable medium of claim 18 , wherein the plurality of probabilities are pre-computed during a model training process configured to train the natural language model to classify documents according to at least said label and said task.Join the waitlist — get patent alerts
Track US2019384809A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.