Document relevancy analysis within machine learning systems
Abstract
Systems and methods that quantify document relevance for a document relative to a training corpus and select a best match or best matches are provided herein. Methods may include generating an example-based explanation for relevancy of a document to a training corpus by executing a support vector machine classifier, the support vector machine classifier performing a centroid classification of a relevant document in a term frequency-inverse document frequency features space relative to training examples in a training corpus, and generating an example-based explanation by selecting a best match for the relevant document from the training examples based upon the centroid classification.
Claims
exact text as granted — not AI-modifiedWhat is claimed is:
1 . A method for quantifying relevancy of a document to a training corpus, the method comprising:
classifying a first training corpus using a classifier to calculate an internal best match score for a relevant document relative to each training example in the first training corpus, the relevant document comprising a member of a document universe, the classification including:
calculating cosine distances between the relevant document and training examples in the first training corpus relative to term frequency-inverse document frequency weights associated with the training examples of the first training corpus; and
determining the training example in the first training corpus having a closest cosine distance to the relevant document; and
providing to an end user the training example from the first training corpus having the closest cosine distance to the relevant document.
2 . The method according to claim 1 , further comprising converting each of the training examples into a high-dimensional feature space using term frequencies.
3 . The method according to claim 1 , further comprising calculating an internal best match score for each of the training examples by multiplying a square root of term frequencies by an inverse document frequency.
4 . The method according to claim 1 , further comprising:
training the classifier on a training corpus that comprises training examples; classifying a set of documents; and determining relevant documents in the set.
5 . The method according to claim 4 , wherein the classifier comprises a support vector machine.
6 . The method according to claim 4 , wherein determining relevant documents in the set comprises determining distances between each document within the set of documents relative to a support vector machine model.
7 . The method according to claim 1 , further comprising outputting a list of the training examples based upon ranked cosine distances between the relevant document and the training examples.
8 . The method according to claim 7 , further comprising applying a relevancy threshold to affect an amount of training examples that are included in the list.
9 . The method according to claim 1 , wherein determining the training example having the closest cosine distance to the relevant document further comprises ranking the training examples by stretching the internal best match scores for the training examples linearly to cover a complete unit interval.
10 . A machine learning system that quantifies relevancy of a document to a training corpus, the system comprising:
at least one server comprising a processor configured to execute instructions that reside in memory, the instructions comprising:
a classifier module that:
calculates an internal best match score for a relevant document relative to each training example in the first training corpus, the relevant document being a member of a document universe, the classification including:
calculating cosine distances between the relevant document and training examples in the first training corpus relative to term frequency-inverse document frequency weights associated with the training examples of the first training corpus; and
determining the training example in the first training corpus having a closest cosine distance to the relevant document; and
a user interface module that provides to an end user the training example from the first training corpus having the closest cosine distance to the relevant document.
11 . The machine learning system according to claim 10 , wherein each of the training examples has been converted into a high-dimensional feature space using term frequencies.
12 . The machine learning system according to claim 10 , wherein the classifier module calculates an internal best match score for each of the training examples by multiplying a square root of term frequencies by an inverse document frequency.
13 . The machine learning system according to claim 10 , wherein the classifier module further:
classifies a set of documents; and determines relevant documents in the set, the classifier module being trained on a training corpus that comprises training examples.
14 . The machine learning system according to claim 13 , wherein the classifier module comprises a support vector machine.
15 . The machine learning system according to claim 13 , wherein the classifier module determines relevant documents in the set by determining cosine distances between each document within the set of documents using a support vector machine model.
16 . The machine learning system according to claim 10 , wherein the user interface module further outputs a list of the training examples based upon ranked cosine distances between the relevant document and the training examples.
17 . The machine learning system according to claim 16 , wherein the classifier module further applies a relevancy threshold to affect an amount of training examples that are included in the list.
18 . The machine learning system according to claim 10 , wherein the classifier module determines the training example having the closest cosine distance to the relevant document by ranking the training examples by stretching the internal best match scores for the training examples linearly to cover a complete unit interval.Join the waitlist — get patent alerts
Track US2019213197A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.