US2015294016A1PendingUtilityA1
System, method and computer program for preparing data for analysis
Assignee: Patent Analytics Holding Pty LtdPriority: Jul 8, 2010Filed: Jun 23, 2015Published: Oct 15, 2015
Est. expiryJul 8, 2030(~4 yrs left)· nominal 20-yr term from priority
Inventors:Doris Spielthenner
G06F 17/30867G06F 17/3089G06F 17/3053G06F 7/24G06F 16/958G06F 16/9535G06F 16/382G06F 16/24578
36
PatentIndex Score
0
Cited by
0
References
0
Claims
Abstract
A method of preparing data for analysis, comprising the steps of receiving an initial data set including a plurality of records, each of the plurality of records including an identifier attribute and an associative attribute that identifies a further one or more records; receiving the further one or more records identified by the associative attribute in each of the plurality of records; and associating the further one or more records with the initial data set to form a final data set.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method of automated preparation of a set of data for analysis, the method comprising the steps of:
(a) retrieving from non-volatile memory or via a communication network a mapping input data set including a plurality of records, each of the plurality of records including an identifier attribute value and at least one associative attribute value, the identifier attribute value being an identifier for the record and each associative attribute value being an identifier of another record; (b) forming one or more networks from records of the mapping input data set by: (c) determining a plurality of direct connections, each one direct connection linking one record of the plurality of data records of the mapping input data set that share at least one common identifier attribute value, to form one or more networks of data records to another one record of the mapping input data set as a pair of data records, where the identifier attribute value of one record of the pair is the same as an associative attribute value of the other record of the pair; and (d) determining a set of records for each of the one or more networks based on the direct connections, where each record of the network forms a pair via a direct connection with at least one other record in the network; and (e) processing at least one of the one or more networks for further analysis or display.
2 . A method in accordance with claim 1 , further comprising preparing the mapping data set by performing the steps of:
(f) retrieving from non-volatile memory or via a communication network an initial input data set including one or more records, each of the records including an identifier attribute value and at least one associative attribute value, the identifier attribute value being an identifier for the record and each associative attribute value being an identifier of another record; (g) identifying one or more further records to retrieve by comparing each associative attribute value from each of the records in the input data set with the identifier attribute values of each of the records in the input data set, to determine any associative attribute values not matching any one of the identifier attributes values of records in the input data set, the associative attribute values not matching any one of the identifier attributes being identifier attributes of the one or more further records to retrieve; (h) retrieving from non-volatile memory or via a communication network the identified further records to retrieve; (i) associating the retrieved one or more further records with the initial data set to update the input data set; and (j) providing the updated initial data set as the mapping input data set.
3 . A method in accordance with claim 2 , further comprising repeating steps (g) to (i) until an end criteria is satisfied.
4 . A method in accordance with claim 2 , wherein initial data set is retrieved in accordance with a search query arranged to extract the initial data set from a database of data records.
5 . A method in accordance with claim 1 , wherein the plurality of data records includes any one or more of: information regarding patents, information regarding registered trade marks and information regarding scientific publications and for each record the identifier attribute value is any one of a serial number, an application number or a publication number, used to identify the document associated with the record.
6 . A method in accordance with claim 5 , wherein for each record each associative attribute value of the at least one associative attribute values is an identifier attribute value of a citation whereby the associative attributes for a record provide a list of citations for the record.
7 . A method in accordance with claim 1 , wherein step (e) comprises for at least one network the steps of:
(k) identifying at least one pair of data records in the network having a direct connection and, for each selected pair of data records; (l) identifying any one or more indirect connections linking the pair of data records via one or more intermediate records, each indirect connection consisting of x direct connections forming an x degree of separation link between the pair of data records, where x is a variable number between 2 and a maximum value n, and the value of x may be different for each indirect connection.
8 . A method in accordance with claim 7 , wherein step (e) further comprises the steps of:
(m) counting the number of indirect connections identified for each pair of records to derive a total count for the links between the data records of the pair; (o) setting a threshold level for the total count value between each pair of data records; (p) removing a direct link of any pair of data records having a total count value below the set threshold level; and (q) removing all data records that then no longer have a direct connection to any other record in the network from the reduced set for the network.
9 . A method in accordance with claim 7 , wherein step (e) further comprises the steps of:
(r) counting the number of instances of indirect links for each of degree of separation value x to derive a separation value for each of the degrees of separation for the pair of data records; (s) multiplying each respective separation value with a multiplier assigned for the degree of separation to derive a count value for each degree of separation; and (t) summing the count value for each degree of separation to derive a total count value.
10 . A method in accordance with claim 8 , wherein the maximum value of n equals 2.
11 . A method in accordance with claim 8 , wherein the threshold level for the total count value between each pair of data records is set based on a user selection.
12 . A method in accordance with claim 7 , further comprising the steps of ascribing a size value to each of the one or more networks, the size value being based on the number of data records in the reduced set for the network, and
selecting one or more networks for further analysis or display based on size value, wherein the selection is based on any one of:
a largest size value;
a set size value threshold; and
a size value ratio value as a function of a largest size value.
12 . A method in accordance with claim 1 , further comprising the step of
ranking the one or more networks wherein ranking is based on any one or more of:
a subject matter value determined for at least one of the one or more networks based on one or more subject matter attribute values for each of the records in the network; and
user input regarding one or more subject matter attribute values,
wherein subject matter attribute values include any one or more of:
number of times one or more key words appear;
number of forward citations;
number of family members; and
international patent classification (IPC) code.
13 . A method in accordance with claim 1 , wherein step (e) comprises displaying at least one of the one or more networks of data records utilising a visualisation methodology.
14 . A method in accordance with claim 13 , wherein at least one of the one or more networks is displayed as a map of interconnected nodes, each record being represented as a node and in response to a user selecting a node record data for the node is displayed.
15 . A method in accordance with claim 1 , wherein step (e) comprises constructing and outputting a related list of documents for one or more of the data networks.
16 . A method in accordance with claim 15 wherein the related list is a ranked list.
17 . A system for the automated preparation of a set of data for analysis comprising:
a processor; and a storage device;
the processor including:
a retrieving module configured to retrieve, from non-volatile memory or via a communication network, and store on the storage device a mapping input data set including a plurality of records, each of the plurality of records including an identifier attribute value and at least one associative attribute value, the identifier attribute value being an identifier for the record and each associative attribute value being an identifier of another record; and a linking module configured to form one or more networks from records of the mapping input data set by: determining a plurality of direct connections, each one direct connection linking one record of the plurality of data records of the mapping input data set to another one record of the mapping input data set as a pair of data records, where the identifier attribute value of one record of the pair is the same as an associative attribute value of the other record of the pair; determining a set of records for each of the one or more networks based on the direct connections, where each record of the network forms a pair via a direct connection with at least one other record in the network; and storing the at least one or more networks on the storage means for retrieval for further analysis or display.
18 . A system according to claim 17 , the processor further including an association module configured to, prepare the mapping input data set by, in response to one or more records being retrieved and stored as an input data set, identify one or more further records to retrieve by comparing each associative attribute value from each of the data records in the input data set with the identifier attribute value of each of the records in the input data set, to determine any associative attribute values not matching any one of the identifier attributes values of records in the input data set, the associative attribute values not matching any one of the identifier attributes being identifier attributes of the one or more further records to retrieve;
the retrieval module being further arranged to retrieve from non-volatile memory or via a communication network and store in the storage device the identified further records to retrieve; the association module being further arranged to associate the retrieved one or more further records with the initial data set to form the mapping input data set and store the mapping input data set in the storage device.
19 . A system according to claim 17 , the linking module being further configured to, in response to the one or more networks being formed, for at least one network:
identify at least one pair of data records in the network having a direct connection and, for each selected pair of data records; and identify any one or more indirect connections linking the pair of data records via one or more intermediate records, each indirect connection consisting of x direct connections forming an x degree of separation link between the pair of data records, where x is a variable number between 2 and a maximum value n, and the value of x may be different for each indirect connection.
20 . A system according to claim 19 , the processor further including a reduction module configured to, in response to the linking module identifying indirect connections for a network:
count the number of indirect connections identified for each pair of records to derive a total count for the links between the data records of the pair; set a threshold level for the total count value between each pair of data records; remove a direct link of any pair of data records having a total count value below the threshold level; and remove all data records that then no longer have a direct connection to any other record in the network, to reduce the number of data records in the set of records for the network.
21 . A system in accordance with claim 20 , wherein the maximum value of n equals 2.
22 . A system in accordance with claim 17 , the processor further including a ranking module configured to rank the one or more records of a network.
23 . A system in accordance with claim 17 , further comprising:
a display; and the processor further including a visualisation module configured to control display on the display of at least one of the one or more selected networks of data records utilising a visualisation methodology.Join the waitlist — get patent alerts
Track US2015294016A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.