Harvesting data from page
Abstract
Among other disclosure, computer-implemented methods and computer program products for obtaining data from a page. A method can include initiating a harvesting process for a page available in a computer system. The method can include identifying a feed representation that has been created for the page. The method can include retrieving and storing, as part of the harvesting process, at least a portion from the page based on information in the identified feed representation. The feed representation can include at least excerpts of content from the page. The feed representation can include at least one representation selected from: an RSS feed, an Atom feed, an XML feed, an RDF feed, a serialized data feed representation, and combinations thereof.
Claims
exact text as granted — not AI-modified1 . A computer-implemented method of automatically mining data, comprising:
identifying a first portion of a page that is included in a feed representation associated with the page and a second portion of the page that is not included in the feed representation; retrieving and storing the second portion and not the first portion of the page.
2 . The computer-implemented method of claim 1 , wherein identifying the first portion includes comparing a part of the feed representation with the page.
3 . The computer-implemented method of claim 2 , wherein the comparing includes using a text recognition technique.
4 . The computer-implemented method of claim 1 , wherein the feed representation includes one or more of an RSS feed, an Atom feed, an XML feed, an RDF feed, and a serialized data feed representation.
5 . A computer-implemented method of automatically mining data, comprising:
identifying a page as a target for content retrieval; identifying a feed representation for the page, the identified feed representation including multiple feed entries matching portions of the page; identifying a first of the portions of the page that match the multiple feed entries and a second portion of the page that does not match any of the multiple feed entries; retrieving and storing the second portion and not the first portion of the page.
6 . The computer-implemented method of claim 5 , wherein the feed representation includes excerpts of content from the page.
7 . The computer-implemented method of claim 5 , wherein the feed representation comprises one or more of an RSS feed, an Atom feed, an XML feed, an RDF feed, and a serialized data feed representation.
8 . The computer-implemented method of claim 5 , wherein identifying the first portion includes performing text recognition.
9 . The computer-implemented method of claim 5 , further comprising:
identifying a second page linked to the first portion of the page; selecting a third portion of the second page that matches one of the feed entries; selecting a fourth portion of the second page that does not match any of the feed entries; retrieving and storing the fourth portion and not the third portion of the second page.
10 . The computer-implemented method of claim 5 , further comprising:
identifying a second page linked to the second portion of the page; selecting a third portion of the second page that matches one of the feed entries; selecting a fourth portion of the second page that does not match any of the feed entries; retrieving and storing the fourth portion and not the third portion of the second page.
11 . A system for automatically mining data, comprising:
a processor and memory, cooperating to function as: an identifying unit configured to identify a first portion of a page that is included in a feed representation associated with the page and a second portion of the page that is not included in the feed representation; a retrieving unit configured to retrieve and storing the second portion and not the first portion of the page.
12 . The system of claim 11 , wherein identifying the first portion includes comparing a part of the feed representation with the page.
13 . The system of claim 12 , wherein the comparing includes using a text recognition technique.
14 . The system of claim 11 , wherein the feed representation includes one or more of an RSS feed, an Atom feed, an XML feed, an RDF feed, and a serialized data feed representation.
15 . A machine-readable storage medium having stored thereon a set of instructions which when executed perform a method, the method compromising:
identifying a page as a target for content retrieval; identifying a feed representation for the page, the identified feed representation including multiple feed entries matching portions of the page; identifying a first of the portions of the page that match the multiple feed entries and a second portion of the page that does not match any of the multiple feed entries; retrieving and storing the second portion and not the first portion of the page.
16 . The machine-readable storage medium of claim 15 , wherein the feed representation includes excerpts of content from the page.
17 . The machine-readable storage medium of claim 15 , wherein the feed representation comprises one or more of an RSS feed, an Atom feed, an XML feed, an RDF feed, a serialized data feed representation.
18 . The machine-readable storage medium of claim 15 , wherein identifying the first portion includes performing text recognition.
19 . The machine-readable storage medium of claim 15 , the method further comprising:
identifying a second page linked to the first portion of the page; selecting a third portion of the second page that matches one of the feed entries; selecting a fourth portion of the second page that does not match any of the feed entries; retrieving and storing the fourth portion and not the third portion of the second page.
20 . The machine-readable storage medium of claim 5 , the method further comprising:
identifying a second page linked to the second portion of the page; selecting a third portion of the second page that matches one of the feed entries; selecting a fourth portion of the second page that does not match any of the feed entries; retrieving and storing the fourth portion and not the third portion of the second page.Join the waitlist — get patent alerts
Track US2015100870A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.