Dynamically constrained, forward scheduling over uncertain workloads
Abstract
Scheduling searchable items such as web pages for crawling involves dynamically scheduling items for downloading based on capacity based on time. The workload is distributed over time, in advance, by anticipating and accounting for the discovery of new links on the particular host. Respective times to download items can be determined based on the current size of the host's crawl corpus relative to the maximum size of the host's crawl corpus. The respective times may be determined based additionally on respective freshness targets for the searchable items, which characterize how often an item's content should be refreshed by re-downloading the item, and on respective politeness factors for the host, which characterize the delay time between consecutive download requests to that host. As such, one can know precisely how the system is performing at any point in time and predict future performance.
Claims
exact text as granted — not AI-modified1 . A machine-implemented method comprising:
downloading a first searchable item from a particular host; determining a quality metric for the first searchable item; determining a freshness target for the first searchable item based at least in part on the quality metric for the first searchable item; and using the freshness target to determine when to re-download the first searchable item.
2 . The method of claim 1 , further comprising using a number of searchable items on the particular host that are scheduled for downloading to determine when to re-download the first searchable item.
3 . The method of claim 2 , further comprising using the number of searchable items on the particular host that are scheduled for downloading relative to a maximum number of searchable items on the particular host that are to be scheduled for downloading to determine when to re-download the first searchable item.
4 . The method of claim 1 , further comprising:
downloading a second searchable item on the particular host after downloading the first searchable item; determining a quality metric for the second searchable item; determining a second freshness target for the second searchable item based at least in part on the quality metric for the second searchable item; and using the second freshness target to determine when to re-download the second searchable item, wherein the second searchable item is re-downloaded before the first searchable item is re-downloaded.
5 . The method of claim 4 , wherein the quality metric for the second searchable item is different than the quality metric for the first searchable item.
6 . The method of claim 4 , further comprising using the number of searchable items on the particular host that are scheduled for downloading relative to a maximum number of searchable items on the particular host that are to be scheduled for downloading to determine when to re-download the second searchable item.
7 . The method of claim 1 , wherein said particular host is a first host, the method further comprising:
downloading a second searchable item on a second host after downloading the first searchable item, wherein the second host is a different host from the first host; determining a quality metric for the second searchable item; determining a second freshness target for the second searchable item based at least in part on the quality metric for the second searchable item; and using the second freshness target to determine when to re-download the second searchable item; wherein the second searchable item is re-downloaded before the first searchable item is re-downloaded.
8 . The method of claim 7 , wherein the first and second searchable items are re-downloaded using shared resources used to crawl a plurality of hosts.
9 . The method of claim 7 , further comprising using a number of searchable items on the second host that are scheduled for downloading to determine when to re-download the second searchable item.
10 . The method of claim 9 , further comprising using the number of searchable items on the second host that are scheduled for downloading relative to a maximum number of searchable items on the second host that are to be scheduled for downloading to determine when to re-download the second searchable item.
11 . The method of claim 1 , further comprising using processing capabilities of the particular host to determine when to re-download the first searchable item.
12 . The method of claim 1 , further comprising:
determining a politeness factor that specifies a delay between consecutive download requests submitted to the particular host; and using the politeness factor to determine when to re-download the first searchable item.
13 . The method of claim 1 , further comprising:
re-downloading the first searchable item; and using the freshness target to determine when to again re-download the first searchable item.
14 . The method of claim 13 , further comprising using the processing capabilities of the particular host to determine when to again re-download the first searchable item.
15 . The method of claim 13 , further comprising using a politeness factor that specifies a delay between consecutive download requests submitted the said particular host to determine when to again re-download the first searchable item.
16 . A machine-implemented method comprising:
downloading a plurality of searchable items hosted by a particular host, reading one or more freshness targets that correspond to the plurality of searchable items, wherein a freshness target specifies how often a corresponding searchable item should be crawled, wherein a freshness target is based at least in part on the quality metric for one or more of the plurality of searchable items; reading a politeness factor that corresponds to the particular host, wherein the politeness factor specifies a rate for submitting consecutive download requests to the particular host; determining, based at least in part on the freshness targets and the politeness factor, whether the particular crawler system has enough processing capacity to crawl the plurality of searchable items in compliance with the freshness targets and the politeness factor; and if determined the particular crawler system does not have enough processing capacity to crawl the plurality of searchable items in compliance with the freshness targets and the politeness factor, then generating a message indicating that the particular crawler system does not have enough processing capacity.
17 . The method of claim 16 , further comprising:
causing a graphical display of a number of searchable items queued for downloading per hour for a certain number of hours in the future.
18 . The method of claim 16 , further comprising:
causing display of a current operational status of the particular crawler system, wherein the operational status includes a number of searchable items queued for downloading for each of one or more hosts.
19 . The method of claim 18 , wherein the operational status includes an estimated time per download for the searchable items for each of the one or more hosts.Join the waitlist — get patent alerts
Track US2009077198A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.