US2008215355A1PendingUtilityA1

Method and System for Predicting Causes of Network Service Outages Using Time Domain Correlation

Assignee: HERRING DAVIDPriority: Nov 28, 2000Filed: Mar 24, 2008Published: Sep 4, 2008
Est. expiryNov 28, 2020(expired)· nominal 20-yr term from priority
H04L 43/106H04L 41/5003H04L 43/00H04L 41/064H04L 41/147G06Q 30/0283H04L 43/0817H04L 41/20H04L 43/10H04L 41/5025
40
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Methods and systems are described for predicting the likely causes of service outages using only time information, and for predicting and the likely costs of service outages. The likely causes are found by defining a narrow likely cause window around an outage based on service quality and/or service usage data, and correlating service events to the likely cause window in the time domain to find a probability distribution for the events. The likely costs are found by measuring usage loss and duration for a given point during an outage and using cost component functions of the time and usage to extrapolate over the outage. These cause and cost predictions supply service administrators with tools for making more informed decisions about allocation of resources in preventing and correcting service outages.

Claims

exact text as granted — not AI-modified
1 . A network monitoring system comprising:
 a service monitor for monitoring quality of service on the network;   a usage meter for measuring usage of the network;   an event detector for detecting a plurality of network events and corresponding times at which the network events occur; and   a probable cause engine, coupled to receive data from the service monitor, usage meter, and the event detector, the probable cause engine including a processing device that, in response to executable instructions, is operative to:   set a service change time window based upon data received from the service monitor or usage meter, the service change time window encompassing at least part of an occurrence of a service outage in the network;   determine which of the network events detected by the event detector is the most likely cause of a service change including computing a probability for each of the detected events that each of the detected events caused the service change based at least in part on a correlation between the event time and service change window;   determining whether one or more other events of a type identical to one of the detected events occurred; and   wherein computing the probability comprises computing the probability using at least in part a false occurrence weighting function which decreases the probability of the detected event as the case of the service change for instances in which the detected event occurred outside the service change time window.   
     
     
         2 . A computer readable medium storing program code for, when executed, causing a computer to perform a method for analyzing a potential cause of a change in a service, wherein service quality of the service is monitored, usage amount of the service is measured, and service events are detected, the method comprising:
 determining a service change time window based at least in part upon a change in service quality between a first working state and a second, non-working state, and upon a change in service usage amount, the service change time window encompassing at least part of a service outage;   retrieving data representing a plurality of detected events and corresponding times in which the event occurred;   computing a probability for each of the detected events that each of the detected events caused the service change based at least in part on a correlation between the event time and the service change time window;   determining whether one or more other events of a type identical to one of the detected events occurred; and   wherein computing the probability comprises computing the probability using at least in part a false occurrence weighting function which decreases the probability of the detected event as the cause of the service change for instances in which the detected event occurred outside the service change time window.   
     
     
         3 . Computer readable media comprising program code that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, the method comprising:
 determining a service change time window encompassing a change of service between a first working state and a service outage, the service change being determined at least in part based on measured service usage levels;   detecting occurrences of a set of events;   retrieving data representing the plurality of detected events and corresponding times in which the events occurred, wherein the set of events are within a given time prior to and during the service change time window, each occurrence of an event being associated with a time at which the event occurred;   computing a probability distribution for the set of events, which probability distribution determines for each event in the set the probability that the detected event caused the service change, the probability distribution being based at least in part on relations between the time of each event occurrence and the service change window;   wherein computing the probability includes using two or more second functions selected from the group consisting of:   a time weighting function which decreases the probability of a given event as the cause of the service change with the distance between the given event time and the service change time window;   a false occurrence weighting function which decreases the probability of a given event as the cause of the service change for instances in which events of the same type as the given event occurred outside the service change time window;   a positive occurrence weighting function which increases the probability of a given event as the cause of the service change based on instances stored in a historical database in which events of the same type as the given event occurred within a prior service change time window; and   a historical weighting function which increases the probability of a given event as the cause of the service change based on instances in the historical database in which events of the same type as the given event were identified as having caused a prior service outage.   
     
     
         4 . The computer readable media comprising program code of  claim 3  that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, wherein computing the probability distribution for the set of event comprises computing the probability distribution using a first weighting function which is the product of two or more second weighting functions. 
     
     
         5 . The computer readable media comprising program code of  claim 3  that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, the method comprising monitoring service quality, an wherein determining the service change time window comprises determining a service failure time window based upon a change in monitored service quality and narrowing the service failure time window to the service change time window based upon the service usage amount measured during the service failure time window. 
     
     
         6 . The computer readable media comprising program code of  claim 3  that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, the method comprising computing the probability distribution such that the total of all probabilities in the distribution is 1. 
     
     
         7 . The computer readable media comprising program code of  claim 3  that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, wherein the service comprises service over a communication network and wherein the detected events comprise network events. 
     
     
         8 . The computer readable media comprising program code of  claim 3  that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, wherein the service comprises service provided by an application program and wherein the detected events comprise application program events. 
     
     
         9 . The computer readable media comprising program code of  claim 3  that, when executed by a programmable microprocessor, causes the programmable microprocessor to execute a method for analyzing potential cause of a service change, wherein the service change is a service outage, comprising determining the service change time window as a change in service from the first working state to the second, non-working state.

Join the waitlist — get patent alerts

Track US2008215355A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.