US2022075676A1PendingUtilityA1

Using a machine learning module to perform preemptive identification and reduction of risk of failure in computational systems

Assignee: IBMPriority: Oct 26, 2018Filed: Nov 15, 2021Published: Mar 10, 2022
Est. expiryOct 26, 2038(~12.2 yrs left)· nominal 20-yr term from priority
G06N 5/046G06F 11/142G06F 11/0781G06F 11/0793G06F 11/008G06F 11/004G06N 3/084G06N 20/00
66
PatentIndex Score
0
Cited by
0
References
0
Claims

Abstract

Input on a plurality of attributes of a computing environment is provided to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within the computing environment. A determination is made as to whether the risk score exceeds a predetermined threshold. In response to determining that the risk score exceeds a predetermined threshold, an indication is transmitted to indicate that potential malfunctioning is likely to occur within the computing environment. A modification is made to the computing environment to prevent the potential malfunctioning from occurring.

Claims

exact text as granted — not AI-modified
1 - 24 . (canceled) 
     
     
         25 . A method, comprising:
 providing input on a plurality of attributes of a computing environment comprising one or more devices to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within a computing environment; and   in response to determining that the risk score exceeds a predetermined threshold, transmitting an indication to indicate that the potential malfunctioning is likely to occur within the computing environment, wherein the indication additionally indicates a level of severity of the potential malfunctioning.   
     
     
         26 . The method of  claim 25 , wherein the computing environment is modified to prevent the potential malfunctioning from occurring. 
     
     
         27 . The method of  claim 25 , wherein the computing environment comprises one or more devices comprising one or more storage controllers, one or more storage drives, and one or more host computing systems, wherein the one or more storage controllers manage the storage drives to allow input/output (I/O) access to the one or more host computing systems. 
     
     
         28 . The method of  claim 27 , wherein an attribute of the plurality of attributes is a measure of a firmware or software level of a device in comparison to a minimum or recommended firmware or software level for the device. 
     
     
         29 . The method of  claim 27 , wherein the plurality of attributes are based on:
 whether a device has reached an end of life cycle; and   a ratio of faulty replaced drives to total number of drives over a period of time.   
     
     
         30 . The method of  claim 27 , wherein an attribute is based on a level of redundancy in the computing environment indicated by Redundant Array of Independent Disks (RAID) configurations. 
     
     
         31 . The method of  claim 27 , wherein the plurality of attributes indicate:
 whether critical policy failures have occurred in the computing environment;   whether one or more devices have missed heartbeats;   an age of a device; and   problems identified with a device.   
     
     
         32 . The method of  claim 25 , wherein the machine learning module is configured to receive the input on the plurality of attributes of the computing environment, wherein the machine learning module is trained to adjust weights within the machine learning module, in response to events occurring in the computing environment, and wherein the events are associated with an expected output value of the machine learning module. 
     
     
         33 . A system, comprising:
 a memory; and   a processor coupled to the memory, wherein the processor performs operations, the operations comprising:
 providing input on a plurality of attributes of a computing environment comprising one or more devices to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within a computing environment; and 
   in response to determining that the risk score exceeds a predetermined threshold, transmitting an indication to indicate that the potential malfunctioning is likely to occur within the computing environment, wherein the indication additionally indicates a level of severity of the potential malfunctioning.   
     
     
         34 . The system of  claim 33 , wherein the computing environment is modified to prevent the potential malfunctioning from occurring. 
     
     
         35 . The system of  claim 33 , wherein the computing environment comprises one or more devices comprising one or more storage controllers, one or more storage drives, and one or more host computing systems, wherein the one or more storage controllers manage the storage drives to allow input/output (I/O) access to the one or more host computing systems. 
     
     
         36 . The system of  claim 35 , wherein an attribute of the plurality of attributes is a measure of a firmware or software level of a device in comparison to a minimum or recommended firmware or software level for the device. 
     
     
         37 . The system of  claim 35 , wherein the plurality of attributes are based on:
 whether a device has reached an end of life cycle; and   a ratio of faulty replaced drives to total number of drives over a period of time.   
     
     
         38 . The system of  claim 35 , wherein an attribute is based on a level of redundancy in the computing environment indicated by Redundant Array of Independent Disks (RAID) configurations. 
     
     
         39 . The system of  claim 35 , wherein the plurality of attributes indicate:
 whether critical policy failures have occurred in the computing environment;   whether one or more devices have missed heartbeats;   an age of a device; and   problems identified with a device.   
     
     
         40 . The system of  claim 33 , wherein the machine learning module is configured to receive the input on the plurality of attributes of the computing environment, wherein the machine learning module is trained to adjust weights within the machine learning module, in response to events occurring in the computing environment, and wherein the events are associated with an expected output value of the machine learning module. 
     
     
         41 . A computer program product, the computer program product comprising a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code configured to perform operations in a computational device, the operations comprising:
 providing input on a plurality of attributes of a computing environment comprising one or more devices to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within a computing environment; and   in response to determining that the risk score exceeds a predetermined threshold, transmitting an indication to indicate that the potential malfunctioning is likely to occur within the computing environment, wherein the indication additionally indicates a level of severity of the potential malfunctioning.   
     
     
         42 . The computer program product of  claim 41 , wherein the computing environment is modified to prevent the potential malfunctioning from occurring. 
     
     
         43 . The computer program product of  claim 41 , wherein the computing environment comprises one or more devices comprising one or more storage controllers, one or more storage drives, and one or more host computing systems, wherein the one or more storage controllers manage the storage drives to allow input/output (I/O) access to the one or more host computing systems. 
     
     
         44 . The computer program product of  claim 43 , wherein an attribute of the plurality of attributes is a measure of a firmware or software level of a device in comparison to a minimum or recommended firmware or software level for the device. 
     
     
         45 . The computer program product of  claim 43 , wherein the plurality of attributes are based on:
 whether a device has reached an end of life cycle; and   a ratio of faulty replaced drives to total number of drives over a period of time.   
     
     
         46 . The computer program product of  claim 43 , wherein an attribute is based on a level of redundancy in the computing environment indicated by Redundant Array of Independent Disks (RAID) configurations. 
     
     
         47 . The computer program product of  claim 43 , wherein the plurality of attributes indicate:
 whether critical policy failures have occurred in the computing environment;   whether one or more devices have missed heartbeats;   an age of a device; and   problems identified with a device.   
     
     
         48 . The computer program product of  claim 41 , wherein the machine learning module is configured to receive the input on the plurality of attributes of the computing environment, wherein the machine learning module is trained to adjust weights within the machine learning module, in response to events occurring in the computing environment, and wherein the events are associated with an expected output value of the machine learning module.

Join the waitlist — get patent alerts

Track US2022075676A1 — get alerts on status changes and closely related new filings.

We store only your email — no account needed. See our privacy policy.