Using a machine learning module to perform preemptive identification and reduction of risk of failure in computational systems
Abstract
Input on a plurality of attributes of a computing environment is provided to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within the computing environment. A determination is made as to whether the risk score exceeds a predetermined threshold. In response to determining that the risk score exceeds a predetermined threshold, an indication is transmitted to indicate that potential malfunctioning is likely to occur within the computing environment. A modification is made to the computing environment to prevent the potential malfunctioning from occurring.
Claims
exact text as granted — not AI-modified1 - 24 . (canceled)
25 . A method, comprising:
providing input on a plurality of attributes of a computing environment comprising one or more devices to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within a computing environment; and in response to determining that the risk score exceeds a predetermined threshold, transmitting an indication to indicate that the potential malfunctioning is likely to occur within the computing environment, wherein the indication additionally indicates a level of severity of the potential malfunctioning.
26 . The method of claim 25 , wherein the computing environment is modified to prevent the potential malfunctioning from occurring.
27 . The method of claim 25 , wherein the computing environment comprises one or more devices comprising one or more storage controllers, one or more storage drives, and one or more host computing systems, wherein the one or more storage controllers manage the storage drives to allow input/output (I/O) access to the one or more host computing systems.
28 . The method of claim 27 , wherein an attribute of the plurality of attributes is a measure of a firmware or software level of a device in comparison to a minimum or recommended firmware or software level for the device.
29 . The method of claim 27 , wherein the plurality of attributes are based on:
whether a device has reached an end of life cycle; and a ratio of faulty replaced drives to total number of drives over a period of time.
30 . The method of claim 27 , wherein an attribute is based on a level of redundancy in the computing environment indicated by Redundant Array of Independent Disks (RAID) configurations.
31 . The method of claim 27 , wherein the plurality of attributes indicate:
whether critical policy failures have occurred in the computing environment; whether one or more devices have missed heartbeats; an age of a device; and problems identified with a device.
32 . The method of claim 25 , wherein the machine learning module is configured to receive the input on the plurality of attributes of the computing environment, wherein the machine learning module is trained to adjust weights within the machine learning module, in response to events occurring in the computing environment, and wherein the events are associated with an expected output value of the machine learning module.
33 . A system, comprising:
a memory; and a processor coupled to the memory, wherein the processor performs operations, the operations comprising:
providing input on a plurality of attributes of a computing environment comprising one or more devices to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within a computing environment; and
in response to determining that the risk score exceeds a predetermined threshold, transmitting an indication to indicate that the potential malfunctioning is likely to occur within the computing environment, wherein the indication additionally indicates a level of severity of the potential malfunctioning.
34 . The system of claim 33 , wherein the computing environment is modified to prevent the potential malfunctioning from occurring.
35 . The system of claim 33 , wherein the computing environment comprises one or more devices comprising one or more storage controllers, one or more storage drives, and one or more host computing systems, wherein the one or more storage controllers manage the storage drives to allow input/output (I/O) access to the one or more host computing systems.
36 . The system of claim 35 , wherein an attribute of the plurality of attributes is a measure of a firmware or software level of a device in comparison to a minimum or recommended firmware or software level for the device.
37 . The system of claim 35 , wherein the plurality of attributes are based on:
whether a device has reached an end of life cycle; and a ratio of faulty replaced drives to total number of drives over a period of time.
38 . The system of claim 35 , wherein an attribute is based on a level of redundancy in the computing environment indicated by Redundant Array of Independent Disks (RAID) configurations.
39 . The system of claim 35 , wherein the plurality of attributes indicate:
whether critical policy failures have occurred in the computing environment; whether one or more devices have missed heartbeats; an age of a device; and problems identified with a device.
40 . The system of claim 33 , wherein the machine learning module is configured to receive the input on the plurality of attributes of the computing environment, wherein the machine learning module is trained to adjust weights within the machine learning module, in response to events occurring in the computing environment, and wherein the events are associated with an expected output value of the machine learning module.
41 . A computer program product, the computer program product comprising a computer readable storage medium having computer readable program code embodied therewith, the computer readable program code configured to perform operations in a computational device, the operations comprising:
providing input on a plurality of attributes of a computing environment comprising one or more devices to a machine learning module to produce an output value that comprises a risk score that indicates a likelihood of a potential malfunctioning occurring within a computing environment; and in response to determining that the risk score exceeds a predetermined threshold, transmitting an indication to indicate that the potential malfunctioning is likely to occur within the computing environment, wherein the indication additionally indicates a level of severity of the potential malfunctioning.
42 . The computer program product of claim 41 , wherein the computing environment is modified to prevent the potential malfunctioning from occurring.
43 . The computer program product of claim 41 , wherein the computing environment comprises one or more devices comprising one or more storage controllers, one or more storage drives, and one or more host computing systems, wherein the one or more storage controllers manage the storage drives to allow input/output (I/O) access to the one or more host computing systems.
44 . The computer program product of claim 43 , wherein an attribute of the plurality of attributes is a measure of a firmware or software level of a device in comparison to a minimum or recommended firmware or software level for the device.
45 . The computer program product of claim 43 , wherein the plurality of attributes are based on:
whether a device has reached an end of life cycle; and a ratio of faulty replaced drives to total number of drives over a period of time.
46 . The computer program product of claim 43 , wherein an attribute is based on a level of redundancy in the computing environment indicated by Redundant Array of Independent Disks (RAID) configurations.
47 . The computer program product of claim 43 , wherein the plurality of attributes indicate:
whether critical policy failures have occurred in the computing environment; whether one or more devices have missed heartbeats; an age of a device; and problems identified with a device.
48 . The computer program product of claim 41 , wherein the machine learning module is configured to receive the input on the plurality of attributes of the computing environment, wherein the machine learning module is trained to adjust weights within the machine learning module, in response to events occurring in the computing environment, and wherein the events are associated with an expected output value of the machine learning module.Join the waitlist — get patent alerts
Track US2022075676A1 — get alerts on status changes and closely related new filings.
We store only your email — no account needed. See our privacy policy.