CN116089155A - Fault processing method, computing device and computer storage medium - Google Patents

Fault processing method, computing device and computer storage medium Download PDF

Info

Publication number
CN116089155A
CN116089155A CN202310384818.7A CN202310384818A CN116089155A CN 116089155 A CN116089155 A CN 116089155A CN 202310384818 A CN202310384818 A CN 202310384818A CN 116089155 A CN116089155 A CN 116089155A
Authority
CN
China
Prior art keywords
fault
processing component
target processing
determining
target
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
CN202310384818.7A
Other languages
Chinese (zh)
Inventor
崔毕轩
王志强
毛文安
曾勇
冯富秋
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alibaba Cloud Computing Ltd
Original Assignee
Alibaba Cloud Computing Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alibaba Cloud Computing Ltd filed Critical Alibaba Cloud Computing Ltd
Priority to CN202310384818.7A priority Critical patent/CN116089155A/en
Publication of CN116089155A publication Critical patent/CN116089155A/en
Pending legal-status Critical Current

Links

Images

Classifications

    • YGENERAL TAGGING OF NEW TECHNOLOGICAL DEVELOPMENTS; GENERAL TAGGING OF CROSS-SECTIONAL TECHNOLOGIES SPANNING OVER SEVERAL SECTIONS OF THE IPC; TECHNICAL SUBJECTS COVERED BY FORMER USPC CROSS-REFERENCE ART COLLECTIONS [XRACs] AND DIGESTS
    • Y02TECHNOLOGIES OR APPLICATIONS FOR MITIGATION OR ADAPTATION AGAINST CLIMATE CHANGE
    • Y02DCLIMATE CHANGE MITIGATION TECHNOLOGIES IN INFORMATION AND COMMUNICATION TECHNOLOGIES [ICT], I.E. INFORMATION AND COMMUNICATION TECHNOLOGIES AIMING AT THE REDUCTION OF THEIR OWN ENERGY USE
    • Y02D10/00Energy efficient computing, e.g. low power processors, power management or thermal management

Landscapes

  • Hardware Redundancy (AREA)

Abstract

The embodiment of the application provides a fault processing method, computing equipment and a computer storage medium. The fault processing method comprises the following steps: acquiring fault notification information; determining, in response to the failure notification information, a failed target processing component from a plurality of processing components deployed in the computing device; judging whether the target processing component meets an isolation condition or not; and executing isolation processing on the target processing component under the condition that the target processing component meets the isolation condition. The technical scheme provided by the embodiment of the invention avoids service interruption and improves the service quality of the computing equipment.

Description

Fault processing method, computing device and computer storage medium
Technical Field
The embodiment of the invention relates to the technical field of computers, in particular to a fault processing method, computing equipment and a computer storage medium.
Background
Cloud computing is an infrastructure in the fields of the internet, the internet of things, artificial intelligence, big data and the like, and various cloud computing products and service layers are endless with the development of the fields. In the process of continuously changing and developing cloud computing platform software, the scale of the cloud computing platform is larger and larger, the number of computing devices for running the cloud computing platform is larger and larger, and the stability of the computing devices is also more and more important.
RAS (reliability, availability, and serviceability) capabilities of a computing device measure the recovery capability of the computing device in the event of a failure, determine the quality of service of the computing device, and also affect the continuity of processes on the cloud.
The inventors have discovered during the implementation of the inventive concept that hardware, and in particular, process component failures, are significant factors that lead to downtime of computing devices. In general, when a computing device detects that a processing component generates a fault, an operating system receives fault notification information, and because the processing component is an operation and control core of the computing device, in order to avoid the fault spreading inside the computing device, the operating system directly downtime the computing device, and all processing components deployed in the computing device are disconnected, so that the service quality of the computing device is reduced.
Disclosure of Invention
The embodiment of the invention provides a fault processing method, a fault processing device, computing equipment and a computer storage medium.
In a first aspect, an embodiment of the present invention provides a fault handling method, including:
acquiring fault notification information;
determining, in response to the failure notification information, a failed target processing component from a plurality of processing components deployed in the computing device;
judging whether the target processing component meets an isolation condition or not;
and executing isolation processing on the target processing component under the condition that the target processing component meets the isolation condition.
In a second aspect, an embodiment of the present invention provides a fault handling apparatus, including:
the notification acquisition module is used for acquiring fault notification information;
a target processing component determining module for determining a target processing component generating a fault from a plurality of processing components deployed in the computing device in response to the fault notification information;
the judging module is used for judging whether the target processing assembly meets the isolation condition or not;
and the isolation module is used for executing isolation processing on the target processing component under the condition that the target processing component meets the isolation condition.
The embodiment of the invention provides a fault processing method, which comprises the following steps: acquiring fault notification information; determining, in response to the failure notification information, a failed target processing component from a plurality of processing components deployed in the computing device; judging whether the target processing component meets an isolation condition or not; under the condition that the target processing assembly meets the isolation condition, the technical scheme of executing the isolation processing on the target processing assembly can locate the failed target processing assembly when at least one processing assembly of the computing equipment fails, and isolate the target processing assembly when the target processing assembly meets the isolation condition, so that the computing equipment can continue to operate by utilizing other processing assemblies which do not fail, service interruption is avoided, and service quality of the computing equipment is improved.
These and other aspects of the invention will be more readily apparent from the following description of the embodiments.
Drawings
In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the following description will briefly explain the drawings used in the embodiments or the description of the prior art, and it is obvious that the drawings in the following description are some embodiments of the present invention, and other drawings can be obtained according to these drawings without inventive effort for a person skilled in the art.
FIG. 1 schematically illustrates a flow chart of a fault handling method provided by one embodiment of the present invention;
FIG. 2 schematically illustrates a fault handling method provided by an embodiment of the present invention;
FIG. 3 schematically illustrates a schematic diagram of a fault handling method provided by an embodiment of the present invention;
FIG. 4 schematically illustrates a schematic diagram of an isolated target processing component provided by an embodiment of the present invention;
FIG. 5 schematically illustrates a block diagram of a fault handling apparatus provided by one embodiment of the present invention;
FIG. 6 schematically illustrates a block diagram of a computing device provided by one embodiment of the invention.
Detailed Description
In order to enable those skilled in the art to better understand the present invention, the following description will make clear and complete descriptions of the technical solutions according to the embodiments of the present invention with reference to the accompanying drawings.
In some of the flows described in the specification and claims of the present invention and in the foregoing figures, a plurality of operations occurring in a particular order are included, but it should be understood that the operations may be performed out of order or performed in parallel, with the order of operations such as 101, 102, etc., being merely used to distinguish between the various operations, the order of the operations themselves not representing any order of execution. In addition, the flows may include more or fewer operations, and the operations may be performed sequentially or in parallel. It should be noted that, the descriptions of "first" and "second" herein are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, and are not limited to the "first" and the "second" being different types.
It should be noted that, the user information (including but not limited to user equipment information, user personal information, etc.) and the data (including but not limited to data for analysis, stored data, presented data, etc.) related to the present invention are information and data authorized by the user or fully authorized by each party, and the collection, use and processing of the related data need to comply with the related laws and regulations and standards of the related country and region, and provide corresponding operation entries for the user to select authorization or rejection.
Cloud computing is an infrastructure in the fields of the internet, the internet of things, artificial intelligence, big data and the like, and various cloud computing products and service layers are endless with the development of the fields. In the process of continuously changing and developing cloud computing platform software, the scale of the cloud computing platform is larger and larger, the number of computing devices for running the cloud computing platform is larger and larger, and the stability of the computing devices is also more and more important.
RAS (reliability, availability, and serviceability) capabilities of a computing device measure the recovery capability of the computing device in the event of a failure, determine the quality of service of the computing device, and also affect the continuity of processes on the cloud.
The inventors have discovered during the implementation of the inventive concept that hardware, and in particular, process component failures, are significant factors that lead to downtime of computing devices. In general, when a computing device detects that a processing component generates a fault, an operating system receives fault notification information, and because the processing component is an operation and control core of the computing device, in order to avoid the fault spreading inside the computing device, the operating system directly downtime the computing device, and all processing components deployed in the computing device are disconnected, so that the service quality of the computing device is reduced.
In order to solve the technical problems in the related art, the embodiment of the invention provides a fault processing method, which comprises the following steps: acquiring fault notification information; determining, in response to the failure notification information, a target processing component that generated the failure from among a plurality of processing components deployed in the computing device; judging whether the target processing component meets the isolation condition or not; under the condition that the target processing components meet the isolation conditions, the technical scheme of executing the isolation processing on the target processing components can locate the failed target processing components when at least one processing component of the computing equipment fails, and isolate the target processing components when the target processing components meet the isolation conditions, so that the computing equipment can continue to operate by utilizing other processing components which do not fail, service interruption is avoided, and service quality of the computing equipment is improved.
The following description of the embodiments of the present invention will be made clearly and completely with reference to the accompanying drawings, in which it is apparent that the embodiments described are only some embodiments of the present invention, but not all embodiments. All other embodiments, which can be made by those skilled in the art based on the embodiments of the invention without making any inventive effort, are intended to fall within the scope of the invention.
Fig. 1 schematically illustrates a flowchart of a fault handling method according to an embodiment of the present invention, where, as shown in fig. 1, the fault handling method may include the following steps:
101, acquiring fault notification information;
102, determining a target processing component generating a fault from a plurality of processing components deployed in the computing device in response to the fault notification information;
103, judging whether the target processing component meets the isolation condition;
104, in the case that the target processing component satisfies the isolation condition, performing the isolation processing on the target processing component.
According to embodiments of the present invention, a plurality of processing components are typically deployed in a computing device, and fault notification information may be generated in the event that at least one of the plurality of processing components fails.
In a preferred embodiment of the present invention, the failure notification information may be generated in the case where there is a hardware failure of at least one of the plurality of processing components, which may be a hardware failure determined to be unrepairable.
According to an embodiment of the invention, the processing component may be an operation and control component of a computing device, such as a CPU, an MCU (micro control unit, microcontroller Unit), a GPU (graphics processor, graphics Processing Unit), etc.
According to an embodiment of the present invention, for example, in the case where the processing component is a CPU, the CPU may be internally composed of an arithmetic logic unit, a controller, a register, and an interrupt system. The fault notification information may be generated in case of a fault of an internal component of the CPU.
In general, when a hardware failure of an internal component of a CPU is detected, the hardware failure needs to be reported to an operating system, and the operating system triggers a downtime of a computing device, because the hardware failure cannot be recovered. When the computing device is down, all CPUs in the computing device are taken off line, resulting in service interruption.
In an embodiment of the present invention, when at least one processing component in the computing device has a hardware fault, a specific target processing component that generates the fault may be first located from a plurality of processing components, and then it is determined whether the fault target processing component satisfies the isolation condition.
If the target processing component meets the isolation condition, the target processing component can be isolated. By isolating the target processing component with faults, the faults of the target processing component can be prevented from being spread to other processing components of the computing equipment, so that the computing equipment can continue to operate by utilizing the other processing components, the integral downtime of the computing equipment caused by the faults of at least one processing component in the computing equipment is avoided, the service interruption is avoided, and the service quality of the computing equipment is improved.
According to the embodiment of the invention, judging whether the target processing component meets the isolation condition can be realized specifically as follows:
determining the fault type and the fault position of the target processing component;
and determining whether the target processing component meets the isolation condition according to the fault type and the fault position.
According to an embodiment of the present invention, the isolation conditions may include, for example, the failure type of the target processing component being a specified failure type, and the failure location being at a specified location.
Fig. 2 schematically illustrates a schematic diagram of a fault handling method provided by an embodiment of the present invention.
A schematic diagram of registers of a computing device is schematically shown in fig. 2. The computing device provides a variety of registers for detecting and handling hardware faults. As shown in fig. 2, the system includes a global register and a plurality of fault report registers, where each fault report register is configured to record and report fault information of a corresponding hardware unit in the CPU. For example, the fault report register 1 is used to record and report fault information of the arithmetic logic unit.
In practical applications, the fault report register can only record the primary fault generated by the corresponding hardware unit, and the recording rule is that the fault with high priority covers the fault with low priority, and the fault with the same priority only records the first fault. The priority may be determined based on the severity of the fault, e.g., the more severe the fault is.
For example, when an arithmetic logic unit generates a low priority fault, such as a repairable fault, the fault report register 1 records the fault and reports the fault to an operating system, and the operating system triggers a repair instruction to repair the repairable fault. When the arithmetic logic unit generates two times of low-priority faults at the same time, because the fault report register records the first fault only, only one fault is recorded and reported, and the other fault is not detected, if the operation system is still reported and the operation of triggering the instruction to recover the fault is still executed, only the recorded fault is recovered, and the fault which is not recorded is continuously executed by the CPU, so that the fault is diffused into other CPUs of the computing equipment.
Thus, when a fault occurs in the CPU internal hardware unit twice or more simultaneously, the fault will be upgraded to an overflow fault, which can be determined as an unrepairable fault. When the operating system detects that the CPU generates overflow faults, only one fault can be detected and repaired in two or more faults, so that the computing equipment is required to be triggered to be down, and service interruption is caused.
According to the embodiment of the invention, according to the fault type and the fault position of the computing equipment, determining whether the target processing component meets the isolation condition can be specifically realized as follows:
determining whether the fault is an overflow fault and whether the fault is a repairable fault according to the fault type;
determining whether a fault occurs inside the target processing component according to the fault position;
when the fault is determined to be an overflow fault and the fault is a repairable fault and the fault occurs inside the target processing component, the target processing component is determined to need to be isolated.
According to an embodiment of the present invention, after determining a failed target processing component, the following determination operation may be performed to determine whether the target processing component can be isolated:
firstly, judging whether the fault of a target processing component is an overflow fault, namely, judging whether the target processing component has the same hardware unit and simultaneously generates two or more faults;
secondly, judging whether the fault recorded by the fault report register is a repairable fault or not; because the recording rule of the fault report register is that the high-priority fault covers the low-priority fault, if the fault recorded by the fault report register is a repairable fault, the repairable fault is the fault with the highest priority, namely, the fault is not the unrepairable fault, and the fault can be repaired after isolating the target processing component;
then, it is determined whether the fault occurred inside the target processing component.
Under the condition that the fault of the target processing component meets the three judging conditions at the same time, the target processing component can be determined to meet the isolating condition, and the target processing component can be isolated.
According to an embodiment of the present invention, the fault handling method further includes:
and executing downtime processing on the device under the condition that the fault is determined not to be an overflow fault and is an irreparable fault.
According to the embodiment of the invention, in the case that the fault of the target processing component is determined to be the irreparable fault, even if the target processing component is isolated, the fault cannot be recovered, the computing device can be directly subjected to downtime so as to overcome the irreparable fault as soon as possible by, for example, replacing the target processing component.
According to an embodiment of the present invention, the fault handling method further includes:
and executing repair processing on the fault in the case that the fault is determined not to be an overflow fault and the fault is a repairable fault.
According to the embodiment of the invention, under the condition that the fault is not an overflow fault and is a repairable fault, the fault can be repaired on line, the computer equipment is not required to be down, and the target processing assembly is not required to be isolated.
According to an embodiment of the present invention, determining the fault type and the fault location of the target processing component may be specifically implemented as:
reading a register of the target processing component;
and determining the fault type and the fault position of the target processing component according to the value of the register.
According to an embodiment of the present invention, determining the fault type and the fault location of the target processing component according to the value of the register may be specifically implemented as:
determining a first bit and a second bit of a register;
and determining the fault type and the fault position according to the values of the first bit and the second bit respectively.
According to the embodiment of the invention, the fault type and the fault position of the target processing component can be determined by reading the value of the fault report register corresponding to the hardware unit in the target processing component.
According to an embodiment of the present invention, different bits of the fault report register may be pre-configured to record different fault information, e.g. a first bit may be used to record the fault type and a second bit may be used to record the fault location.
According to an embodiment of the invention, the second bit may be, for example, a value of 0 or 1,0 may be used to characterize that the fault is occurring outside the target processing component and 1 may be used to characterize that the fault is occurring inside the target processing component.
According to an embodiment of the present invention, the first bit may comprise, for example, a first value and a second value, the first value may be used to indicate whether the recorded fault is an overflow fault, for example, the first value may be 0 or 1,0 may indicate that the recorded fault is not an overflow fault, and 1 may indicate that the recorded fault is an overflow fault. The second value may be used to characterize whether the recorded fault is a repairable fault, e.g., the second value may be 0 or 1,0 may characterize that the recorded fault is not a repairable fault, and 1 may characterize that the recorded fault is a repairable fault.
According to an embodiment of the present invention, determining a failed target processing component from a plurality of processing components deployed in a computing device may be specifically implemented as:
sequentially retrieving values of fault reporting registers of a plurality of processing components;
and determining the target processing component from the plurality of processing components according to the value of the fault reporting register.
According to the embodiment of the invention, the fault information of the hardware units of the processing components recorded in the fault report registers of the processing components is used for determining the target processing components in a mode of sequentially searching the values of the fault report registers of the processing components.
According to the embodiment of the invention, the values of the fault report registers of the processing components can be sequentially searched as follows:
determining a first processing component from a plurality of processing components;
sending an inspection instruction to the first processing component such that the first processing component performs the following operations in response to the inspection instruction:
reading a global register of the first processing component, and determining whether a fault occurs in the first processing component according to the value of the global register;
if yes, outputting the identification information of the first processing component and the value of the fault reporting register;
if not, scanning global registers of other processing components except the first processing component in turn until the target processing component is determined, and outputting identification information of the target processing component and the value of a fault reporting register of the target processing component.
Fig. 3 schematically illustrates a schematic diagram of a fault handling method provided by an embodiment of the present invention.
According to embodiments of the present invention, when at least one processing component in a computing device generates an unrepairable hardware fault, the operating system may generate a check instruction notification to all processing components to cause all processing components to suspend task processing for fault checking.
According to an embodiment of the present invention, the first processing component may be the processing component that first responds to the inspection instruction, or the processing component that is randomly determined from a plurality of processing components.
According to an embodiment of the present invention, after the first processing component receives the inspection instruction, first, self-inspection may be performed. Specifically, firstly, the check bit of the fault report register can be read, whether the fault report register is effective is determined through the value of the check bit, if the fault report register is effective, the fault report register can be continuously read, whether the fault triggered by the processing component is judged to be an irreparable fault or not is judged according to the value of the fault report register, if the fault triggered by the processing component is not judged to be an irreparable fault according to the value of the fault report register, the fault can be repaired online, and then a fault repairing flow can be executed; if the fault triggered by the processing component is judged to be the irreparable fault according to the value of the fault reporting register, the global register can be read, whether the fault occurs locally to the first processing component or not is judged according to the value of the global register, and if the fault occurs locally to the first processing component, the identification information of the first processing component and the value of the fault reporting register are recorded, so that fault information is generated.
If the fault does not occur locally in the first processing component, the first processing component may sequentially scan global registers of other processing components except the first processing component until a target processing component generating a hardware clap is determined according to the value of the global register, and then record identification information of the target processing component and the value of a fault reporting register of the target processing component.
According to an embodiment of the present invention, performing isolation processing on a target processing component may be specifically implemented as:
the target processing component is off-line from the computing device such that the target processing component cannot continue to receive processing requests.
According to an embodiment of the present invention, after performing the isolation processing on the target processing component, the fault processing method further includes:
acquiring delay time preset by a user;
and when the isolation time is detected to reach the delay time, downtime processing is carried out on the computing equipment.
FIG. 4 schematically illustrates a schematic diagram of an isolated target processing component provided by an embodiment of the present invention.
According to the embodiment of the invention, the target processing component is isolated, so that the target processing component is disconnected from the computing device, and the fault is prevented from being spread among a plurality of processing components of the computing device due to the fact that the fault is prevented from continuously running.
According to an embodiment of the present invention, isolating the target processing component may be performed by an isolation module, which may be a functional module of a software isolated processing component, including steps of deleting the target processing component from a processing component topology of the computing device at a software level, releasing a resource held by the target processing component, stopping interaction with the failed target processing component, and configuring a subsequent policy.
Specifically, the isolation module may first determine whether the target processing component is a main processing component of the computing device, where the main processing component is a core processing component of the computing device, and cannot perform isolation, and if the target processing component is the core processing component, the isolation module may print fault information and perform downtime processing on the computing device.
Second, the quarantine module may shut down the software timeout detection dog and context resources of the target processing component and set the online status of the target processing component to off, and then shut down the Inter-processor interrupt (Inter-Processor Interrupt) of the target processing component. Meanwhile, the target processing component can be deleted from the processing component topological graph of the computing device, and at the moment, the target processing component can be determined to be successfully isolated at the operating system level.
According to the embodiment of the invention, after the target processing component is successfully isolated, a corresponding processing mode can be adopted according to the setting of a user.
Specifically, for example, for the isolation module, after the isolation is successful, fault information may be printed, where the fault information may include, for example, identification information of the target processing component, information of a fault report register, and the like; for the delay downtime mode, delay time configured by a user through a configuration interface in advance can be acquired, after the target processing component is successfully isolated, the isolation time is recorded, and when the isolation time reaches the delay time set by the user, downtime processing is carried out on the computing equipment; for the immediate downtime mode, the computing equipment can be directly downtime, and fault information is printed.
According to the embodiment of the invention, after the target processing component is successfully isolated, the fault information of the target processing component can be reported by using the fault reporting module, and the fault information can comprise the fault position, the fault type and the isolation success information obtained by reading the fault reporting register.
Fig. 5 schematically illustrates a block diagram of a fault handling apparatus according to an embodiment of the present invention, where, as shown in fig. 5, the fault handling apparatus may include:
a notification acquisition module 501, configured to acquire fault notification information;
a target processing component determination module 502 for determining a failed target processing component from a plurality of processing components deployed in the computing device in response to the failure notification information;
a judging module 503, configured to judge whether the target processing component meets an isolation condition;
an isolation module 504, configured to perform isolation processing on the target processing component if the target processing component meets an isolation condition.
According to an embodiment of the present invention, the judging module 503 includes:
the first determining submodule is used for determining the fault type and the fault position of the target processing component;
and the second determining submodule is used for determining whether the target processing assembly meets the isolation condition according to the fault type and the fault position.
According to an embodiment of the invention, the second determination submodule comprises:
the first determining unit is used for determining whether the fault is an overflow fault or not and whether the fault is a repairable fault or not according to the fault type;
the second determining unit is used for determining whether the fault occurs in the target processing assembly according to the fault position;
and the third determining unit is used for determining that the target processing assembly needs to be subjected to isolation processing when the fault is determined to be an overflow fault and the fault is a repairable fault and the fault occurs in the target processing assembly.
According to an embodiment of the present invention, the fault handling apparatus further includes:
and the first determining module is used for executing downtime processing on the computing equipment under the condition that the fault is determined not to be an overflow fault and is an irreparable fault.
According to an embodiment of the present invention, the fault handling apparatus further includes:
and the second determining module is used for executing repairing processing on the computing equipment fault under the condition that the fault is not an overflow fault and the fault is a repairable fault.
According to an embodiment of the invention, the first determination submodule comprises:
a reading unit for reading a register of the target processing component;
and the fourth determining unit is used for determining the fault type and the fault position of the target processing component according to the value of the register.
According to an embodiment of the present invention, the fourth determination unit includes:
a first determination subunit configured to determine a first bit and a second bit of the register;
and the second determining subunit is used for determining the fault type and the fault position according to the values of the first bit and the second bit respectively.
According to an embodiment of the invention, the target processing component determination module 502 includes:
the retrieval sub-module is used for sequentially retrieving values of fault reporting registers of the plurality of processing components;
and the target processing component determining submodule is used for determining the target processing component from the plurality of processing components according to the value of the fault reporting register.
According to an embodiment of the invention, the retrieval submodule comprises:
a first component determination unit configured to determine a first processing component from a plurality of processing components;
an instruction transmitting unit for transmitting a check instruction to the first processing component so that the first processing component performs the following operations in response to the check instruction:
the register reading unit is used for reading a global register of the first processing component and determining whether a fault occurs in the first processing component according to the value of the global register;
the first output unit is used for outputting the identification information of the first processing component and the value of the fault report register under the condition that the fault occurs in the first processing component;
and the scanning unit is used for scanning global registers of other processing components except the first processing component in sequence until the target processing component is determined under the condition that the fault does not occur in the first processing component, and outputting the identification information of the target processing component and the value of the fault reporting register of the target processing component.
According to an embodiment of the invention, the isolation module 204 includes:
and the isolation unit is used for disconnecting the target processing component from the computing device so that the target processing component cannot continuously receive the processing request.
According to an embodiment of the present invention, the fault handling apparatus further includes:
the time acquisition module is used for acquiring delay time preset by a user;
and the downtime module is used for downtime processing the computing equipment when the isolation time is detected to reach the delay time.
The fault handling apparatus shown in fig. 5 may perform the fault handling method described in the embodiment shown in fig. 1, and its implementation principle and technical effects are not repeated. The specific manner in which the individual modules, units, and operations of the apparatus of the above embodiment are performed has been described in detail in connection with the embodiment of the method, and will not be described in detail here.
In one possible design, the fault handling apparatus provided by the embodiments of the present invention may be implemented as a computing device, as shown in fig. 6, which may include a storage component 601, a plurality of processing components 602, and a management component 603;
the storage component 601 stores one or more computer instructions, where the one or more computer instructions are used for the management and control component 603 to call and execute, so as to implement the fault handling method provided by the embodiment of the present invention.
Of course, the computing device may necessarily include other components, such as input/output interfaces, communication components, and the like. The input/output interface provides an interface between the processing component and a peripheral interface module, which may be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the computing device and other devices, and the like.
The computing device may be a physical device or an elastic computing host provided by the cloud computing platform, and at this time, the computing device may be a cloud server, and the processing component, the storage component, and the like may be a base server resource rented or purchased from the cloud computing platform.
When the computing device is a physical device, the computing device may be implemented as a distributed cluster formed by a plurality of servers or terminal devices, or may be implemented as a single server or a single terminal device.
The embodiment of the invention also provides a computer readable storage medium which stores a computer program, and the computer program can realize the fault processing method provided by the embodiment of the invention when being executed by a computer.
The embodiment of the invention also provides a computer program product, which comprises a computer program, wherein the computer program can realize the fault processing method provided by the embodiment of the invention when being executed by a computer.
Wherein the processing components of the respective embodiments above may include one or more processors to execute computer instructions to perform all or part of the steps of the methods described above. Of course, the processing component may also be implemented as one or more Application Specific Integrated Circuits (ASICs), digital Signal Processors (DSPs), digital Signal Processing Devices (DSPDs), programmable Logic Devices (PLDs), field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic elements for executing the methods described above.
The storage component is configured to store various types of data to support operation in the device. The memory component may be implemented by any type or combination of volatile or nonvolatile memory devices such as Static Random Access Memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic or optical disk.
It will be clear to those skilled in the art that, for convenience and brevity of description, specific working procedures of the above-described systems, apparatuses and units may refer to corresponding procedures in the foregoing method embodiments, which are not repeated herein.
The apparatus embodiments described above are merely illustrative, wherein the elements illustrated as separate elements may or may not be physically separate, and the elements shown as elements may or may not be physical elements, may be located in one place, or may be distributed over a plurality of network elements. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art will understand and implement the present invention without undue burden.
From the above description of the embodiments, it will be apparent to those skilled in the art that the embodiments may be implemented by means of software plus necessary general hardware platforms, or of course may be implemented by means of hardware. Based on this understanding, the foregoing technical solution may be embodied essentially or in a part contributing to the prior art in the form of a software product, which may be stored in a computer readable storage medium, such as ROM/RAM, a magnetic disk, an optical disk, etc., including several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the method described in the respective embodiments or some parts of the embodiments.
Finally, it should be noted that: the above embodiments are only for illustrating the technical solution of the present invention, and are not limiting; although the invention has been described in detail with reference to the foregoing embodiments, it will be understood by those of ordinary skill in the art that: the technical scheme described in the foregoing embodiments can be modified or some technical features thereof can be replaced by equivalents; such modifications and substitutions do not depart from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims (13)

1. A method of fault handling comprising:
acquiring fault notification information;
determining, in response to the failure notification information, a failed target processing component from a plurality of processing components deployed in the computing device;
judging whether the target processing component meets an isolation condition or not;
and executing isolation processing on the target processing component under the condition that the target processing component meets the isolation condition.
2. The method of claim 1, wherein said determining whether the target processing component satisfies an isolation condition comprises:
determining a fault type and a fault position of the target processing component;
and determining whether the target processing component meets an isolation condition according to the fault type and the fault position.
3. The method of claim 2, wherein determining whether the target processing component satisfies an isolation condition based on the fault type and the fault location comprises:
determining whether the fault is an overflow fault or not and whether the fault is a repairable fault or not according to the fault type;
determining whether the fault occurs inside the target processing component according to the fault position;
when the fault is determined to be an overflow fault and the fault is a repairable fault and the fault occurs inside a target processing component, determining that isolation processing is required for the target processing component.
4. A method according to claim 3, characterized in that the method further comprises:
and executing downtime processing on the computing equipment under the condition that the fault is determined not to be an overflow fault and is an irreparable fault.
5. The method according to claim 4, wherein the method further comprises:
and executing repair processing on the fault when the fault is determined not to be an overflow fault and the fault is a repairable fault.
6. The method of claim 2, wherein the determining the fault type and fault location of the target processing component comprises:
reading a register of the target processing component;
and determining the fault type and the fault position of the target processing component according to the value of the register.
7. The method of claim 6, wherein determining the fault type and fault location of the target processing component based on the register value comprises:
determining a first bit and a second bit of the register;
and determining the fault type and the fault position according to the values of the first bit and the second bit respectively.
8. The method of claim 1, wherein determining a failed target processing component from a plurality of processing components deployed in a computing device comprises:
sequentially retrieving values of fault reporting registers of the plurality of processing components;
and determining the target processing component from a plurality of processing components according to the value of the fault reporting register.
9. The method of claim 8, wherein sequentially retrieving values of fault reporting registers of the plurality of processing components comprises:
determining a first processing component from the plurality of processing components;
sending a check instruction to the first processing component so that the first processing component performs the following operations in response to the check instruction:
reading a global register of a first processing component, and determining whether a fault occurs in the first processing component according to the value of the global register;
if yes, outputting the identification information of the first processing component and the value of the fault reporting register;
if not, scanning global registers of other processing components except the first processing component in turn until a target processing component is determined, and outputting identification information of the target processing component and the value of a fault reporting register of the target processing component.
10. The method of claim 1, wherein performing isolation processing on the target processing component comprises:
the target processing component is off-line from the computing device such that the target processing component cannot continue to receive processing requests.
11. The method of claim 1, wherein after performing the quarantine treatment on the target processing component, the method further comprises:
acquiring delay time preset by a user;
and when the isolation time is detected to reach the delay time, downtime processing is carried out on the computing equipment.
12. A computing device comprising a management and control component, a plurality of processing components, and a storage component;
the storage component stores one or more computer instructions; the one or more computer instructions are operable to be invoked by a management component to implement a fault handling method as claimed in any one of claims 1 to 11.
13. A computer storage medium, characterized in that a computer program is stored, which, when executed by a computer, implements the fault handling method according to any of claims 1 to 11.
CN202310384818.7A 2023-04-11 2023-04-11 Fault processing method, computing device and computer storage medium Pending CN116089155A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
CN202310384818.7A CN116089155A (en) 2023-04-11 2023-04-11 Fault processing method, computing device and computer storage medium

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
CN202310384818.7A CN116089155A (en) 2023-04-11 2023-04-11 Fault processing method, computing device and computer storage medium

Publications (1)

Publication Number Publication Date
CN116089155A true CN116089155A (en) 2023-05-09

Family

ID=86208708

Family Applications (1)

Application Number Title Priority Date Filing Date
CN202310384818.7A Pending CN116089155A (en) 2023-04-11 2023-04-11 Fault processing method, computing device and computer storage medium

Country Status (1)

Country Link
CN (1) CN116089155A (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120723515A (en) * 2025-06-24 2025-09-30 中兴通讯股份有限公司 Processor fault handling method, electronic device, storage medium and program product

Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190129788A1 (en) * 2017-10-31 2019-05-02 Paypal, Inc. Automated, adaptive, and auto-remediating system for production environment
CN109783262A (en) * 2018-12-24 2019-05-21 新华三技术有限公司 Fault data processing method, device, server and computer readable storage medium
CN109947586A (en) * 2019-03-20 2019-06-28 浪潮商用机器有限公司 A method, apparatus and medium for isolating faulty equipment
CN110691234A (en) * 2019-09-09 2020-01-14 北京达佳互联信息技术有限公司 Fault processing method, device, server and storage medium
CN113918375A (en) * 2021-12-13 2022-01-11 苏州浪潮智能科技有限公司 Fault processing method and device, electronic equipment and storage medium
CN115328684A (en) * 2022-06-30 2022-11-11 超聚变数字技术有限公司 Memory fault reporting method, BMC and electronic equipment
CN115766405A (en) * 2023-01-09 2023-03-07 苏州浪潮智能科技有限公司 A fault handling method, device, equipment and storage medium

Patent Citations (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20190129788A1 (en) * 2017-10-31 2019-05-02 Paypal, Inc. Automated, adaptive, and auto-remediating system for production environment
CN109783262A (en) * 2018-12-24 2019-05-21 新华三技术有限公司 Fault data processing method, device, server and computer readable storage medium
CN109947586A (en) * 2019-03-20 2019-06-28 浪潮商用机器有限公司 A method, apparatus and medium for isolating faulty equipment
CN110691234A (en) * 2019-09-09 2020-01-14 北京达佳互联信息技术有限公司 Fault processing method, device, server and storage medium
CN113918375A (en) * 2021-12-13 2022-01-11 苏州浪潮智能科技有限公司 Fault processing method and device, electronic equipment and storage medium
CN115328684A (en) * 2022-06-30 2022-11-11 超聚变数字技术有限公司 Memory fault reporting method, BMC and electronic equipment
CN115766405A (en) * 2023-01-09 2023-03-07 苏州浪潮智能科技有限公司 A fault handling method, device, equipment and storage medium

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN120723515A (en) * 2025-06-24 2025-09-30 中兴通讯股份有限公司 Processor fault handling method, electronic device, storage medium and program product

Similar Documents

Publication Publication Date Title
US7328376B2 (en) Error reporting to diagnostic engines based on their diagnostic capabilities
US8225142B2 (en) Method and system for tracepoint-based fault diagnosis and recovery
US9804917B2 (en) Notification of address range including non-correctable error
JP2017517060A (en) Fault processing method, related apparatus, and computer
CN111881014A (en) System test method, device, storage medium and electronic equipment
CN112905480B (en) Test script generation method, device, storage medium and electronic device
CN116841585A (en) Firmware upgrading method and device
CN115658373A (en) Server-based memory processing method and device, processor and electronic equipment
CN116089155A (en) Fault processing method, computing device and computer storage medium
CN112068935A (en) Method, device and equipment for monitoring deployment of kubernets program
CN117472657A (en) Error repair methods, devices, equipment and storage media
CN116643906A (en) Cloud platform fault processing method and device, electronic equipment and storage medium
CN119814529B (en) Fault alarm method, device, computer equipment and storage medium
CN115640236B (en) Script quality detection method and computing device
CN115658470B (en) Automatic testing method and device for failure recovery mechanism of distributed system
CN108415788B (en) Data processing apparatus and method for responding to non-responsive processing circuitry
CN116401118A (en) Method and device for monitoring Samba of file sharing service
CN115484267A (en) Multi-cluster deployment processing method and device, electronic equipment and storage medium
CN111813786B (en) Defect detection/processing method and device
CN118550758B (en) Troubleshooting methods and computing devices
CN116991710B (en) Automatic testing method and system, electronic device, and storage medium
CN120929319B (en) Memory stress testing methods and electronic devices
CN119728481B (en) PaaS application anomaly detection method, device and equipment
CN120256187B (en) Device failure processing method, electronic device, storage medium, and program product
CN120723508A (en) Fault detection method, device and equipment, readable storage medium, and program product

Legal Events

Date Code Title Description
PB01 Publication
PB01 Publication
SE01 Entry into force of request for substantive examination
SE01 Entry into force of request for substantive examination
RJ01 Rejection of invention patent application after publication

Application publication date: 20230509

RJ01 Rejection of invention patent application after publication