JP2004355424A - Failure management method for information processing equipment - Google Patents

Failure management method for information processing equipment Download PDF

Info

Publication number
JP2004355424A
JP2004355424A JP2003153705A JP2003153705A JP2004355424A JP 2004355424 A JP2004355424 A JP 2004355424A JP 2003153705 A JP2003153705 A JP 2003153705A JP 2003153705 A JP2003153705 A JP 2003153705A JP 2004355424 A JP2004355424 A JP 2004355424A
Authority
JP
Japan
Prior art keywords
failure
fault
information processing
maintenance
component
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP2003153705A
Other languages
Japanese (ja)
Inventor
Daiki Abe
大輝 阿部
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hitachi Ltd
Original Assignee
Hitachi Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hitachi Ltd filed Critical Hitachi Ltd
Priority to JP2003153705A priority Critical patent/JP2004355424A/en
Publication of JP2004355424A publication Critical patent/JP2004355424A/en
Pending legal-status Critical Current

Links

Images

Landscapes

  • Test And Diagnosis Of Digital Computers (AREA)
  • Computer And Data Communications (AREA)
  • Management, Administration, Business Operations System, And Electronic Commerce (AREA)

Abstract

【課題】情報処理装置の障害管理方式において、保守作業に連動して自動的に障害コード辞書を更新することにより、人為的な障害コード辞書の更新作業を行うことなく、被疑部品の的中率を向上させる。
【解決手段】情報処理装置の部品の障害回復をシステムファームウェアの動作テストで検出し、サービスプロセッサ上で動作する交換部品特定プログラムが部品のシリアル番号の変化から保守のために交換した部品を特定し、交換した部品の情報をインターネット経由で保守センタの保守管理サーバへ送信し、前記情報を受信した保守管理サーバ上で動作する障害コード辞書更新プログラムがハードディスクに格納されている障害コード辞書ファイルを更新する。
【選択図】 図1
In a fault management system for an information processing apparatus, a fault code dictionary is automatically updated in conjunction with a maintenance work, so that a hit rate of a suspected part can be accurately calculated without performing an artificial fault code dictionary update work. Improve.
An information processing apparatus detects a failure recovery of a component by an operation test of a system firmware, and a replacement component specifying program operating on a service processor specifies a component replaced for maintenance based on a change in a serial number of the component. The information of the replaced parts is transmitted to the maintenance management server of the maintenance center via the Internet, and the failure code dictionary update program operating on the maintenance management server receiving the information updates the failure code dictionary file stored on the hard disk. I do.
[Selection diagram] Fig. 1

Description

【0001】
【発明の属する技術分野】
本発明は、情報処理装置の障害管理方式に関し、特に障害コードを基に被疑部品を指摘する方式に関する。
【0002】
【従来の技術】
従来の情報処理装置の障害管理方式においては、障害コードと被疑部品の対応を示す障害コード辞書を予め作成し、障害発生時には障害コードをキーとして障害コード辞書を検索し、被疑部品の指摘を行っている。
【0003】
ただし、障害コードと被疑部品の関係は必ずしも1対1ではなく、1つの障害コードに対して複数の部品が被疑の対象となる場合がある。このため、被疑部品と被疑順位がセットで障害コード辞書に設定され、障害発生時には被疑順位に従った順番で被疑部品の交換を行っている。
【0004】
ただし、被疑順位の設定は情報処理装置の設計段階で行われるため、被疑順位には実際の部品の故障率および品質が反映されておらず、情報処理装置の出荷後に保守作業の回数を積まない限り、被疑順位の設定誤りが顕在化しないという問題があった。このため、情報処理装置の出荷後も被疑順位の設定が実状と相違ないか監視し、必要に応じて障害コード辞書の更新を行っていた。
【0005】
【特許文献1】
特開平10−320241号公報
【0006】
【発明が解決しようとする課題】
前記方式では、前述した通り、被疑部品の的中率を向上させるため、情報処理装置の出荷後も障害コード辞書の見直しを継続的に行わなければならないという問題があった。
【0007】
本発明は、前記問題点を鑑みてなされたものであり、保守作業に連動して自動的に障害コード辞書を更新することにより、人為的な障害コード辞書の更新作業を行うことなく、被疑部品の的中率を向上させることを目的とする。
【0008】
【課題を解決するための手段】
本発明の情報処理装置の障害管理方式について、図1を参照して説明する。
【0009】
本発明の障害管理方式は、情報処理装置(101)の部品(103)の障害発生を障害発生検出手段(105)が検出し、サービスプロセッサ(107)上で動作する障害診断プログラム(108)が障害コードを公衆回線(110)経由で保守センタ(111)の保守管理サーバ(114)へ送信し、前記障害コードを受信した保守管理サーバ(114)上で動作する被疑部品表示プログラム(115)がハードディスク(112)に格納されている障害コード辞書ファイル(113)を基に推定した被疑部品を出力手段(117)に表示する情報処理装置の障害管理方式において、情報処理装置(101)の部品(103)の障害回復を障害回復検出手段(106)が検出し、サービスプロセッサ(107)上で動作する交換部品特定プログラム(109)が保守のために交換した部品の種別情報を公衆回線(110)経由で保守センタ(111)の保守管理サーバ(114)へ送信し、前記情報を受信した保守管理サーバ(114)上で動作する障害コード辞書更新プログラム(116)がハードディスク(112)に格納されている障害コード辞書ファイル(113)を更新することを特徴とする。
【0010】
また、サービスプロセッサ(107)上で動作する交換部品特定プログラム(109)が交換された部品(103)を特定する手段として、情報処理装置(101)の部品(103)に一意に付加されているシリアル番号(104)をシリアル番号採取インタフェース(102)経由で参照し、障害発生前と障害回復後のシリアル番号(104)を比較することを特徴とする。
【0011】
さらに、障害回復検出手段(106)として、情報処理装置のブート時にシステムファームウェアが行う初期部品動作テストを用いることを特徴とする。
【0012】
なお、保守センタを設けず、情報処理装置の部品の障害発生を障害発生検出手段が検出し、サービスプロセッサ上で動作する障害診断プログラムが作成した障害コードを基にサービスプロセッサ上で動作する被疑部品表示プログラムが推定した被疑部品を出力手段に表示する情報処理装置の障害管理方式において、情報処理装置の部品の障害回復を障害回復検出手段が検出し、サービスプロセッサ上で動作する交換部品特定プログラムが保守のために交換した部品の種別情報を基にサービスプロセッサ上で動作する障害コード辞書更新プログラムがハードディスクに格納されている障害コード辞書ファイルを更新するという構成も、前記課題を解決するための手段として取り得る。
【0013】
【発明の実施の形態】
以下、本発明の実施形態について、図面を参照して説明する。図2は、本発明の一実施形態の構成を示すブロック図である。
【0014】
保守対象サーバ(201)は、保守センタ(223)と保守サービスの契約を結んでいる情報処理装置である。保守対象サーバ(201)の交換可能な部品は、CPU(205)、DIMM(207)、システムボード(202)の3つであり、各部品には、一意のシリアル番号が付加されている。なお、CPU(205)とDIMM(207)は、システムボード(202)上に搭載されているが、自由に着脱することが可能である。
【0015】
CPU(205)のシリアル番号は、PIROM(206)に記憶されている。PIROM(206)は、CPU(205)に内蔵されているSEEPROMであり、CPU(205)に関する情報を記憶している。
【0016】
DIMM(207)のシリアル番号は、SPD(208)に記憶されている。SPD(208)は、DIMM(207)に内蔵されているSEEPROMであり、DIMM(207)に関する情報を記憶している。
【0017】
システムボード(202)のシリアル番号は、FRU−ROM(204)に記憶されている。FRU−ROM(204)は、システムボード(202)上に搭載されているSEEPROMであり、システムボード(202)に関する情報を記憶している。
【0018】
システムボード(202)上には、CPU(205)、DIMM(207)、FRU−ROM(204)の他に、チップセット(211)とBIOS−ROM(209)が搭載されている。
【0019】
チップセット(211)は、CPU(205)からDIMM(207)、BIOS−ROM(209)、NVRAM(214)へのアクセスを制御し、CPU(205)、DIMM(207)、チップセット(211)の内部処理および通信処理において発生した障害を検出する機能を備えている。前記機能は、障害を検出した場合、チップセット(211)に内蔵されている障害状態レジスタ(212)に障害情報を記憶する。
【0020】
BIOS−ROM(209)は、保守対象サーバ(201)のシステムファームウェアのコードが記憶されているEEPROMである。前記システムファームウェアは、CPU(205)上で動作し、保守対象サーバ(201)のブート処理を行う。なお、前記ブート処理には、CPU(205)、DIMM(207)、システムボード(202)が正常に動作するかテストする初期部品動作テスト機能が含まれている。障害回復確認プログラム(210)は、前記初期部品動作テスト機能を利用して障害が回復したか確認するプログラムである。
【0021】
図8は、障害回復確認プログラム(210)の処理を示す流れ図である。障害回復確認プログラム(210)は、前記初期部品動作テスト機能を用いてCPU(205)、DIMM(207)、システムボード(202)が正常に動作するかテストを行い(801)、前記テストの結果より障害が回復したか確認し(802)、障害の回復を確認できた場合、障害回復フラグを“1”にセットし(803)、ブート処理を続行する(804)。一方、障害の回復を確認できなかった場合、ブート処理を中断する(805)。
【0022】
保守対象サーバ(201)内には、システムボード(202)の他に、サービスプロセッサボード(213)が装着されている。サービスプロセッサボード(213)上には、NVRAM(214)とマイコン(219)が搭載されている。
【0023】
NVRAM(214)は、CPU(205)とマイコン(219)の両方からアクセスが可能な不揮発性メモリであり、障害発生フラグ(215)、障害回復フラグ(216)、RC待避変数(217)、シリアル番号表(218)が配置されている。
【0024】
障害発生フラグ(214)は、障害の発生状態を示す2値変数である。障害発生フラグ(214)=“0”は、障害が発生していないことを意味し、障害発生フラグ(214)=“1”は、障害が発生したことを意味する。
【0025】
障害回復フラグ(215)は、障害の回復状態を示す2値変数である。障害回復フラグ(215)=“0”は、障害が回復していないことを意味し、障害回復フラグ(215)=“1”は、障害が回復したことを意味する。
【0026】
RC退避変数(216)は、障害診断プログラム(220)が作成したRCを記憶しておくために使用する変数である。なお、RCとは“ReferenceCode”の略語であり、障害コードに相当する用語である。
【0027】
シリアル番号表(218)は、CPU(205)、DIMM(207)、システムボード(202)のシリアル番号を記憶するために使用する配列変数である。図3は、シリアル番号表(218)の構成を示す表である。配列の添数は、部品の種別に対応し、“1”=CPU(205)、“2”=DIMM(207)、“3”=システムボード(202)と定義する。また、配列の要素には、添数に対応する部品のシリアル番号が記憶される。
【0028】
マイコン(219)は、CPU(205)と独立して動作する組み込みコントローラであり、LAN通信機能、IICバスアクセス機能、内蔵ROMを備えている。
【0029】
前記LANアクセス機能は、インターネット(222)経由の通信を行うための機能である。マイコン(219)は、インターネット(222)経由で、保守管理サーバ(227)との通信を行うことができる。
【0030】
前記IICバスアクセス機能は、IICバス(203)経由のアクセスを行うための機能である。マイコン(219)は、IICバス(203)経由で、FRU−ROM(204)、PIROM(206)、SPD(208)、障害状態レジスタ(212)をアクセスすることができる。
【0031】
前記内蔵ROMには、障害診断プログラム(220)と交換部品特定プログラム(221)が格納されている。
【0032】
障害診断プログラム(220)は、障害が発生したことを検出し、RCを保守管理サーバ(227)に通知するプログラムである。図7は、障害診断プログラム(220)の処理を示す流れ図である。障害診断プログラム(220)は、IICバス経由で障害状態レジスタ(212)を参照し(701)、障害状態レジスタ(212)に障害情報が記憶されているか判定し(702)、障害状態レジスタ(212)に障害情報が記憶されていた場合、前記障害情報を基にRCを作成し(703)、前記RCをRC待避変数(217)に記憶し(704)、前記RCをインターネット(222)経由で保守管理サーバ(227)に送信し(705)、障害回復フラグ(216)の値を“0”にクリアし(706)、障害発生フラグ(215)の値を“1”にセットする(707)。
【0033】
交換部品特定プログラム(221)は、障害が回復したことを検出し、保守のために交換した部品の種別情報を保守管理サーバ(227)に通知するプログラムである。図9は、交換部品特定プログラム(221)の処理を示す流れ図である。交換部品特定プログラム(221)は、障害発生フラグ(215)と障害回復フラグ(216)の値が両方共“1”であるか判定し(901)、障害発生フラグ(215)と障害回復フラグ(216)の値が両方共“1”の場合、IICバス(203)経由でPIROM(206)、SPD(208)、FRU−ROM(204)に記憶されている現在のCPU(205)、DIMM(207)、システムボード(202)のシリアル番号を参照し(902、906、910)、シリアル番号表(218)に記憶されている障害発生時のCPU(205)、DIMM(207)、システムボード(202)のシリアル番号と差違があるか判定し(903、907、911)、現在のシリアル番号と障害発生時のシリアル番号に差違があった場合、シリアル番号に差違のあった部品の種別情報をRC待避変数に記憶しておいたRCと共にインターネット(222)経由で保守管理サーバ(227)に送信し(904、908、912)、シリアル番号表(218)を現状に沿うように更新する(905、909、913)。最後に、障害発生フラグ(215)の値を“0”にクリアする(914)。
【0034】
保守センタ(223)は、保守対象サーバ(201)の保守作業を行う保守員の在籍する建物であり、ディスプレイ(224)、ハードディスク(225)、保守管理サーバ(227)が設置されている。
【0035】
ディスプレイ(224)は、保守管理サーバ(227)の出力を表示する表示装置である。
【0036】
ハードディスク(225)は、保守管理サーバ(227)のファイルを保存する記憶装置であり、RC辞書ファイル(226)を保存している。
【0037】
RC辞書ファイル(226)は、RCと被疑部品の対応を示すファイルである。図4は、RC辞書ファイル(226)の構成を示した表である。RC辞書ファイル(226)の項目は、RC、部品、交換回数、優先順位の4つである。なお、キー項目は、RCと部品である。交換回数は、該RCに対して該部品を交換して障害が回復した回数を示している。優先順位は、該RCに対して交換回数の等しい部品が複数存在した場合の被疑の優先順位を示し、保守対象サーバ(201)の設計者によって予め設定されている。なお、優先順位は、“1”が最も高く、数値が増加するほど低くなる。また、優先順位が“0”の場合は、該部品が被疑対象外であることを示している。
【0038】
保守管理サーバ(227)は、保守対象サーバ(201)を管理するサーバであり、被疑部品表示プログラム(228)とRC辞書更新プログラム(229)が格納されている。
【0039】
被疑部品表示プログラム(228)は、マイコン(219)からインターネット(222)経由でRCを受信し、前記RCをキーにRC辞書ファイル(226)を検索し、ディスプレイ(224)に被疑部品を表示するプログラムである。なお、被疑部品が複数存在する場合は、夫々の被疑部品に被疑順位を付けて表示する。図5は、被疑部品表示プログラム(228)によるディスプレイ(224)の表示を示す図であり、図4のRC辞書ファイル(226)と対応している。図5の(A)は、RC=“AAAAAAAA”の場合の表示である。図5より、DIMM(207)の交換回数は“8”であり、システムボード(202)の交換回数である“6”より多い。ここで、交換回数に差がある場合、優先順位は使用しない。よって、図5の(A)に示す通り、被疑順位1位はDIMM(207)となり、被疑順位2位はシステムボード(202)となる。なお、CPU(205)は、優先順位が“0”のため、被疑対象外である。図7の(B)は、RC=“BBBBBBBB”の場合の表示である。図4より、CPU(205)の交換回数は“4”であり、システムボード(202)の交換回数である“4”と等しい。交換回数が等しい場合は、優先順位を用いて被疑順位を付ける。図4より、CPU(205)の優先順位は“1”であり、システムボード(202)の優先順位である“2”より高い。よって、図5の(B)に示す通り、被疑順位1位はCPU(205)となり、被疑順位2位はシステムボード(202)となる。なお、DIMM(207)は、優先順位が“0”のため、被疑対象外である。
【0040】
RC辞書更新プログラム(229)は、マイコン(219)からインターネット(222)経由で交換した部品の種別情報とRCを受信し、該RCと該部品をキーにRC辞書ファイル(226)を検索し、対応するレコードの交換回数を“+1”するプログラムである。
【0041】
次に、図6の流れ図を参照して障害の発生からRC辞書ファイル(226)の更新までの動作について説明する。
【0042】
図10は、障害発生前の障害発生フラグ(214)、障害回復フラグ(215)、RC退避変数(216)、シリアル番号表(217)、実際の部品のシリアル番号、RC辞書ファイル(226)、ディスプレイ(229)の表示を示す図である。図10の状態において、システムボード(202)に起因する障害が発生したと仮定する(601)。チップセット(211)は、前記障害を検出し、障害状態レジスタ(212)に障害情報を記憶する(602)。障害状態レジスタ(212)に前記障害情報が記憶されたことを受け、マイコン(219)上で動作する障害診断プログラム(220)が処理を開始する(603)。
【0043】
ここで、説明を図7の障害診断プログラム(220)の流れ図に移す。障害診断プログラム(220)は、IICバス(203)経由で障害状態レジスタ(212)を参照し(701)、障害状態レジスタ(212)に前記障害情報が記憶されていることを検出し(702)、前記障害情報を基にRCを作成し(703)、前記RCをRC待避変数(217)に記憶し(704)、前記RCをインターネット(222)経由で保守管理サーバ(227)に送信し(705)、障害回復フラグ(216)の値を“0”にクリアし(706)、障害発生フラグ(215)の値を“1”にセットする(707)。なお、前記ステップ703において作成したRCは、“AAAAAAAA”であったと仮定する。
【0044】
説明を図6の流れ図に戻す。保守管理サーバ(227)は、前記RCをマイコン(219)からインターネット(222)経由で受信し、被疑部品表示プログラム(228)を起動する(604)。被疑部品表示プログラム(228)は、前記RCをキーにRC辞書ファイル(226)を検索し、ディスプレイ(224)に被疑部品を表示する。図10のRC辞書ファイル(226)より、前記RC=“AAAAAAAA”に対応する部品の交換回数は、DIMM(207)とシステムボード(202)が共に“2”である。交換回数が等しい場合は、優先順位を用いて被疑順位を付ける。優先順位は、DIMM(207)が“1”であり、システムボード(202)の“2”よりも高い。よって、被疑順位1位はDIMM(207)となり、被疑順位2位はシステムボード(202)となる。図11は、現時点の障害発生フラグ(215)、障害回復フラグ(216)、RC退避変数(217)、シリアル番号表(218)、実際の部品のシリアル番号、RC辞書ファイル(226)、ディスプレイ(224)の表示を示す図である。なお、図11の網掛け箇所は、図10との差分を示している。
【0045】
次に、保守センタ(223)に在籍する保守員は、ディスプレイ(224)の前記表示を見て、最も被疑順位が高いDIMM(207)を交換し(605)、保守対象サーバ(201)を再起動する(606)。保守対象サーバ(201)を再起動することにより、情報処理装置(201)のブート処理が始まり、CPU(205)上で動作する障害回復確認プログラム(210)が処理を開始する。
【0046】
ここで、説明を図8の障害診断プログラム(210)の流れ図に移す。障害回復確認プログラム(210)は、システムファームウェアの初期部品動作テスト機能を用いて部品の動作テストを行うが(801)、障害の原因はDIMM(207)ではなくシステムボード(202)であるため、前記初期部品動作テストの結果はNGとなり(802)、ブート処理を中断する(805)。
【0047】
説明を図6の流れ図に戻す。保守員は、保守対象サーバ(201)のブート処理の中断を見て、障害が回復しなかったと判断し(608)、交換したDIMM(207)を元に戻し、次に被疑順位が高いシステムボード(202)を交換し(605)、保守対象サーバ(201)を再起動する(606)。保守対象サーバ(201)を再起動することにより、情報処理装置(201)のブート処理が始まり、CPU(205)上で動作する障害回復確認プログラム(210)が処理を開始する。
【0048】
ここで、説明を図8の障害診断プログラム(210)の流れ図に移す。障害回復確認プログラム(210)は、システムファームウェアの初期部品動作テスト機能を用いて部品の動作テストを行い(801)、障害の原因はシステムボード(202)であったため、前記初期部品動作テストの結果はOKとなり(802)、障害回復フラグを“1”にセットし(803)、ブート処理を続行する(804)。
【0049】
説明を図6の流れ図に戻す。保守員は、保守対象サーバ(201)のブート処理の正常終了を見て、障害が回復したと判断する(608)。なお、前記ステップ605において交換した新しいシステムボード(202)のシリアル番号は、“4444444444”であったと仮定する。図12は、現時点の障害発生フラグ(215)、障害回復フラグ(216)、RC退避変数(217)、シリアル番号表(218)、実際の部品のシリアル番号、RC辞書ファイル(226)、ディスプレイ(224)の表示を示す図である。なお、図12の網掛け箇所は、図11との差分を示している。
【0050】
障害回復フラグが“1”にセットされたことを受け、マイコン(219)上で動作する交換部品特定プログラム(221)が処理を開始する(609)。
【0051】
ここで、説明を図9の交換部品特定プログラム(221)の流れ図に移す。交換部品特定プログラム(221)は、障害発生フラグと障害回復フラグが共に“1”であることを確認し(901)、IICバス(203)経由でPIROM(206)、SPD(208)、FRU−ROM(204)に記憶されている現在のCPU(205)、DIMM(207)、システムボード(202)のシリアル番号を参照し(902、906、910)、シリアル番号表(218)に記憶されている障害発生時のCPU(205)、DIMM(207)、システムボード(202)のシリアル番号と差違があるか判定する(903、907、911)。図12より、CPU(205)、DIMM(207)、システムボード(202)のシリアル番号のうち、障害発生時と現在で差違があるのはシステムボード(202)のシリアル番号だけである。よって、システムボード(202)の交換が行われたものと判断し、交換した部品はシステムボード(202)であるという情報とRC待避変数に記憶しておいたRC=“AAAAAAAA”をインターネット(222)経由で保守管理サーバ(227)に送信し(912)、シリアル番号表(218)のシステムボード(202)のシリアル番号を現状に沿うように“4444444444”に更新する(913)。最後に、障害発生フラグ(216)の値を“0”にクリアする(914)。
【0052】
説明を図6の流れ図に戻す。保守管理サーバ(227)上で動作するRC辞書更新プログラム(229)は、マイコン(219)からインターネット(222)経由で交換した部品はシステムボード(202)であるという情報とRC=“AAAAAAAA”を受信し、RC=“AAAAAAAA”と部品=“システムボード(202)”をキーにRC辞書ファイル(226)を検索し、該当するレコードの交換回数を“2”から“3”へ“+1”する。図13は、現時点の障害発生フラグ(215)、障害回復フラグ(216)、RC退避変数(217)、シリアル番号表(218)、実際の部品のシリアル番号、RC辞書ファイル(226)、ディスプレイ(224)の表示を示す図である。なお、図13の網掛け箇所は、図12との差分を示している。
【0053】
以上説明したように、RC辞書ファイル(226)の更新は、保守員が意識することなく行われる。そして、今度また同様の障害が発生した場合は、図14に示す通り、被疑順位1位はシステムボード(202)となり、被疑順位2位はDIMM(207)となり、最初からシステムボード(202)を交換することになり、余分なDIMM(207)の交換作業を省くことができる。これは、保守対象サーバ(201)のダウンタイムの短縮に繋がる。
【0054】
【発明の効果】
以上説明したように、本発明は、実際の保守作業の内容に基づいて障害コード辞書を更新するため、被疑部品の的中率が向上し、ダウンタイムの短縮が図られ、情報処理装置の稼働率を向上することができる。
【0055】
また、保守作業に連動して自動的に障害コード辞書を更新するため、新たな作業を生じることなく、前記効果を得ることができる。
【図面の簡単な説明】
【図1】情報処理装置の障害管理方式の構成を示すブロック図である。
【図2】本発明の一実施形態の構成を示すブロック図である。
【図3】シリアル番号表(218)の構成を示す表である。
【図4】RC辞書ファイル(226)の構成を示す表である。
【図5】被疑部品表示プログラム(228)によるディスプレイ(224)の表示を示す図である。
【図6】本発明の動作を示す流れ図である。
【図7】障害診断プログラム(220)の処理を示す流れ図である。
【図8】障害回復確認プログラム(210)の処理を示す流れ図である。
【図9】交換部品特定プログラム(221)の処理を示す流れ図である。
【図10】各変数とシリアル番号とRC辞書ファイル(226)とディスプレイ(224)表示の初期状態を示す図である。
【図11】各変数と各シリアル番号とRC辞書ファイル(226)とディスプレイ(224)表示のステップ604を終えた時点の状態を示す図である。
【図12】各変数と各シリアル番号とRC辞書ファイル(226)とディスプレイ(224)表示のステップ607を終えた時点の状態を示す図である。
【図13】各変数と各シリアル番号とRC辞書ファイル(226)とディスプレイ(224)表示のステップ609を終えた時点の状態を示す図である。
【図14】各変数と各シリアル番号とRC辞書ファイル(226)とディスプレイ(224)表示の再度同様の障害が発生した後のステップ604を終えた時点の状態を示す図である。
【符号の説明】
101…情報処理装置、102…シリアル番号採取インターフェース、103…部品、104…シリアル番号、105…障害発生検出手段、106…障害回復検出手段、107…サービスプロセッサ、108…障害診断プログラム、109…交換部位特定プログラム、110…公衆回線、111…保守センタ、112…ハードディスク、113…障害コード辞書ファイル、114…保守管理サーバ、115…被疑部品表示プログラム、116…障害コード辞書更新プログラム、117…出力手段、201…保守対象サーバ、202…システムボード、203…IICバス、204…FRU−ROM、205…CPU、206…PIROM、207…DIMM、208…SPD、209…BIOS−ROM、210…障害回復確認プログラム、211…チップセット、212…障害状態レジスタ、213…サービスプロセッサボード、214…NVRAM、215…障害発生フラグ、216…障害回復フラグ、217…RC退避変数、218…シリアル番号表、219…マイコン、220…障害診断プログラム、221…交換部品特定プログラム、222…インターネット、223…保守センタ、224…ディスプレイ、225…ハードディスク、226…RC辞書ファイル、227…保守管理サーバ、228…被疑部品表示プログラム、229…RC辞書更新プログラム。
[0001]
TECHNICAL FIELD OF THE INVENTION
The present invention relates to a failure management method for an information processing apparatus, and more particularly to a method for pointing out a suspected component based on a failure code.
[0002]
[Prior art]
In the conventional fault management method of an information processing device, a fault code dictionary indicating correspondence between a fault code and a suspected component is created in advance, and when a fault occurs, the fault code dictionary is searched using the fault code as a key, and the suspected component is identified. ing.
[0003]
However, the relationship between the fault code and the suspected component is not necessarily one-to-one, and a plurality of components may be suspected for one fault code. Therefore, the suspected component and the suspected order are set as a set in the fault code dictionary, and when a fault occurs, the suspected component is replaced in an order according to the suspected order.
[0004]
However, since the setting of the suspicion order is performed in the design stage of the information processing apparatus, the suspicion order does not reflect the actual failure rate and quality of the parts, and does not accumulate the number of maintenance work after the information processing apparatus is shipped. As far as possible, there has been a problem that a setting error of the suspicion order does not become apparent. For this reason, even after shipment of the information processing apparatus, it is monitored whether or not the setting of the suspected order is different from the actual situation, and the fault code dictionary is updated as necessary.
[0005]
[Patent Document 1]
JP-A-10-320241
[0006]
[Problems to be solved by the invention]
In the above method, as described above, there is a problem that the fault code dictionary must be continuously reviewed even after the information processing apparatus is shipped in order to improve the hit rate of the suspected part.
[0007]
The present invention has been made in view of the above problems, and automatically updates a failure code dictionary in conjunction with maintenance work, thereby eliminating the need for an artificial failure code dictionary update operation, thereby reducing the possibility of a suspect component. The aim is to improve the accuracy of the target.
[0008]
[Means for Solving the Problems]
A failure management method for an information processing apparatus according to the present invention will be described with reference to FIG.
[0009]
According to the fault management method of the present invention, a fault occurrence detecting means (105) detects occurrence of a fault in a component (103) of an information processing apparatus (101), and a fault diagnosis program (108) operating on a service processor (107) is used. The failure code is transmitted to the maintenance management server (114) of the maintenance center (111) via the public line (110), and the suspected part display program (115) operating on the maintenance management server (114) receiving the failure code is transmitted. In a fault management method for an information processing apparatus, in which a suspected component estimated based on a fault code dictionary file (113) stored in a hard disk (112) is displayed on an output unit (117), a part ( The fault recovery detecting means (106) detects the fault recovery of (103), and the replacement part specifying program operating on the service processor (107). The maintenance management server (114) having transmitted the type information of the parts replaced by the gram (109) for maintenance via the public line (110) to the maintenance management server (114) of the maintenance center (111) and receiving the information. The failure code dictionary update program (116) operating on the above updates the failure code dictionary file (113) stored in the hard disk (112).
[0010]
Further, a replacement part specifying program (109) operating on the service processor (107) is uniquely added to the part (103) of the information processing apparatus (101) as means for specifying the replaced part (103). It is characterized in that the serial number (104) is referred to via the serial number collection interface (102), and the serial number (104) before the occurrence of the failure and after the recovery from the failure are compared.
[0011]
Furthermore, an initial component operation test performed by the system firmware when the information processing apparatus is booted is used as the failure recovery detection means (106).
[0012]
A maintenance center is not provided, and the failure occurrence detecting means detects the occurrence of a failure in a component of the information processing apparatus, and the suspected component that operates on the service processor based on the failure code created by the failure diagnosis program that operates on the service processor. In the fault management method for an information processing device, in which a suspected component estimated by a display program is displayed on an output unit, a fault recovery detecting unit detects fault recovery of a component of the information processing device, and a replacement component specifying program operating on a service processor is used. A configuration in which a failure code dictionary update program that operates on a service processor based on type information of parts replaced for maintenance updates a failure code dictionary file stored in a hard disk is also a means for solving the above problem. Can be taken as
[0013]
BEST MODE FOR CARRYING OUT THE INVENTION
Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 2 is a block diagram showing a configuration of one embodiment of the present invention.
[0014]
The maintenance target server (201) is an information processing apparatus that has a maintenance service contract with the maintenance center (223). The replaceable parts of the maintenance target server (201) are the CPU (205), the DIMM (207), and the system board (202), and a unique serial number is added to each part. The CPU (205) and the DIMM (207) are mounted on the system board (202), but can be freely attached and detached.
[0015]
The serial number of the CPU (205) is stored in the PIROM (206). The PIROM (206) is a EEPROM built in the CPU (205) and stores information on the CPU (205).
[0016]
The serial number of the DIMM (207) is stored in the SPD (208). The SPD (208) is a EEPROM built in the DIMM (207), and stores information on the DIMM (207).
[0017]
The serial number of the system board (202) is stored in the FRU-ROM (204). The FRU-ROM (204) is a SEEPROM mounted on the system board (202), and stores information on the system board (202).
[0018]
On the system board (202), a chipset (211) and a BIOS-ROM (209) are mounted in addition to the CPU (205), the DIMM (207), and the FRU-ROM (204).
[0019]
The chip set (211) controls access from the CPU (205) to the DIMM (207), the BIOS-ROM (209), and the NVRAM (214), and the CPU (205), the DIMM (207), and the chip set (211). It has a function of detecting a failure that has occurred in the internal processing and communication processing of the. When detecting a failure, the function stores failure information in a failure status register (212) built in the chipset (211).
[0020]
The BIOS-ROM (209) is an EEPROM in which the code of the system firmware of the maintenance target server (201) is stored. The system firmware operates on the CPU (205) and performs boot processing of the maintenance target server (201). The boot process includes an initial component operation test function for testing whether the CPU (205), the DIMM (207), and the system board (202) operate normally. The failure recovery confirmation program (210) is a program for confirming whether the failure has been recovered by using the initial component operation test function.
[0021]
FIG. 8 is a flowchart showing the processing of the failure recovery confirmation program (210). The failure recovery confirmation program (210) tests whether the CPU (205), the DIMM (207), and the system board (202) operate normally by using the initial component operation test function (801). It is further confirmed whether the failure has been recovered (802). If the recovery of the failure has been confirmed, the failure recovery flag is set to "1" (803), and the boot process is continued (804). On the other hand, if the recovery from the failure cannot be confirmed, the boot process is interrupted (805).
[0022]
In the maintenance target server (201), a service processor board (213) is mounted in addition to the system board (202). An NVRAM (214) and a microcomputer (219) are mounted on the service processor board (213).
[0023]
The NVRAM (214) is a non-volatile memory accessible from both the CPU (205) and the microcomputer (219), and includes a failure occurrence flag (215), a failure recovery flag (216), an RC save variable (217), a serial A number table (218) is arranged.
[0024]
The failure occurrence flag (214) is a binary variable indicating a failure occurrence state. The failure flag (214) = "0" means that no failure has occurred, and the failure flag (214) = "1" means that a failure has occurred.
[0025]
The failure recovery flag (215) is a binary variable indicating a failure recovery state. The failure recovery flag (215) = "0" means that the failure has not been recovered, and the failure recovery flag (215) = "1" means that the failure has been recovered.
[0026]
The RC save variable (216) is a variable used to store the RC created by the failure diagnosis program (220). Note that RC is an abbreviation for “ReferenceCode” and is a term corresponding to a failure code.
[0027]
The serial number table (218) is an array variable used to store the serial numbers of the CPU (205), the DIMM (207), and the system board (202). FIG. 3 is a table showing the configuration of the serial number table (218). The index of the array corresponds to the type of the component, and is defined as “1” = CPU (205), “2” = DIMM (207), and “3” = system board (202). In the elements of the array, the serial numbers of the parts corresponding to the subscripts are stored.
[0028]
The microcomputer (219) is an embedded controller that operates independently of the CPU (205), and has a LAN communication function, an IIC bus access function, and a built-in ROM.
[0029]
The LAN access function is a function for performing communication via the Internet (222). The microcomputer (219) can communicate with the maintenance management server (227) via the Internet (222).
[0030]
The IIC bus access function is a function for performing access via the IIC bus (203). The microcomputer (219) can access the FRU-ROM (204), the PIROM (206), the SPD (208), and the fault status register (212) via the IIC bus (203).
[0031]
The built-in ROM stores a fault diagnosis program (220) and a replacement part specifying program (221).
[0032]
The failure diagnosis program (220) is a program that detects the occurrence of a failure and notifies the maintenance management server (227) of the RC. FIG. 7 is a flowchart showing the processing of the failure diagnosis program (220). The fault diagnosis program (220) refers to the fault status register (212) via the IIC bus (701), determines whether fault information is stored in the fault status register (212) (702), and checks the fault status register (212). If the failure information is stored in ()), an RC is created based on the failure information (703), the RC is stored in an RC save variable (217) (704), and the RC is transmitted via the Internet (222). The value is transmitted to the maintenance management server (227) (705), the value of the failure recovery flag (216) is cleared to "0" (706), and the value of the failure occurrence flag (215) is set to "1" (707). .
[0033]
The replacement part specifying program (221) is a program that detects that the failure has recovered, and notifies the maintenance management server (227) of the type information of the part replaced for maintenance. FIG. 9 is a flowchart showing the processing of the replacement part specifying program (221). The replacement part specifying program (221) determines whether both the value of the failure occurrence flag (215) and the value of the failure recovery flag (216) are "1" (901), and determines the value of the failure occurrence flag (215) and the failure recovery flag (215). If both values of “216” are “1”, the current CPU (205) stored in the PIROM (206), SPD (208), and FRU-ROM (204) via the IIC bus (203), and the DIMM ( 207), refer to the serial numbers of the system board (202) (902, 906, 910), and store the CPU (205), DIMM (207), and system board (207) at the time of the failure stored in the serial number table (218). It is determined whether there is a difference from the serial number of step 202) (903, 907, 911). Then, the type information of the part having the difference in the serial number is transmitted to the maintenance management server (227) via the Internet (222) together with the RC stored in the RC save variable (904, 908, 912), and the serial number table is displayed. (218) is updated so as to conform to the current state (905, 909, 913). Finally, the value of the failure occurrence flag (215) is cleared to "0" (914).
[0034]
The maintenance center (223) is a building in which maintenance personnel who perform maintenance work of the maintenance target server (201) are registered, and a display (224), a hard disk (225), and a maintenance management server (227) are installed.
[0035]
The display (224) is a display device that displays an output of the maintenance management server (227).
[0036]
The hard disk (225) is a storage device for storing files of the maintenance management server (227), and stores an RC dictionary file (226).
[0037]
The RC dictionary file (226) is a file indicating the correspondence between the RC and the suspected part. FIG. 4 is a table showing the configuration of the RC dictionary file (226). The RC dictionary file (226) has four items: RC, parts, number of replacements, and priority. The key items are RC and parts. The number of replacements indicates the number of times that the failure has been recovered by replacing the component with the RC. The priority indicates the priority of the suspected case where a plurality of parts having the same number of replacements exist for the RC, and is set in advance by the designer of the maintenance target server (201). Note that the priority is “1”, which is the highest, and lower as the numerical value increases. When the priority is "0", it indicates that the part is not a suspected part.
[0038]
The maintenance management server (227) is a server that manages the maintenance target server (201), and stores a suspected part display program (228) and an RC dictionary update program (229).
[0039]
The suspected part display program (228) receives the RC from the microcomputer (219) via the Internet (222), searches the RC dictionary file (226) using the RC as a key, and displays the suspected part on the display (224). It is a program. When there are a plurality of suspected parts, each of the suspected parts is displayed with a suspected rank. FIG. 5 is a diagram showing a display on the display (224) by the suspected part display program (228), and corresponds to the RC dictionary file (226) in FIG. FIG. 5A is a display in the case of RC = “AAAAAAAAA”. 5, the number of replacements of the DIMM (207) is “8”, which is larger than the number of replacements of the system board (202), “6”. Here, if there is a difference in the number of exchanges, the priority is not used. Therefore, as shown in FIG. 5A, the first place in the suspected order is the DIMM (207), and the second place in the suspected order is the system board (202). The CPU (205) is not a suspect because the priority is “0”. FIG. 7B shows a display when RC = "BBBBBBBBB". From FIG. 4, the number of replacements of the CPU (205) is "4", which is equal to "4" which is the number of replacements of the system board (202). If the number of exchanges is equal, a priority order is assigned using the priority order. 4, the priority of the CPU (205) is “1”, which is higher than the priority of the system board (202), “2”. Therefore, as shown in FIG. 5B, the first place of the suspicion rank is the CPU (205) and the second place of the suspicion rank is the system board (202). Since the priority of the DIMM (207) is “0”, the DIMM (207) is not a suspect.
[0040]
The RC dictionary update program (229) receives from the microcomputer (219) the type information of the replaced parts and the RC via the Internet (222), searches the RC dictionary file (226) using the RC and the parts as keys, This is a program for incrementing the number of exchanges of the corresponding record by “+1”.
[0041]
Next, an operation from occurrence of a failure to updating of the RC dictionary file (226) will be described with reference to the flowchart of FIG.
[0042]
FIG. 10 shows a failure occurrence flag (214), a failure recovery flag (215), an RC save variable (216), a serial number table (217), an actual part serial number, an RC dictionary file (226), It is a figure showing a display of display (229). In the state of FIG. 10, it is assumed that a failure due to the system board (202) has occurred (601). The chipset (211) detects the fault and stores the fault information in the fault status register (212) (602). When the fault information is stored in the fault status register (212), the fault diagnosis program (220) running on the microcomputer (219) starts processing (603).
[0043]
Here, the explanation is shifted to the flowchart of the fault diagnosis program (220) in FIG. The fault diagnosis program (220) refers to the fault status register (212) via the IIC bus (203) (701), and detects that the fault information is stored in the fault status register (212) (702). An RC is created based on the failure information (703), the RC is stored in an RC save variable (217) (704), and the RC is transmitted to the maintenance management server (227) via the Internet (222) ( 705), the value of the failure recovery flag (216) is cleared to "0" (706), and the value of the failure occurrence flag (215) is set to "1" (707). It is assumed that the RC created in step 703 is “AAAAAAAAA”.
[0044]
The description returns to the flowchart of FIG. The maintenance management server (227) receives the RC from the microcomputer (219) via the Internet (222) and activates the suspected part display program (228) (604). The suspected part display program (228) searches the RC dictionary file (226) using the RC as a key, and displays the suspected part on the display (224). From the RC dictionary file (226) in FIG. 10, the number of replacements of the component corresponding to the RC = “AAAAAAAAA” is “2” for both the DIMM (207) and the system board (202). If the number of exchanges is equal, a priority order is assigned using the priority order. The priority is "1" for the DIMM (207), which is higher than "2" for the system board (202). Therefore, the first place in the suspicion order is the DIMM (207), and the second place in the suspicion order is the system board (202). FIG. 11 shows a current failure occurrence flag (215), a failure recovery flag (216), an RC save variable (217), a serial number table (218), an actual part serial number, an RC dictionary file (226), and a display ( FIG. 224 is a diagram showing a display of FIG. The shaded portions in FIG. 11 indicate the differences from FIG.
[0045]
Next, the maintenance staff at the maintenance center (223) looks at the display on the display (224), replaces the DIMM (207) with the highest suspicion rank (605), and re-installs the maintenance target server (201). Activate (606). By restarting the maintenance target server (201), boot processing of the information processing apparatus (201) starts, and the failure recovery confirmation program (210) running on the CPU (205) starts processing.
[0046]
Here, the explanation is shifted to the flowchart of the failure diagnosis program (210) in FIG. The failure recovery confirmation program (210) performs an operation test of parts using the initial part operation test function of the system firmware (801). However, since the cause of the failure is not the DIMM (207) but the system board (202), The result of the initial component operation test is NG (802), and the boot process is interrupted (805).
[0047]
The description returns to the flowchart of FIG. The maintenance worker sees the interruption of the boot process of the server to be maintained (201), determines that the failure has not been recovered (608), returns the replaced DIMM (207), and returns the system board with the next highest suspicion rank. (202) is replaced (605), and the maintenance target server (201) is restarted (606). By restarting the maintenance target server (201), boot processing of the information processing apparatus (201) starts, and the failure recovery confirmation program (210) running on the CPU (205) starts processing.
[0048]
Here, the explanation is shifted to the flowchart of the failure diagnosis program (210) in FIG. The failure recovery confirmation program (210) performs a component operation test using the initial component operation test function of the system firmware (801). Since the cause of the failure is the system board (202), the result of the initial component operation test is obtained. Is OK (802), the failure recovery flag is set to "1" (803), and the boot process is continued (804).
[0049]
The description returns to the flowchart of FIG. The maintenance worker sees the normal end of the boot processing of the maintenance target server (201) and determines that the failure has been recovered (608). It is assumed that the serial number of the new system board (202) replaced in step 605 is "44444444444". FIG. 12 shows a failure occurrence flag (215), a failure recovery flag (216), an RC save variable (217), a serial number table (218), a serial number of an actual part, an RC dictionary file (226), and a display ( FIG. 224 is a diagram showing a display of FIG. Note that the hatched portions in FIG. 12 indicate the differences from FIG.
[0050]
In response to the failure recovery flag being set to "1", the replacement part specifying program (221) operating on the microcomputer (219) starts processing (609).
[0051]
Here, the description is shifted to the flowchart of the replacement part specifying program (221) in FIG. The replacement part identification program (221) confirms that both the failure occurrence flag and the failure recovery flag are "1" (901), and via the IIC bus (203) the PIROM (206), SPD (208), FRU- The current CPU (205), DIMM (207) and system board (202) serial numbers stored in the ROM (204) are referred to (902, 906, 910) and stored in the serial number table (218). It is determined whether there is a difference between the serial numbers of the CPU (205), DIMM (207), and system board (202) at the time of occurrence of the fault (903, 907, 911). From FIG. 12, among the serial numbers of the CPU (205), the DIMM (207), and the system board (202), only the serial number of the system board (202) is different from that at the time of occurrence of the failure. Therefore, it is determined that the system board (202) has been replaced, and information that the replaced part is the system board (202) and RC = “AAAAAAAAA” stored in the RC save variable are transmitted to the Internet (222). ) To the maintenance management server (227) (912), and updates the serial number of the system board (202) in the serial number table (218) to “44444444444” so as to conform to the current state (913). Finally, the value of the failure flag (216) is cleared to "0" (914).
[0052]
The description returns to the flowchart of FIG. The RC dictionary update program (229) operating on the maintenance management server (227) includes information that the component replaced from the microcomputer (219) via the Internet (222) is the system board (202) and RC = “AAAAAAAAA”. The RC dictionary file (226) is searched using the RC = “AAAAAAAAA” and the component = “system board (202)” as keys, and the number of times of exchange of the corresponding record is “+1” from “2” to “3”. . FIG. 13 shows a failure occurrence flag (215), a failure recovery flag (216), an RC save variable (217), a serial number table (218), a serial number of an actual part, an RC dictionary file (226), and a display ( FIG. 224 is a diagram showing a display of FIG. Note that the shaded portions in FIG. 13 indicate the differences from FIG.
[0053]
As described above, the update of the RC dictionary file (226) is performed without the care of the maintenance staff. Then, when the same failure occurs again, as shown in FIG. 14, the first place in the suspected order becomes the system board (202), the second place in the suspected order becomes the DIMM (207), and the system board (202) is replaced from the beginning. Since the replacement is performed, the replacement work of the extra DIMM (207) can be omitted. This leads to a reduction in downtime of the maintenance target server (201).
[0054]
【The invention's effect】
As described above, according to the present invention, the failure code dictionary is updated based on the contents of the actual maintenance work, so that the hit rate of the suspected part is improved, downtime is reduced, and the operation of the information processing apparatus is reduced. Rate can be improved.
[0055]
Further, since the fault code dictionary is automatically updated in conjunction with the maintenance work, the above-described effect can be obtained without generating a new work.
[Brief description of the drawings]
FIG. 1 is a block diagram illustrating a configuration of a fault management system of an information processing apparatus.
FIG. 2 is a block diagram showing a configuration of one embodiment of the present invention.
FIG. 3 is a table showing a configuration of a serial number table (218).
FIG. 4 is a table showing a configuration of an RC dictionary file (226).
FIG. 5 is a diagram showing a display on a display (224) by a suspected part display program (228).
FIG. 6 is a flowchart showing the operation of the present invention.
FIG. 7 is a flowchart showing processing of a failure diagnosis program (220).
FIG. 8 is a flowchart showing processing of a failure recovery confirmation program (210).
FIG. 9 is a flowchart showing processing of a replacement part specifying program (221).
FIG. 10 is a diagram showing an initial state of each variable, a serial number, an RC dictionary file (226), and a display (224).
FIG. 11 is a diagram showing a state at the end of step 604 for displaying each variable, each serial number, an RC dictionary file (226), and a display (224).
FIG. 12 is a diagram showing a state at the end of step 607 for displaying each variable, each serial number, an RC dictionary file (226), and a display (224).
FIG. 13 is a diagram showing a state at the time when step 609 of displaying each variable, each serial number, the RC dictionary file (226), and the display (224) is completed.
FIG. 14 is a diagram showing a state at the time when step 604 is ended after a similar failure again occurred on each variable, each serial number, the RC dictionary file (226), and the display (224).
[Explanation of symbols]
DESCRIPTION OF SYMBOLS 101 ... Information processing apparatus, 102 ... Serial number collection interface, 103 ... Parts, 104 ... Serial number, 105 ... Fault detection means, 106 ... Fault recovery detection means, 107 ... Service processor, 108 ... Fault diagnosis program, 109 ... Replacement Part identification program, 110 public line, 111 maintenance center, 112 hard disk, 113 failure code dictionary file, 114 maintenance management server, 115 suspicious part display program, 116 failure code dictionary update program, 117 output means , 201: maintenance target server, 202: system board, 203: IIC bus, 204: FRU-ROM, 205: CPU, 206: PIROM, 207: DIMM, 208: SPD, 209: BIOS-ROM, 210: failure recovery confirmation program, 11 chip set, 212 fault status register, 213 service processor board, 214 NVRAM, 215 fault occurrence flag, 216 fault recovery flag, 217 RC save variable, 218 serial number table, 219 microcomputer, 220 ... Fault diagnosis program, 221 ... Replacement part identification program, 222 ... Internet, 223 ... Maintenance center, 224 ... Display, 225 ... Hard disk, 226 ... RC dictionary file, 227 ... Maintenance management server, 228 ... Suspicious part display program, 229 ... RC dictionary update program.

Claims (4)

情報処理装置の部品の障害発生を障害発生検出手段が検出し、サービスプロセッサ上で動作する障害診断プログラムが障害コードを公衆回線経由で保守センタの保守管理サーバへ送信し、前記障害コードを受信した保守管理サーバ上で動作する被疑部品表示プログラムがハードディスクに格納されている障害コード辞書ファイルを基に推定した被疑部品を出力手段に表示する情報処理装置の障害管理方式において、情報処理装置の部品の障害回復を障害回復検出手段が検出し、サービスプロセッサ上で動作する交換部品特定プログラムが保守のために交換した部品の種別情報を公衆回線経由で保守センタの保守管理サーバへ送信し、前記情報を受信した保守管理サーバ上で動作する障害コード辞書更新プログラムがハードディスクに格納されている障害コード辞書ファイルを更新することを特徴とする情報処理装置の障害管理方式。The fault occurrence detecting means detects the occurrence of a fault in a part of the information processing apparatus, and the fault diagnosis program operating on the service processor transmits the fault code to the maintenance management server of the maintenance center via the public line, and receives the fault code. In a fault management method for an information processing device, a suspected component display program operating on a maintenance management server displays a suspected component on an output unit based on a fault code dictionary file stored on a hard disk. Failure recovery is detected by the failure recovery detection means, and the replacement part identification program operating on the service processor transmits the type information of the part replaced for maintenance to the maintenance management server of the maintenance center via a public line, and transmits the information. The received fault code dictionary update program that runs on the maintenance management server is stored on the hard disk. Fault management method of an information processing apparatus and updates the fault code dictionary file are. サービスプロセッサ上で動作する交換部品特定プログラムが交換された部品を特定する手段として、情報処理装置の部品に一意に付加されているシリアル番号をシリアル番号採取インタフェース経由で参照し、障害発生前と障害回復後のシリアル番号を比較することを特徴とする請求項1記載の情報処理装置の障害管理方式。As a means for the replacement part identification program running on the service processor to identify the replaced part, the serial number uniquely added to the part of the information processing unit is referenced via the serial number collection interface, and before and after the failure occurs. 2. The fault management system according to claim 1, wherein the serial numbers after the recovery are compared. 障害回復検出手段として、情報処理装置のブート時にシステムファームウェアが行う初期部品動作テストを用いることを特徴とする請求項1記載の情報処理装置の障害管理方式。2. The fault management method for an information processing apparatus according to claim 1, wherein an initial component operation test performed by system firmware when the information processing apparatus is booted is used as the fault recovery detection means. 情報処理装置の部品の障害発生を障害発生検出手段が検出し、サービスプロセッサ上で動作する障害診断プログラムが作成した障害コードを基にサービスプロセッサ上で動作する被疑部品表示プログラムが推定した被疑部品を出力手段に表示する情報処理装置の障害管理方式において、情報処理装置の部品の障害回復を障害回復検出手段が検出し、サービスプロセッサ上で動作する交換部品特定プログラムが保守のために交換した部品の種別情報を基にサービスプロセッサ上で動作する障害コード辞書更新プログラムがハードディスクに格納されている障害コード辞書ファイルを更新することを特徴とする情報処理装置の障害管理方式。Failure occurrence detection means detects the occurrence of a failure in a component of the information processing device, and based on the failure code created by the failure diagnosis program operating on the service processor, the suspected component estimated by the suspected component display program operating on the service processor based on the failure code. In the fault management method of the information processing device displayed on the output means, the fault recovery detecting means detects the fault recovery of the component of the information processing device, and the replacement component specifying program running on the service processor identifies the component replaced for maintenance. A failure management method for an information processing apparatus, wherein a failure code dictionary update program operating on a service processor updates a failure code dictionary file stored on a hard disk based on type information.
JP2003153705A 2003-05-30 2003-05-30 Failure management method for information processing equipment Pending JP2004355424A (en)

Priority Applications (1)

Application Number Priority Date Filing Date Title
JP2003153705A JP2004355424A (en) 2003-05-30 2003-05-30 Failure management method for information processing equipment

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
JP2003153705A JP2004355424A (en) 2003-05-30 2003-05-30 Failure management method for information processing equipment

Publications (1)

Publication Number Publication Date
JP2004355424A true JP2004355424A (en) 2004-12-16

Family

ID=34048556

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2003153705A Pending JP2004355424A (en) 2003-05-30 2003-05-30 Failure management method for information processing equipment

Country Status (1)

Country Link
JP (1) JP2004355424A (en)

Cited By (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008527554A (en) * 2005-01-18 2008-07-24 インターナショナル・ビジネス・マシーンズ・コーポレーション Method and system for fault diagnosis and maintenance in computer systems (history-based prioritization of suspicious components)
JP2010231666A (en) * 2009-03-27 2010-10-14 Nec Personal Products Co Ltd Repair part instruction system, repair part instruction apparatus, repair information management apparatus, method and program thereof
JP2011175513A (en) * 2010-02-25 2011-09-08 Nec Computertechno Ltd Fault management system and method
CN103197999A (en) * 2013-03-22 2013-07-10 北京百度网讯科技有限公司 Method and device for automatically positioning internal memory fault
CN114322202A (en) * 2021-12-20 2022-04-12 青岛海尔空调器有限总公司 Fault self-diagnosis method and system based on cloud server

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008527554A (en) * 2005-01-18 2008-07-24 インターナショナル・ビジネス・マシーンズ・コーポレーション Method and system for fault diagnosis and maintenance in computer systems (history-based prioritization of suspicious components)
JP2010231666A (en) * 2009-03-27 2010-10-14 Nec Personal Products Co Ltd Repair part instruction system, repair part instruction apparatus, repair information management apparatus, method and program thereof
JP2011175513A (en) * 2010-02-25 2011-09-08 Nec Computertechno Ltd Fault management system and method
CN103197999A (en) * 2013-03-22 2013-07-10 北京百度网讯科技有限公司 Method and device for automatically positioning internal memory fault
CN103197999B (en) * 2013-03-22 2016-08-03 北京百度网讯科技有限公司 A kind of memory failure automatic positioning method and device
CN114322202A (en) * 2021-12-20 2022-04-12 青岛海尔空调器有限总公司 Fault self-diagnosis method and system based on cloud server

Similar Documents

Publication Publication Date Title
JP3200661B2 (en) Client / server system
US7287193B2 (en) Methods, systems, and media to correlate errors associated with a cluster
US10761926B2 (en) Server hardware fault analysis and recovery
EP0333620B1 (en) On-line problem management for data-processing systems
CN100451977C (en) System and method to detect errors and predict potential failures
CN109308252B (en) Fault positioning processing method and device
EP0474058A2 (en) Problem analysis of a node computer with assistance from a central site
CN112948182B (en) Method and system for recovering and upgrading emergency backup of set top box
CN100382041C (en) Computer system and fault computer replacement method applied to same
CN119127318A (en) Operating system OS startup method, device and medium
US7730029B2 (en) System and method of fault tolerant reconciliation for control card redundancy
JP2004355424A (en) Failure management method for information processing equipment
CN120255970B (en) Baseboard management controller starting method, computer equipment, medium and product
WO2021061146A1 (en) Lifecycle change cryptographic ledger
CN118694698B (en) Method for detecting network equipment state and correcting errors by domestic operating system
CN119046051A (en) Fault processing method and product of computer system
Bowman et al. 1a processor: Maintenance software
CN116340031B (en) Computer systems and methods for detecting deviations and non-transitory computer-readable media
KR20170070568A (en) System and method for managing servers totally
CN114860494A (en) SAS expander configuration self-adaptive system
JP2010198314A (en) Information management device
CN112667167A (en) Configuration file updating method and device
CN114996045B (en) A kind of automatic recovery system and method of in-memory database fault
JP5439736B2 (en) Computer management system, computer system management method, and computer system management program
JP2011159234A (en) Fault handling system and fault handling method