Translate

Sunday, August 2, 2026

Server chronicles - another repair (UPDATED)

 The entropy can only increase, so after some time I had the server shut down, when I restarted something went wrong.
I didn't notice any problem until I did walk close to it.
I did see the Red LED blinking, indicating a problem.

OK, so we need to repair something.

Step 1 - find out the problem

The blinking LED is not much as information, so as the first step I went on the ILo system, looking for clues.

(on my server the ILo is another IP assigned to it, so I just opened a browser on that address)

On the System Information - Health Summary, only a warning sign on the Storage !
Oh no !!!! Now what ?

Ok, Clicking on Storage (where the warning sign is) reported few more details :


So the  problem is on the HPE Smart Array.
It is the module that basically handle the RAID.

But what exactly is the problem ?

The cache, a memory that stores data before to be written on the hard disks, is "degraded".
There are different possible problems, so I started to look around for a more specific way to diagnose the problem.

I run on the server Ubuntu 22.04 LTE.
A search for tools to diagnose the cache for RAID, for the HP Proliant DL360p Gen8, did show I needed to have a tool called ssacli (Smart Storage Administrator Command Line Interface).

Of course is not in the standard repo of Ubuntu.

Step 2 - install ssacli

To install ssacli we need to add an HPE repo to our list

  • sudo curl -fsSL https://downloads.linux.hpe.com/SDR/hpePublicKey2048_key1.pub | sudo gpg --dearmor -o /usr/share/keyrings/hpePublicKey2048_key1.gpg

  • echo "deb [signed-by=/usr/share/keyrings/hpePublicKey2048_key1.gpg] http://downloads.linux.hpe.com/SDR/repo/mcp jammy/current non-free" | sudo tee /etc/apt/sources.list.d/hp-mcp.list

After that a good update and upgrade

  • sudo apt update
  • sudo apt upgrade
and finally 
  • sudo apt install ssacli

Step 3 - diagnose the problem

Now with the new tool, did run a general diagnostic

$ sudo ssacli ctrl all show status

Smart Array P420i in Slot 0 (Embedded)
   Controller Status: OK
   Cache Status: Permanently Disabled
   Battery/Capacitor Status: Failed (Replace Batteries/Capacitors)

Well !!! I would say a clear diagnostic !!

A more deeper search on the web, and I found that the Smart Array P420i (the RAID controller basically) has two components :

  • a DIMM memory/controller   (and the status is OK)
  • a battery/capacitor - that is the problem
The same search did show what to look on eBay (or other places), the battery/capacitor I need is  this one :

(On eBay for 24$) 

Ordered one, and when it will arrive, I'll change it.

A safety note

Looking for the problem, I found out that the server is still safe to be used.
When this error happens, the RAID is degraded and thus not used with the cache.
It means of course that the server operates under performance, but the data will be safe.

So matter of few days and I could even took the server offline again for a while.

What to expect

Assuming the only problem is the battery/capacitor, when the replacement arrives :

  • I will change it
  • restart the server
On forum and other places, people said to wait until 24 hours before to consider the repair done.
For a while is possible the warning sign on ILo will remain on until the POST will clear the error.
No need to reset something or issue commands.


Step 4 - repair

Arrived the replacement for the battery/capacitor and finally, in a hot afternoon, substituted the damaged capacitor on the server.
Not really a complicate thing but without completely detaching the server from all the cables and put it on a table, i.e. repairing on site, it was kind of fun.

VERY IMPORTANT !
Before to work on the server, remove BOTH the power cables (the server has 2 power supply) and wait at least 5 minutes to facilitate the discharge of capacitors in the server !

Some pictures.

The old battery is close to the fan number 3
However the Raid board is on the back of the server
The cable entering the Raid card
To extract the RAID card is possible to raise little bit the metal block visible behind the cable

The RAID card - not the cable connector.
The new battery installed

The old battery.
Is not easy to see in the photo, but the capacitors are bulged !
Definitively gone!

After the change, I did reconnect the power cables, did wait 1 minute and then I did power on the server.
No more blinking RED light.
ILO was happy, everything green now.

However still a note in the log :

POST Error: 1705-Slot X Drive Array - Please replace Cache Module Super-Cap. Caching will be enabled once Super-Cap has been replaced and charged.


This error should be gone in a day or two, until the new capacitor will be full charged and the BIOS POST will fully recognize it.

To be sure I did run a RAID check :

$ sudo ssacli ctrl all show config
Smart Array P420i in Slot 0 (Embedded)    (sn: 0014380334BAFF0)


   Internal Drive Cage at Port 1I, Box 1 (Index 0), OK

   Port Name: 1I
   Port Name: 2I
   Array A (SATA, Unused Space: 0  MB)
      logicaldrive 1 (3.64 TB, RAID 1, OK)
      physicaldrive 1I:1:1 (port 1I:box 1:bay 1, SATA HDD, 4 TB, OK)
      physicaldrive 1I:1:2 (port 1I:box 1:bay 2, SATA HDD, 4 TB, OK)
   SEP (Vendor ID PMCSIERA, Model SRCv8x6G) 380  (WWID: 50014380334BAFFF)
 


The RAID seems back to the original setting.
Repair done !

No comments:

Post a Comment