manawize

CASE STUDY – Lessons learnt from a Friday kernel crash

When Friday afternoon meets a server failure – how we got our client’s critical system back up and running.

One of our clients’ central servers was running several mission-critical services on a single machine: these included, amongst others, the CRM system and the company chat service. The machine was an older Lenovo server and had recently experienced several power cuts; although daily backups were taken, a potential hardware failure would have meant days or weeks of downtime. 

 

The problem

On Friday afternoon, the company’s staff experienced slow performance and delayed message delivery – this is what the fault report referred to. After the reboot, the server failed to load: it got stuck whilst loading the kernel (the runtime environment that forms the basis of the operating system), and remote access was also unavailable. A colleague of the client’s at the site attempted to restore operations following instructions given over the phone and via video call – with little success. By the end of Friday, the entire company’s operations had therefore come to a standstill.

 

The solution

On Saturday, we rectified the error in several stages:

  • Whilst still on site, we made a backup copy of the most recent backup onto an external hard drive.
  • We then transported the server to our office in comfort.
  • Following a more detailed analysis, the kernel appeared to be the source of the problem.
  • Fortunately, we had an earlier version of the kernel, so we restored the server to that.
  • Even after restarting later, we encountered further problems caused by the BIOS reset. (The battery had run flat.)
  • Following a battery replacement, a power cut simulation and several restarts, the system proved to be stable.

 

On Monday, the services were up and running again; on Tuesday, the server was moved back to its original location.

Lessons learnt and our next steps

Thanks to the rapid troubleshooting, we managed to ensure the client could start work on Monday. However, the problem highlighted that relying on a single ageing server to run the business poses a very serious risk.

The next steps are therefore as follows:
  • Planning and proposing alternative infrastructure (Cloud, Hybrid or a new on-premises environment),
  • Designing for redundant, scalable and secure operation.
  • And if our client ultimately decides to go with an on-premises solution after all, and there is no room for a second server or a move to the cloud, then we would consider relocating the server closer to the client and storing spare hardware components in a warehouse. This would significantly reduce the risk of downtime.

 

This latest incident also highlighted to the client that the time required for troubleshooting can in itself pose a significant risk. As all services were running on a single physical machine, we were unable to ensure continuity of service on an alternative system whilst the investigation was underway.

In a well-designed infrastructure, fault rectification and business operations do not come at the expense of one another: the standby environment takes over operations, whilst the technical team can calmly rectify the original fault.

If you feel that your IT infrastructure isn’t on a particularly solid footing, get in touch with us and we’ll help you find a cost-effective solution.

Our latest blog posts:

Share it with others!

How can we help your company?

Have a question?
Would you like to give us a try?
Feel free to write to me!

Are you ready for the next step? Request a personalised quote now!
We will get back to you within 24 hours.

IT Service Request Form – New Gen
Adatvédelmi áttekintés

Ez a weboldal sütiket használ, hogy a lehető legjobb felhasználói élményt nyújthassuk. A cookie-k információit tárolja a böngészőjében, és olyan funkciókat lát el, mint a felismerés, amikor visszatér a weboldalunkra, és segítjük a csapatunkat abban, hogy megértsék, hogy a weboldal mely részei érdekesek és hasznosak.