Introduction
Imagine waking up to a digital apocalypse – flights grounded, call centers silenced, and surgeries suddenly canceled.
That’s the scene many faced on Friday, July 19, as Windows systems worldwide crashed, leaving chaos in their wake. The culprit? A botched update from cybersecurity giant CrowdStrike.
Now, CEO George Kurtz has stepped forward with a heartfelt apology and a promise to make things right.
Kurtz addressed the gravity of the situation head-on. “I want to sincerely apologize directly to all of you for today’s outage,” he wrote, acknowledging the widespread disruption caused by the incident. He also clarified that this wasn’t the result of a cyber attack.
So then, what actually went wrong?
The crash was triggered by a defect in a Falcon content update for Windows hosts.
This defective update, rolled out at 04:09 UTC (00:09 Eastern Time), wasn’t fixed until 79 minutes later. By then, it was already too late. Systems running Falcon sensor for Windows 7.11 and above – which downloaded the update during that window – crashed spectacularly.
As soon as the error was identified, CrowdStrike’s engineers acted fast to remove the defective content from Channel 291, the file that was storing the defective update. This stopped further crashes, but for many systems, the damage had already been done.
To aid recovery, CrowdStrike provided detailed instructions for detecting and fixing affected systems. This included guidance on remotely recovering systems and temporary workarounds for both physical machines and virtual servers.
Despite these efforts, the fallout was severe, with many customers facing significant operational disruptions.
What can we learn from this?
So, what is the best way to respond to – or perhaps prevent being impacted by – such an event? Here are some things that you can learn from this.
1- Restore operations securely
Use trusted sources and official updates to restore your systems. CrowdStrike is actively assisting customers impacted by the faulty update. Customers are advised to visit the support portal for the latest updates.
2- Reinforce data protections
Regular backups should be maintained and securely stored to protect against data loss. Additionally, sensitive data should be encrypted to prevent unauthorized access. These measures can significantly reduce the risk of data breaches so your information remains secure even in the event of a system failure.
3- Watch for phishing emails
Be vigilant regarding phishing emails pretending to be CrowdStrike support. Educate your employees on how to recognize and report suspicious emails so they do not share any sensitive information or authentication details.
Best-practices for effective vendor management
To reduce future risks, it’s crucial to have proper vendor management practices in place.
1- Review vendor access
Start by assessing which vendors have deep access to your systems and data. Implement strict access controls and conduct regular audits of vendor activities to ensure they are complying with your security standards.
2- Diversify technology providers
Avoid relying on a single vendor for critical systems. By using multiple vendors, you can mitigate the impact of any single vendor’s failure and improve the overall resilience of your technology stack.
3- Regularly evaluate third-party vendors
Continuous monitoring and evaluation of vendors can help you stay ahead of potential issues and maintain a secure and reliable network of technology partners.
What’s next for CrowdStrike customers?
It’s hard to say what to do next. There’s a workaround, but it’s not scalable because it needs to be done manually for each system. Doing so could take hours or more to fix.
According to Adam Harrison, managing director at FTI Cybersecurity, the issue is tough to resolve once systems enter a reboot loop. “Manual fixes will take time for system admins to apply, and CrowdStrike can’t push a new update remotely. Each system needs manual intervention.”
Looking ahead, this incident serves as a distinct reminder of the delicate balance in cybersecurity. Even the most well-meaning updates can have unexpected consequences.
For now, the priority is clear: fix the systems, support the customers, and learn from the experience to prevent future mishaps.