CybersecurityJuly 24, 2026· via BleepingComputer

Microsoft 365 outage traced to automated maintenance bug

Microsoft 365 outage traced to automated maintenance bug

Image : BleepingComputer

A routine automated maintenance request at Microsoft spiraled into a continent-wide service blackout Thursday, leaving millions of Microsoft 365 and Azure customers without access for several hours. The company confirmed the incident stemmed from a bug in its network maintenance system, which incorrectly removed IP routes from more devices than intended and disrupted core cloud services.

How the routing error unfolded

Microsoft’s automated network maintenance tool is designed to streamline updates by pushing out routing changes across its global data-centre fabric. On Thursday, however, the system misfired, stripping valid IP routes from a wider set of devices than it should have. Without those routes, traffic destined for Microsoft 365 and Azure services could not reach its destination, triggering widespread timeouts and service degradation. The company said the issue was contained within hours, but the ripple effects lasted well into the evening for some users.

Impact on users and workloads

The outage hit businesses running Exchange Online, SharePoint Online and Teams, as well as customers relying on Azure compute and storage. Reports of failed logins, delayed emails and stalled virtual machines surfaced across social media and enterprise monitoring channels. While Microsoft restored services by rolling back the unintended routing changes, the episode underscores how a single misconfigured automation step can cascade into a major incident for cloud-scale platforms.

Why it matters

Automation is supposed to reduce human error, yet when a maintenance script goes wrong the blast radius can be enormous. This outage shows that even the most sophisticated cloud providers must pair automated tooling with rigorous safeguards—such as canary deployments, circuit breakers and instant rollback mechanisms—to prevent a single bad change from taking down services for millions of users. For enterprises, the lesson is clear: validate every automated change as if it were a manual one, and ensure your own disaster-recovery runbooks account for the same kinds of systemic failures.


Source: BleepingComputer. AI-assisted editorial synthesis — TechnoExpress.

Read the original source on BleepingComputer →

← Back to home