Hoppa till innehåll

2022-09-07: DEGRADED PERFORMANCE AND POSSIBLE TIMEOUTS ON SMVS

Description

At 12:15 it was first noticed that monitoring tools showed failed requests. Investigation was performed by Solidsoft Reply and it showed that all national systems hosted in Microsoft Azure environment was impacted. A call was logged with Microsoft to investigate possible Microsoft issues and at 14:20 Microsoft confirmed an Azure CosmosDB problem in Northern Europe that caused degraded performance and possible timeouts. Microsoft was resolving the problem by applying their mitigations. At 19:30 all monitoring tools showed successful requests and SMVS was back to full service and performance.

Timeline of incident (CEST)

  • At 12:15: The incident was first identified due to monitoring tools showed failed requests. Investigation was immediately started, and analysis showed that all national systems hosted in Microsoft Azure was impacted. Decreased CosmosDB performance could be the cause and further investigation was performed.
  • At 13:30: A call was logged with Microsoft to investigate possible Microsoft issues.
  • At 14:20: Microsoft confirmed an Azure CosmosDB problem in Northern Europe that may result in timeouts.
  • At 14:53: Continued to work with Microsoft on resolving the problem, and an improvement was noted in the latency of the CosmosDB in Northern Europe, however end user requests was still intermittent on whether they timeout or not.
  • At 17:00: Gradual improvements in different services for each hour going forward as Microsoft continued to apply their mitigations.
  • At 19:30: All monitoring showed successful requests and national systems were stable and operating in full service.