On 20 August GitHub published a post-mortem of the 17 August failure, which held the platform down for seven hours and 47 minutes and took out github.com, authentication, Actions, APIs and Copilot. Per the company the cause was not a change in code or configuration but capacity that did not arrive in time.
- The 17 August outage lasted 7 hours and 47 minutes and hit github.com, authentication, Actions, APIs, pull requests, issues and Copilot.
- Per GitHub neither it nor the 6 August failure came from a change in code or configuration. Both are capacity failures.
- Monthly commits have grown from 1.4 billion in April to 2.9 billion in August.
Seven hours and 47 minutes. That is how long the platform half the world ships software through stayed down, and three days later GitHub's chief technology officer sat down and wrote the post-mortem with a sentence you rarely read in corporate text: if you were trying to ship software that day, we let you down.
Texts like this usually start with "we observed elevated error rates". This one starts with an admission and goes on with numbers. It is worth reading carefully, because the numbers are more interesting than the apology.
Where the heavy part of this post-mortem is
A capacity outage is the more uncomfortable of the two possibilities. Broken code gets rolled back. A bad configuration is fixed in minutes, sometimes by one person with one key. Capacity running out means the system worked exactly as it was built, only the world came out bigger than expected.
I have had nights where more people turned up than we had counted on paper, and right at the entrance it was clear the problem would not be solved with more effort inside. A narrow entrance does not widen during the event.
And here comes the part few people talk about. A share of that growth is agentic. Code written with a machine next to you gets committed more often and in smaller pieces, and every piece wants authentication, a check, a run. I have not seen GitHub write it that way in the text. This is my reading and I say so plainly.
The retry loop
Here is the detail I would pin on the wall. The Copilot client got an error and tried again. And again. Thousands of installations around the world were doing the same thing at the same time and poured extra traffic onto a system that was just getting back on its feet.
This is a classic of the craft and it still bites. The first of GitHub's two immediate measures is aimed exactly at it: limits, budgets and variable waits on retries between services. The second is a review of the processor and memory alarms, which were low priority and were not calling loudly enough.
If you are writing something that talks to somebody else's API, this is the free lesson of the day. Check what your client does when the other side starts returning errors.
What they promise
More than 3 million processor cores and 120 petabytes of fast storage have already been added, and Azure today carries about 58 percent of the platform's load against 12 percent in May. All of that is GitHub data about GitHub itself, so I take it as a claim and wait to see it verified under pressure. They also promise to isolate the critical systems and remove the shared dependencies between them. Until then we have a company that admitted in public it was late to scale, and wrote down in numbers how late. That is rare and worth noting. I will believe it when the next traffic peak passes.
For you the practical part is one thing and it is uncomfortable. If your development, your delivery and your authentication all go through the same place, your three providers are actually one. Either you know what you do on the day it is not there, or you will be making it up live.