we_are_coded.by CODE · The world, decoded
БГ
GitHub

GitHub explained the 17 August outage: capacity did not keep up with its own growth

GitHub BlogInfra

On 20 August GitHub published a post-mortem of the 17 August failure, which held the platform down for seven hours and 47 minutes and took out github.com, authentication, Actions, APIs and Copilot. Per the company the cause was not a change in code or configuration but capacity that did not arrive in time.

In short
  • The 17 August outage lasted 7 hours and 47 minutes and hit github.com, authentication, Actions, APIs, pull requests, issues and Copilot.
  • Per GitHub neither it nor the 6 August failure came from a change in code or configuration. Both are capacity failures.
  • Monthly commits have grown from 1.4 billion in April to 2.9 billion in August.
Checked on21 August 2026Responsible editorTsvetelin IvanovHow we workMethod · Corrections

Seven hours and 47 minutes. That is how long the platform half the world ships software through stayed down, and three days later GitHub's chief technology officer sat down and wrote the post-mortem with a sentence you rarely read in corporate text: if you were trying to ship software that day, we let you down.

Texts like this usually start with "we observed elevated error rates". This one starts with an admission and goes on with numbers. It is worth reading carefully, because the numbers are more interesting than the apology.

The facts: on 20 August 2026 GitHub published a post-mortem of the 17 August outage. It lasted 7 hours and 47 minutes and disrupted github.com, authentication, GitHub Actions, APIs, pull requests and issues, and Copilot. Per GitHub the traffic reached a new peak and a critical infrastructure component in the Central US data centre failed to scale with it; the pressure on capacity spread through the systems and took authentication down. The recovery went in stages, and part of the Copilot services came back last: the errors in them triggered a retry loop on the client side which pushed traffic up exactly while the machine was coming back. The company states that neither this outage nor the Actions failure of 6 August came from a change in code or configuration, and that both are capacity failures. Per its own data monthly commits have grown from 1.4 billion in April to 2.9 billion.

Where the heavy part of this post-mortem is

A capacity outage is the more uncomfortable of the two possibilities. Broken code gets rolled back. A bad configuration is fixed in minutes, sometimes by one person with one key. Capacity running out means the system worked exactly as it was built, only the world came out bigger than expected.

I have had nights where more people turned up than we had counted on paper, and right at the entrance it was clear the problem would not be solved with more effort inside. A narrow entrance does not widen during the event.

Twice the commits in four months. The growth everybody celebrates is the same one that cracks the structure.

And here comes the part few people talk about. A share of that growth is agentic. Code written with a machine next to you gets committed more often and in smaller pieces, and every piece wants authentication, a check, a run. I have not seen GitHub write it that way in the text. This is my reading and I say so plainly.

The retry loop

Here is the detail I would pin on the wall. The Copilot client got an error and tried again. And again. Thousands of installations around the world were doing the same thing at the same time and poured extra traffic onto a system that was just getting back on its feet.

This is a classic of the craft and it still bites. The first of GitHub's two immediate measures is aimed exactly at it: limits, budgets and variable waits on retries between services. The second is a review of the processor and memory alarms, which were low priority and were not calling loudly enough.

If you are writing something that talks to somebody else's API, this is the free lesson of the day. Check what your client does when the other side starts returning errors.

What they promise

More than 3 million processor cores and 120 petabytes of fast storage have already been added, and Azure today carries about 58 percent of the platform's load against 12 percent in May. All of that is GitHub data about GitHub itself, so I take it as a claim and wait to see it verified under pressure. They also promise to isolate the critical systems and remove the shared dependencies between them. Until then we have a company that admitted in public it was late to scale, and wrote down in numbers how late. That is rare and worth noting. I will believe it when the next traffic peak passes.

For you the practical part is one thing and it is uncomfortable. If your development, your delivery and your authentication all go through the same place, your three providers are actually one. Either you know what you do on the day it is not there, or you will be making it up live.

The visual is generated code art. No third-party images.
Follow usFacebookLinkedIn
Official primary sources
→GitHub Blog: The August 17 outage, and the work ahead, Vlad Fedorov (20.08.2026)
Original: https://wearecoded.com/en/articles/github-outage-aug-17-postmortem.html
ShareFacebookXLinkedInTelegramWhatsApp
← Back to all news