Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home Developer AI & Open-Source Ecosystem

GitHub availability report: Might 2026

Future News 24 by Future News 24
June 12, 2026
in Developer AI & Open-Source Ecosystem
0 0
0
GitHub availability report: Might 2026
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


In March and April we shared updates on GitHub’s availability and infrastructure investments. As that work continues and we method some main milestones, we needed to start out sharing extra common updates in our month-to-month availability reviews.  So earlier than we dive into incidents from Might, right here’s how we’re monitoring with our ongoing work to make GitHub extra dependable.  

Our progress in making GitHub extra resilient

The quick model: GitHub’s site visitors is rising quickly, pushed largely by AI-assisted and agentic improvement workflows, and we’ve been reworking our infrastructure to maintain up with it. Meaning shifting to Azure for elastic capability, breaking our monolith aside into remoted providers, and eliminating the shared failure factors which have pushed previous incidents. 

Right here’s the place we stand. We’re now serving 40% of monolith site visitors from Azure (up from 8% in February), with Git site visitors at 30% and repository replication at 99%. We’ve greater than doubled our efficient capability in 4 months. On the similar time, we’re finishing the isolation of our major database cluster: splitting customers, authentication, and authorization into impartial domains in order that an issue in a single can now not cascade throughout the platform. Our new customers service is absolutely minimize over, and dealing with double the site visitors at considerably decrease database value. Stateless authentication tokens are additionally rolling out, eliminating per-request database lookups that amplified strain throughout site visitors spikes. 

We’re making structural adjustments that completely take away failure modes. We acknowledge that we have work to do, however we’re dedicated to getting it performed and making GitHub dependable when and the place you want it. The precept guiding our determination is easy: availability, then capability, then options. 

Thanks to your partnership as we hold constructing GitHub’s reliability and resilience. 

In Might, we skilled 9 incidents that resulted in degraded efficiency throughout GitHub providers.

Might 04 15:45 UTC (lasting 55 minutes)

On Might 4, 2026, between 15:34 and 16:40 UTC, github.com skilled a service disruption that produced elevated latency and an elevated price of request failures throughout a broad set of customer-facing providers. Whole buyer influence lasted roughly one hour and 6 minutes. 

Essentially the most considerably impacted service was pull requests, which was statused Purple at some stage in peak influence. Points, actions, webhooks, and Git operations skilled elevated latency and intermittent errors. Quite a few dependent providers—together with Codespaces, Pages, Packages, OAuth and GitHub Apps, Market, and Copilot—additionally noticed various levels of degraded efficiency because of shared information dependencies. At peak, roughly 1.3% of requests returned a 5xx response, averaging round 0.46% throughout the period of the incident. 

The disruption was triggered by a routine on-line schema migration working towards a big, heavily-accessed database desk. The migration had been progressing with out difficulty for a number of hours, however as site visitors ramped up towards the weekly peak, the mixed load from the migration and regular manufacturing site visitors saturated database connection capability. This produced question competition on a major database and cascading timeouts throughout providers that rely upon it. 

The incident was detected inside roughly three minutes of the primary indicators of influence by way of a mix of automated monitoring and on-call statement. As soon as the contributing migration was recognized, it was paused, and dependent providers recovered shortly thereafter. Time to mitigation was roughly 33 minutes, and full decision adopted roughly half-hour later 

As follow-up, we’re implementing a number of enhancements to scale back the probability and blast radius of an identical occasion. Migrations towards massive, high-traffic tables will likely be extra tightly aligned with low-traffic home windows and can use dynamic throttling that adapts to dwell cluster load. We’re including automated circuit breakers that may pause in-flight migrations when latency or connection utilization on the underlying database crosses secure thresholds, and we’re extending our monitoring in order that migration-induced strain (i.e., write price, lock time, and connection saturation) triggers alerts earlier than buyer influence happens. In parallel, we’re reviewing connection-pool capability to make sure satisfactory headroom is maintained whereas migrations are working.  

Might 05 13:37 UTC (lasting 3 hours and 49 minutes)

Might 06 07:19 UTC (lasting 2 hours and 25 minutes)

On Might 5 and Might 6, GitHub Actions was degraded throughout two associated incidents affecting hosted runners. The 2 occasions have been linked: remediation work carried out after the Might 5 incident launched the configuration difficulty that triggered the Might 6 incident. 

On Might 5, 2026, from 13:22 to 17:05 UTC, GitHub Actions hosted runners within the East US area have been degraded. Roughly 13.5% of jobs requesting a normal runner failed and ~16% of requested bigger runners with personal networking pinned to East US failed or have been delayed by greater than 5 minutes. Copilot code evaluate requests have been additionally impacted. Roughly 8,500 code evaluate requests timed out throughout this window. Affected customers noticed an error touch upon their pull requests and have been capable of retry by rerequesting a evaluate. Most runner requests have been picked up by different areas mechanically, however a portion of requests nonetheless routing to East US have been impacted. 

This was triggered by a scale-up operation for hosted runner VMs within the East US area. This can be a common operation, however the VM create load hit an inner price restrict when VM creates pull photos from storage. Present backoff logic was not triggered due to the response code returned on this case. The speed limiting and VM creation failures have been mitigated by decreasing load to permit for restoration and permitting queued work to be processed. By 15:34 UTC, queued and failed job assignments have been largely mitigated, with lower than 0.5% of runner assignments impacted between 15:34 and full restoration at 17:05. 

On Might 6, 2026, from 06:45 to 09:15 UTC, GitHub Actions Normal Ubuntu hosted runners have been once more degraded, and roughly 17.1% of jobs requesting a normal runner failed. The difficulty was brought on by surprising configuration information launched throughout remediation work for the day gone by’s incident, which blocked new allocations as each day load ramped up. We eliminated the problematic information at 08:51 UTC, permitting allocations to renew and runner swimming pools to cut back up and get better.  

We’re bettering our system’s throttling habits when limits happen, bettering our controls to extra rapidly mitigate related conditions sooner or later, and reviewing all limits end-to-end for related operations. As well as, we’re updating the filter logic for this allocation information to be resilient to irregular information shapes and bettering monitoring to alert when allocations are blocked, permitting the workforce to reply earlier than buyer influence begins. 

Might 06 11:21 UTC (lasting 38 minutes)

On Might 6, 2026, between 11:02 and 11:13 UTC, customers have been unable to start out or view Copilot cloud agent or distant periods. Throughout this time, all requests to the session API returned errors, stopping customers from creating new periods or viewing current ones. The difficulty was brought on by a configuration change to the service’s community routing that inadvertently eliminated the ingress path for the service. The workforce reverted the change at 11:13 UTC which restored service. The incident remained open till 11:59 UTC whereas the workforce verified full restoration. We’re taking steps to enhance our deployment validation course of to stop related configuration adjustments from impacting manufacturing site visitors sooner or later. 

Might 06 15:25 UTC (lasting 3 hours and 39 minutes)

On Might 6, 2026, between 15:12 and 19:02 UTC, creation of latest pull request evaluate threads on github.com failed. This included new line feedback and file feedback on pull requests. Present pull requests and beforehand created feedback have been unaffected. 

This incident was brought on by a 32-bit integer key reaching its most worth in a Vitess lookup desk used throughout pull request thread creation. The first desk was migrated to a 64-bit integer key, however the Vitess lookup desk remained 32-bit. As soon as the values within the major desk handed the obtainable 32-bit ID house makes an attempt to create new evaluate threads started failing, leading to a close to 100% failure price for brand new thread creation requests. We mitigated the difficulty by updating the impacted lookup desk definitions throughout all shards to make use of 64-bit integer column varieties, growing the obtainable ID vary, and restoring regular operation. Service was absolutely restored as soon as the schema adjustments competed globally. 

To assist forestall related incidents, we’re increasing current monitoring of database columns to incorporate Vitess lookup tables and enabling earlier detection of any tables approaching a column dimension restrict.    

Might 07 05:02 UTC (lasting 1 hour and 54 minutes)

On Might 7, 2026, between 04:12 and 06:13 UTC, Copilot cloud agent and Copilot code evaluate agent periods for pull requests have been delayed or did not begin.   

Through the influence window, new Copilot coding agent periods triggered by pull requests weren’t created, whereas evaluate agent periods dropped by ~50% from the baseline.   

The difficulty was brought on by follow-up restoration work from a separate pull request incident. As a part of that restoration, we ran a big database migration, which precipitated replication delays on a number of duplicate hosts. 

Though these replicas weren’t serving person site visitors, our safeguards accurately handled the elevated replication lag as a sign to decelerate writes to the affected database cluster. Because of this, some pull request background processing was briefly delayed. That processing is answerable for sending the inner occasions that Copilot brokers use to start work, so affected brokers didn’t begin till the database replicas caught up. 

The system recovered as soon as replication lag returned to regular and pull request processing resumed. We’re making crucial background processing extra resilient to replication lag so it degrades gracefully and retains important occasions flowing below pressure, decreasing the prospect of comparable secondary influence throughout future restoration work. 

Might 15 08:13 UTC (lasting 35 minutes)

On Might 15, 2026, from 07:43 to 08:48 UTC, GitHub Actions skilled a degradation that precipitated workflow runs to fail or expertise delayed begins for a subset of consumers. The incident was triggered by a deliberate failover of supporting infrastructure utilized by GitHub Actions. Throughout that operation, an automatic service discovery replace didn’t propagate accurately, which precipitated site visitors to be routed incorrectly and elevated request timeouts in a core dependency for workflow orchestration. 

At peak influence, 42% of Actions runs failed. Downstream providers that rely upon Actions workflow execution have been additionally impacted, together with GitHub Pages and Copilot cloud providers. At 08:12 UTC, responders manually corrected the service discovery routing difficulty. Timeout and failure charges recovered shortly after, and we continued monitoring till full stabilization was confirmed throughout all affected providers. The incident was marked resolved at 08:48 UTC. 

To forestall recurrence, we’re taking a number of steps. First, we’re implementing failover guardrails that validate service discovery state earlier than finishing failover operations. Second, we’re strengthening verification checks. Lastly, we’re bettering dependency resilience to scale back timeout cascades throughout infrastructure occasions. 

Might 26 10:57 UTC (lasting 2 hours and 21 minutes)

On Might 26, 2026, between 10:40 and 12:56 UTC, GitHub Actions jobs have been degraded. From 10:40 to 12:16 UTC, all newly queued runs did not begin. From 12:16 to 12:56 UTC, runs that required downloading actions for his or her workflows continued to fail. GitHub Pages, Copilot code evaluate, Copilot coding agent, Octoshift, and GitHub Enterprise Importer have been additionally impacted because of their dependency on GitHub Actions. 

This was brought on by our automated account evaluate system incorrectly suspending the service account utilized by GitHub Actions to authenticate workflow runs and obtain actions. 

We mitigated by restoring the account at 12:16 UTC, marking it exempt from additional automated evaluate at 12:20 UTC, and redeploying a associated service at 12:48 UTC to flush cached account state. Full restoration was confirmed at 12:56 UTC. 

Throughout this incident, a small variety of points, pull requests, feedback, and discussions have been marked as hidden when the service account was disabled. No information was misplaced. All content material hidden due to this incident has been restored and full search index restoration was accomplished. 

To stop a recurrence, we’ve added an allowlist of all service accounts that can’t be suspended by automated techniques, and guaranteeing these protections are enforced persistently throughout all account administration tooling. We’re additionally bettering diagnostic tooling for accounts and decreasing cache propagation delays to shorten time to mitigate related incidents sooner or later. 

Might 28 19:01 UTC (lasting 1 hour and 40 minutes)

On Might 28, 2026, between 18:27 and 20:41 UTC, the GitHub Copilot service was degraded because of a problem with the Responses API of an upstream supplier affecting the GPT-5.2, GPT-5.3-Codex, GPT-5.4, and GPT-5.5 fashions. Requests routed to those fashions by way of the Responses API returned elevated error charges, which additionally affected Copilot coding agent and Copilot code evaluate. No different fashions have been impacted. 

We mitigated the incident by shifting site visitors away from the affected fashions whereas the upstream supplier deployed a repair. 

GitHub is working to enhance automated failover for the affected fashions and strengthen monitoring to stop related incidents sooner or later. 

Observe our standing web page for real-time updates on standing adjustments and post-incident recaps. To be taught extra about what we’re engaged on, try the engineering part on the GitHub Weblog.

Written by

Jakub Oleksy



Source link

Tags: availabilityGitHubreport
Previous Post

An weight problems drug deep-dive, and peptides transfer mainstream

Next Post

Claude Fable is relentlessly proactive

Next Post
Claude Fable is relentlessly proactive

Claude Fable is relentlessly proactive

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb