The Pulse: Firebase’s global outage & poor responseThe Firebase iOS SDK crashed after a backend change, crashing all apps which used Firebase analytics for 2-6 hours. Google didn’t update the status page, but did offer a postmortem 4 days later.Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics from last week’s issue of The Pulse. Full subscribers received the article below seven days ago. If you’ve been forwarded this email, you can subscribe here. Firebase – built by Google – has had a nasty outage this week with shockingly poor incident management at odds with how Google itself usually deals with high-severity incidents. The outage started on Tuesday (29 Sep) at 5:41pm (PDT), when iOS apps using the Firebase SDK started to crash upon first opening; every iOS app that uses the Firebase SDK with analytics enabled was affected in this way. Developers of affected apps opened a GitHub ticket, in the absence of much else to do. On the ticket, the message “it’s crashing for me too!” was oft-repeated. Devs reporting their apps crashing. Source: GitHub 6:51pm (PDT): acknowledgement. An hour and ten minutes after the crashes started, an engineer on the Firebase team acknowledged that they were aware of the outage. Just over an hour into the incident, the Firebase team became aware of the outage. Source: GitHub It’s unclear if the Firebase team was alerted via this ticket with 100+ comments by devs, or if Google’s own monitoring tool showed the issue. I asked Google/Firebase two days ago and haven’t had a response. Not having anything better to do than wait for Google to resolve the issue, the memes began: Memes while waiting More memes Others attempted to help the Firebase team by pinpointing the potential issue. Indeed, before a Google engineer acknowledged the incident, an external developer found the root cause at 6:37pm PDT; it was a zero-length entry that was crashing the SDK: Given the flags are shipped by the backend, the offending change was a backend one, and the easiest resolution would be to roll it back, which the community practically begged Google to do: Frustrating: Understanding the problem and how to solve it, but nothing to do but post. Source: GitHub Here’s a neat summary of the incident from another dev: |