Hi, this is Gergely with a bonus, free issue of the Pragmatic Engineer Newsletter. In every issue, I cover Big Tech and startups through the lens of senior engineers and engineering leaders. Today, we cover one out of four topics of last week's The Pulse issue. Full subscribers received the article below seven days ago. If you’ve been forwarded this email, you can subscribe here.
You can no longer watch The Pragmatic Engineer Podcast as video in the Spotify app (only as audio) because I have quit publishing video on that streaming platform. This comes after I decided that reliability takes a back seat within that team – and across much of Spotify. Unlike on other platforms such as YouTube, Apple Podcasts, and Substack, I’ve recently encountered a series of reliability issues around Spotify being unable to process video episodes. Even though I enjoyed a direct link with the Podcasts team there, things haven’t improved.
So from now, I will no longer be publishing video episodes on Spotify. You can find videos of my in-depth chats with guests only on YouTube. Apologies for any inconvenience this change causes! Audio episodes of the podcast can still be found on Spotify via the RSS podcast feed hosted on Substack.
Honestly, the decision to quit the streaming giant wasn’t hard, and I reckon there’s a point here about the risk of deprioritizing reliable operations at major companies in order to push on things like AI adoption, as Spotify seems to be doing.
Some context: for the first two years of The Pragmatic Engineer Podcast, it was published on three podcast platforms:
- Substack’s podcast platform (audio): this is where the “master” RSS feed is served to the likes of Apple Podcasts, the web, Overcast, Pocket Casts, etc
- YouTube (video): video episodes uploaded individually
- Spotify (video + audio): every video episode was uploaded individually and then served as video or audio episodes from the platform.
As someone hosting a podcast, there are good reasons to bother doing three separate uploads:
- Most podcast platforms don’t support video. There will always be a need for a platform that serves the master RSS feed for audio versions while the video ones are elsewhere.
- YouTube doesn’t integrate with anything. YouTube is the leader in video podcast distribution, and uploading there directly makes sense.
- I had a direct line to the Spotify team, which was a big plus. Starting out the podcast, I had the unusual privilege of contact with the podcasts team, thanks to the newsletter gaining a decently-size audience. I was persuaded to take the plunge with them.
For eighteen months, nothing major went wrong. The admin portal for podcast publishers (called ‘Spotify Creators’) was pretty wonky; it gave intermittent errors, and was unable to remember me when I signed in, so, each Wednesday, I’d have to sign in with a code sent to my email to publish an episode.
But overall, things worked, until it all went suddenly downhill…
Unable to publish Spotify podcast episodes 3 out of 5 weeks
From late May, I did not include links to Spotify on new episode announcements because their podcasts product or platform seemingly had outages every time one published on Wednesdays at around 9am PST / 12pm EST / 6pm EU time.
Outage #1 (20 May): podcast publishing broke, my episode would not process on Spotify for 2+ hours. When uploading a video file to Spotify, there’s a processing pipeline that runs to create chunks of the podcast in different video and audio formats. This pipeline appeared to stop running, meaning new episodes were not published.
It was not just the publishing that broke: the Creator portal looked absurd, with NaN% values everywhere, during the outage:

I emailed the Spotify team to alert them about the outage and also complained online. I got a response, confirming the outage and pledging to do better:
“The issue was in one of our podcast publishing metadata pipelines. A small subset of episodes completed normal media processing but then missed a downstream publish update because a newly introduced validation signal was not correctly wired into the logic that wakes up the publishing path. In simpler terms: the episode could become eligible to publish, but the final propagation step was not reliably triggered for that class of episodes.We identified the root cause, deployed a fix, and reprocessed the affected episodes with all-clear called early this morning. We’re also tightening the system so that fields used for publishing eligibility cannot be added without also triggering the relevant downstream updates.
Separately, we’re reviewing how partial creator-impacting publishing delays are surfaced, because even when this is not a broad platform outage, it is still a bad experience for publishers like yourself.
Apologies again that you hit this. It was a real bug, not a wide outage, but it hit some of our most relevant creators.”
Outage #2 (17 June): Spotify down. Four weeks later, when attempting to publish a video episode, all of Spotify went down for many users, including myself.

Spotify does not maintain a status page, so it’s impossible to tell how widespread the outage was. I didn’t include a Spotify link in that week’s announcement either.
Outage #3 (24 June): podcast publishing broke – again. Outage #3 in five weeks; deja vu. This time, it was episode publishing not working, yet again. After waiting two hours for the episode to publish on Spotify, I yet again sent out the announcement with no Spotify link.
I also emailed the Spotify Podcasts team, who confirmed the outage. I said I was considering stopping publishing video episodes, and to switch to audio-only publishing (which means pointing Spotify to my master RSS feed.) I said that an apology was appreciated but it wasn’t enough to make it worth publishing video episodes there.
I also asked for the incident review because I had the feeling that reliability was not all that important on this podcast product. For the first outage I got a vague description of what happened, and promises of improvements that were never done – e.g. during this second outage, there was no improved communications to creators, which I was told would happen, after outage #1.
Internally, Spotify’s team surely conducted an incident review as per usual, so I figured I’d hear back in about two weeks’ time, and assumed a reply would be forthcoming because I’d made clear I was ready to leave Spotify Podcasts if reliability didn’t improve.
No incident review three weeks later, so I quit Spotify
The incident review had never arrived as promised by three weeks later, even though there had been time for it to be completed. It was yet another sign of a platform that has become unreliable. Also, the creator portal occasionally threw up this error:

I checked my Spotify stats: stream plays had been trending downwards unsurprisingly, given the ongoing outages, while the other podcast platforms didn’t show the decline. It made me decide “enough is enough” and to move off Spotify.
Staying on their platform depended on seeing an incident review, but they didn’t prioritize transparency, still had no status page, and nobody had built a feature for episode-processing status like YouTube has had for years. So, I pulled the plug and left:

After I made the switch away from Spotify, the platform’s creators portal became buggier than ever, as in these examples:

Comments disappeared:

… even though other parts of the UI showed dozens of comments:

Episode links directed to 404 pages:

A day or two later, these issues disappeared: I assume no one had tested the flow of moving away from Spotify Podcasts to an RSS feed, and it’s why the experience was so poor.
Incident review finally published, but with a wrong timeline
A few days after offboarding from Spotify, their team published the incident report for outage #3. Reading through it, something did not add up in the timeline:

My email account confirmed that I mailed the Spotify team at around 17:30 about the outage. So, after weeks of creating this report, why did the incident report downplay the fact that customers alerted the team before their own automated alerts fired?I complained to the Podcasts team, and to their credit, the incident report was updated:

I didn’t like how high-level the report is, and how vague the promised improvements were. Specifically, this one:
“During this incident, many creators learned something was wrong from their audiences before they heard anything from us. We are improving our processes and technical capabilities so creators get notified as soon as possible when things aren’t working.”
Overall, I don’t regret the choice to leave, particularly when the focus of Spotify’s leadership is on AI, not reliability.
Does Spotify have “AI psychosis?”
Previously, I used the term “AI psychosis” differently from the usual way of describing when someone starts believing everything an AI model tells them, however outlandish. I applied it to Meta’s rush to develop its own AI model at the cost of the reliability of its profitable business activities. This was based on Instagram’s most embarrassing-ever account takeover incident, which occurred when the team responsible for Instagram’s Trust & Safety was slashed. Soon after, AI-generated, AI-reviewed code caused the hacking of a former US president’s account.
At Spotify, it should have gone the other way. In March, I had the opportunity to meet its Head of Technology & Platforms, Tyson Singer, who said the company puts reliability far ahead of AI adoption, and doesn’t adopt AI for its own sake. So, it was somewhat surprising to read the summary below of a podcast Spotify did with Anthropic:
“Spotify now ships 4,500 production deploys a day, and 73% of PRs are now AI-assisted.Niklas Gustavsson (VP of Engineering at Spotify) keeps 5 to 10 Claude sessions running in tmux, one per git worktree, agents working in the background. All of it inside a 20M+ line monorepo. He expected agents to struggle at that size, but it’s worked well.
Spotify’s migration codemods grew into thousands of lines of edge cases. Code has too much API surface for static rewrites. Early LLMs barely did better. Adding a judge took PR success from ~25% to 80%.
All of this leans on verification, the single most important thing when agents are used and the place most companies underinvest
Spotify rebuilt their test automation around it so engineers can confidently guide and supervise agents, rather than manually execute repetitive tasks.”
It seems to me that all the talk is about usage of AI, and none about reliability, all while Spotify’s platform becomes less reliable than ever, at the same time as the streamer is going all-in on AI; with AI judges and devs running 5-10 parallel Claude sessions.
All things considered, it’s worth asking if Spotify has the corporate variant of “AI psychosis”, whereby the reliability of a successful operation gets torched in the chase for the next big thing by executives. I don’t even think Spotify is all that different from Meta and other companies in this!
Things look bad, based on the quality and reliability degradation of products. Annoyingly, in many cases, customers don’t really have the choice of going elsewhere. My podcast is an exception, as video podcasts on Spotify never truly took off, so quitting the platform wasn’t a big deal. Even so, I’m particularly disappointed that Spotify has prioritized AI usage over reliability. I know some executives there pushed against this, but I feel safe in assuming that they lost that battle.
Value of staying reliable & “sucking less”
Max Kanat-Alexander, distinguished engineer at Capital One, has written about how a software project can become wildly successful just by “sucking less” in his reflections upon the success of the Bugzilla project, (2004-2009):
“All you have to do to succeed in software is to consistently suck less with every release.Nobody would say that Bugzilla 2.18 was awesome, but everybody would say that it sucked less than Bugzilla 2.16 did. Bugzilla 2.20 wasn’t perfect, but without a doubt, it sucked less than Bugzilla 2.18. And then Bugzilla 3.0 fixed a whole lot of sucking in Bugzilla, and it got a whole lot more downloads.
Why is it that this worked?
As long as you consistently suck less with every release, you will retain most of your users. You’re fixing the things that bother them, so there’s no reason for them to switch away. Even if you didn’t fix everything in this release, if you sucked less, your users will have faith that eventually, the things that bother them will be fixed. New users will find your software, and they’ll stick with it too. And in this way, your user count will increase steadily over time.
But what happens if you release frequently, but instead of fixing the things in your software that suck, you just add new features that don’t fix the sucking? Well, eventually the patience of the individual user is going to run out. They’re not going to wait forever for your software to stop sucking.”
Personally, I got tired of Spotify’s Podcasts product continually going in the wrong direction on Max’s scale: the poor reliability, frequent errors on the Creators site, and the sense that they don’t really care about improving existing things.
Read the full issue of last week's The Pulse. The full The Pulse additionally covers:
- Will Kimi K3 trigger US push for closed-source AI models? Moonshot AI’s latest open model, Kimi K3, is on par with Anthropic’s Fable 5. Could it lead to the US government regulating or banning Chinese open models to protect US labs?
- AWS laughs off “heart attack” billing error. AWS customers were billed trillions more than they should have been, due to what was likely a conversion error. But instead of sharing an incident report, AWS saw the funny side.
- Industry pulse. OpenAI’s unreleased model tried to hack HuggingFace to improve its test scores, X took more than a year to develop its new Android app, Google’s new AI model flops, and more.
Subscribe to my weekly newsletter to get articles like this in your inbox. It's a pretty good read - and the #1 software engineering newsletter on Substack.
0 Comments
Log in to join the conversation.No comments yet. Be the first to share your thoughts.