What's New in Rust's "Devblog 80 - Postmortem" Update
Two days after Devblog 80 shipped, Facepunch posted an unusually frank postmortem. Sound pooling caused the freezes, a save bug knocked buildings down, and Garry laid out what went wrong with testing, communication and patch-day pressure.
Two days after Devblog 80 went out on 1 October 2015, Facepunch posted a follow-up that wasn't a patch at all. The Thursday release had gone badly wrong, and rather than quietly hotfixing it, Garry wrote a postmortem laying out each failure and, unusually, what caused it. It remains one of the more candid posts in Rust's devblog history - the summary of it was simply that the patch was broken and they hadn't communicated about it.
Freezing and Crashes
The freezes and crashes traced back to sound pooling, a last-day, untested addition that seemed harmless. It wasn't. Garry called this the usual shape of shipping bugs: you make a change, assume something about it, and the assumption turns out to be wrong. The change was disabled and an update posted, which appeared to fix what most players were hitting.
The Save Bug
The nastiest one was a save bug in the original server release. If you had been running that server for a while and then updated it to the released patch, a lot of your buildings would fall down. Subsequent restarts were fine, but the damage was already done - and losing base progress is about the worst class of bug Rust can ship. Garry took it directly: his mistake, from tinkering on patch day.
The Memory Leak
Reports of a memory leak were widespread, and Facepunch had seen it themselves, though never as extreme as some players described. Garry had watched Rust climb to 6GB of memory and then drop back to 1GB. Unity's memory profiler didn't account for it, and nothing on the native side explained it either. It appeared to have started after the move to Unity 5.2, though that wasn't confirmed. The diagnostic problem was structural: Unity's memory tools only report on the C# side and offered no runtime memory profiling that could flag issues back to the team, which meant Unity Support might need to get involved.
Everything Else
The rest of the bugs were attributed simply to how much the patch touched. A lot changed, so a lot broke. The team was working through them and felt on top of it, but warned to expect a couple more patches. Some performance fixes went out in the hot patches after the main release, with the caveat that they only ever hear about performance when it's bad - hence a poll asking players to rate performance on the current version specifically.
Communication
The patch was broken and nobody was told. Players didn't know Facepunch was working on it, or even that they were aware of the problems. They had posted in a couple of places, but Garry's point was that players shouldn't have to go hunting for that - it should go on Twitter like everything else. The team had been running on the assumption that if something is broken, everyone knows a fix is coming. He accepted that isn't a given, and that when a patch goes this wrong a simple heads-up is worth a lot.
Undertested
The most interesting section was on testing. Facepunch had no internal QA team - Garry considered the concept outdated - and instead relied on players on the dev branch. That didn't work here, for a few compounding reasons. There was no report-back mechanism for players on either the dev branch or main to flag problems. Most of the issues in this update were visible on the dev branch, but nobody was playing it, precisely because of those bugs. The team fixed a batch of problems before the update, but not early enough to get meaningful dev branch testing in. And internal error reporting didn't distinguish between dev and main branches, so dev branch errors were drowned out by the day-to-day volume from main, making automatic reporting useless for this. The conclusion was that they needed to get more people onto the dev branch.
Patch Day Pressure
Garry also explained the schedule logic. Patches go out on a Thursday rather than a Friday specifically because bugs are expected, so the team wants a day to fix them before the weekend. But these weren't the kind of bugs you ship and patch out, and there had been a sense that it might all go wrong before they posted it. His framing of the bind: delay the patch a week and you're a dick because you had two weeks to fix it; ship it and you're a dick because you had two weeks to test it. That pressure is sharpest on a wipe patch, when server owners have already prepared around the date. His conclusion was that the team needed to feel less bound by patch day, and to delay and absorb the criticism when something isn't ready.
The closing note was the honest one: this had happened a couple of times recently, so it needed serious consideration. Being an alpha in active development isn't an excuse to push anything at the main branch - things should stay on dev until they've had a reasonable amount of testing, even if that means patches land on an irregular schedule.