Skip to content
All case studies

Distributed Systems + Platform Architecture

One Coordination Decision Across Three Production Systems

Three production systems could each only run one worker, for what looked like three unrelated reasons. They turned out to be the same problem, so I solved it once and applied it three times, then deliberately left a fourth system alone.

3

production systems unpinned from a single worker

100%

in-flight work survival across restarts, up from none

5

production systems audited and classified

Architecture decisions recorded alongside their rejected alternatives

The problem

Three systems in production could not scale past a single worker process. Locally each looked like its own problem. One had background jobs that would double-run if a second copy started. One held queued work in memory and lost all of it on every deploy. One delivered events only to clients that happened to be connected to the same process that produced them. Treated separately they were three infrastructure projects. None of them had a performance baseline either, so there was no safe way to tell whether any change had actually helped.

What I built

Audited the state each system held in process and sorted it by what kind of coordination it genuinely needed. That reframing did most of the work, because it showed that one shared mechanism covered the real cases and that several things which looked like problems needed nothing done to them at all. Chose the approach against the hardest deployment reality rather than the textbook answer, since one of these systems ships to a customer site with no operator, where every additional running service becomes a permanent support cost instead of a one-time setup. Built the shared mechanism, proved each capability independently, then removed the single-worker restrictions and re-ran the benchmarks to confirm nothing had regressed.

Technical approach

  • Started by classifying in-process state rather than by choosing technology, which is what revealed that three separate scaling projects were one problem wearing three costumes
  • The most useful output of that audit was the list of things needing no coordination at all, since duplicated work that is harmless costs less to leave alone than to synchronize
  • Infrastructure chosen against the hardest deployment target rather than the easiest, because a dependency that is trivial in a managed environment becomes a permanent liability on a box nobody visits
  • Background jobs that must run exactly once are elected to a single owner, with the election supervised so a network interruption produces a predictable handover rather than two owners or none
  • Queued work is made durable before it is acknowledged, so a deploy or a crash no longer silently drops whatever happened to be in flight
  • One capability was deliberately kept out of the shared mechanism. High-frequency streaming has different characteristics, and forcing it through the same channel would have been using the wrong instrument in order to look consistent
  • Shared counters were settled by measuring real traffic rather than by reflex, and the measurement showed the simple answer was sufficient, which avoided standing up infrastructure that would then need maintaining forever
  • Restrictions came off only after each mechanism was proven on its own, with failure behaviour tested by actually killing processes rather than by simulating it

Visuals

State classification across the systems, including what needed nothing
Single-owner election, and what happens when the owner dies
Work lifecycle: durable before acknowledged, reclaimed on failure
Where the pattern was deliberately not applied, and the reasoning