A client called me not because anything had broken, but because something almost did. Their order-processing system — the one piece of software that decides whether a sale actually ships — had run without a single major incident for six years. Then the one developer who truly understood it, let's call him Diego, because that's not his real name, handed in his two weeks' notice on a Friday afternoon.
Nobody panicked right away. The system still worked. It had worked all week, all month, all year. It would keep working on Monday, too, with or without Diego. That was exactly the problem: "still works" told them nothing about what would happen the first time it needed to change, or break, after he was gone.
I hear some version of "if it still works, why touch it?" from almost every business owner I talk to about modernization, and it's a completely reasonable question. Working software is valuable, and change carries real risk. But the question hides an assumption that's worth examining: that "still works" and "fine" mean the same thing. In my experience, they rarely do.
What "still works" actually measures
"Still works" means the system handled today's transactions the way it handled yesterday's. That's it. It says nothing about whether the technology underneath is still supported, whether one person's knowledge is the only thing standing between "working" and "broken," or what it will cost to make the next necessary change. Those questions live on a completely different axis, and a system can score badly on all of them while still processing every order correctly, every single day, for years.
The gap between running and healthy
Four things tend to separate a system that's merely running from one that's actually healthy, and none of them show up in daily operation until they do:
Unsupported dependencies. A library, a database version, or a framework the system depends on stops receiving security updates. Nothing visibly changes the day that happens. The system runs exactly the same the day after support ends as it did the day before. The risk is silent until a vulnerability surfaces that nobody is going to patch.
Single points of failure. Diego's situation is the human version of this, but it shows up in infrastructure too — one server with no failover, one undocumented integration only one person knows how to fix. The system has no visible weak point, right up until the one piece it depends on disappears.
Mounting change cost. Every system accumulates small shortcuts over years: a workaround here, a field repurposed for something it wasn't designed for there. None of it breaks anything on its own. But each new change has to work around everything that came before it, so the same kind of request that took a day five years ago quietly takes a week now, and nobody budgeted for that drift.
A shrinking hiring pool. Older frameworks and languages still work perfectly well, but fewer developers want to build a career in them. That's not a problem while your current team is in place. It becomes a very expensive problem the day you need to hire, and discover the pool of people who can confidently work in the system is a fraction of what it used to be.
Why this risk stays invisible until it isn't
All four of these share the same trait: they cost nothing on any given ordinary day. A system with an unsupported dependency runs identically to one without it — until the day it doesn't. A system with one irreplaceable expert runs fine — until that person is unavailable. This is exactly why "if it still works, why touch it?" feels like a reasonable question in the moment it's asked: there's no visible symptom to point to yet. The cost of these risks doesn't show up on a timeline until the day it shows up all at once.
A short, honest checklist
None of this means every working system needs an overhaul, and it's not a reason to panic. It's a reason to look, clearly, at a few specific questions:
- Is there more than one person who could fix this system under pressure, this month?
- Are the core dependencies (framework, database, major libraries) still receiving security updates?
- Has a recent "small" change taken noticeably longer than it would have a few years ago?
- If you posted a job today for someone to maintain this system, how long would that search realistically take?
If most of those answers are comfortable, that's genuinely good news — it means the system is running and healthy, not just running. If one or two answers make you pause, that's not an emergency. It's useful information, and the kind of thing a short, focused audit can turn into an actual plan instead of a vague worry.
The reassuring part
A system that still works is an asset, not a liability — the years it's run without failing are real evidence it was built reasonably well. The point of asking whether it's actually healthy, not just running, isn't to talk anyone into a rewrite they don't need. It's to replace a vague unease with a specific, answerable list, so a decision to keep the system as it is — which is very often the right call — is a decision made on purpose, not just a thing that kept not getting looked at.
Diego's employer, for what it's worth, used his notice period well: they spent it documenting what he knew and making sure a second person could step in. The system didn't need a rewrite. It needed exactly that: someone to notice the single point of failure while there was still time to fix it, and not a moment after.