Reliability

Why millions of people tolerate terrible software — but your customers are quietly planning to leave

May 12, 20264 min read
Why millions of people tolerate terrible software — but your customers are quietly planning to leave

Millions of salaried Indians are held captive by the Employees' Provident Fund Organisation (EPFO) website, one of the most widely used, and notoriously unreliable, government portals. Logins fail, forms time out, and transactions get stuck in digital limbo. Yet, users are forced to return, not by loyalty, but by a government monopoly on where their mandatory retirement savings go. You don't get to opt out. This level of unreliability is tolerable only when no alternative exists. But the moment a private, working alternative appears, users would switch in a heartbeat, not for better features, but simply for reliability. This broken website highlights a core truth for every tech company that isn't a government monopoly.

Most software products don't have an EPFO-style monopoly. If your product is unreliable, your customers have options, and they'll take them. Unreliable software is an instrument of Satan, made to torture you in your day-to-day life. The best engineering leaders I've worked with understand this instinctively. Their order of operations is security first, reliability second, new features third. Everyone else has those reversed and wonders why their NPS is bleeding.

Here's the part that most leaders miss: customers only report issues promptly when they believe you'll fix them. Once that trust is gone, they stop reporting and start planning their exit. The silence isn't peace. It's the sound of churn loading.

The cost no one is putting on the spreadsheet

I keep hearing software companies complain about the cost of observability tools. The cost of the vendor. The cost of the SRE headcount. The cost of the platform team.

Nobody is calculating the cost of losing a customer because of reliability issues.

I wish there was an easy way to translate "we lost three enterprise accounts last year because of repeat incidents" into a line item that sits right next to the Datadog bill. Because the conversation about which tool to buy and how much to spend is always secondary. The primary question is: what is it costing you to lose customers to reliability issues? Until leadership can answer that, every observability budget conversation is theater.

There's a second hidden cost, even bigger. When your engineers don't have a reliability safety net, when they can't push a change without flinching, they slow down. They add review cycles. They batch deployments. They build coping mechanisms instead of features. You are paying for innovation velocity that you are not getting, and it doesn't show up anywhere in your dashboards.

The fragmentation trap

In large organizations, "caring about reliability" usually means top leadership has punted the problem to individual departments. Each team picks their own observability tool, whatever someone championed, whatever was on sale, whatever the last principal engineer had a preference for. Everybody has something. Nobody has the same thing.

This is a disaster dressed up as autonomy.

Without a single place for observability data across your company, you cannot enforce high standards centrally. Every department has to rebuild best practices from scratch, and pay for the privilege. And even if you assume every team independently nails it, a heroic assumption, the moment you get an incident that spans two departments, you're stitching together dashboards from three tools at 2am while the customer waits.

The simplest observability tool, uniformly implemented across your company, beats your departments divided across three best-in-class ones. Every time.

Observability infrastructure, OpenTelemetry instances, self-hosted tool backends, the shared schemas, the conventions, needs central ownership. Not central control of every dashboard. Central ownership of the platform that enables everyone else.

A Question for Every Leader

If your organization has a fragmented observability setup across teams, consider this: what's stopping you from consolidating?

Is it the migration cost? The political cost of telling a team their favorite tool is going away? The fact that nobody can quantify what fragmentation is actually costing you, so the status quo always wins the budget meeting?

I suspect it's the last one. Until leadership can translate the silent cost of unreliability into a concrete number, the cost of lost customers, and the cost of engineering velocity; reliability will always be treated as a secondary concern. The number exists; you just haven't been forced to calculate it yet.

Author

Ankesh Khemani

Ankesh Khemani

Founder and CEO

Why millions of people tolerate terrible software — but your customers are quietly planning to leave