Field note No. 02 · The shiny object problem

The Fix Was Already in the Building

Buying a platform is the organizational equivalent of a goalkeeper diving. It feels like the responsible move, the research says standing still was better, and the fix you needed was already being run by someone three levels down.

Ben Siegel9 min readStrategy & Experience

Buying is the organizational equivalent of diving. Standing still is usually better, and it always feels worse.

If you've watched your organization approve a large platform purchase while an obvious, cheaper fix sat one floor down going unexamined, you saw that correctly. And if you were the one who approved it, the research is on your side more than you'd expect. What you did is the most studied decision in behavioral economics, and almost everyone does it.

Start with goalkeepers, because the evidence there is unusually clean.

Why buying feels safer than looking

Michael Bar-Eli and colleagues analyzed 286 penalty kicks from top leagues and championships worldwide, published in the Journal of Economic Psychology in 2007. Given how kicks are actually distributed, the best thing a goalkeeper can do is stay in the center of the goal. Goalkeepers almost never do. They dive left or right, nearly every time.

These are elite professionals with enormous incentive to get it right. The researchers' explanation draws on norm theory, from Kahneman and Miller in 1986: because diving is what goalkeepers are expected to do, a goal conceded while standing still feels far worse than the same goal conceded mid-dive. The outcome is identical. The blame isn't. So they dive, and they know they dive, which a follow-up survey of 32 professional goalkeepers confirmed.

That's action bias, and a boardroom is a better habitat for it than a penalty box.

The organizational version has been measured too. Florian Artinger and Gerd Gigerenzer, working out of the Max Planck Institute for Human Development, surveyed 950 managers across every level of a public-sector organization and published the results in Business Research in 2018. They asked about the ten most important decisions each manager had made in the previous twelve months. Around 80% reported that at least one of those decisions had been defensive, meaning they'd knowingly chosen an option that was not the best one for the organization, because it was the safest one for them. On average, about a quarter of those top decisions weren't in the organization's interest.

The authors' finding about causes is the part worth sitting with. The main drivers were psychological rather than technical, and the largest was how the organization treats failure. Where a failure produces a search for someone to blame rather than a search for the cause, defensive decisions rise.

Put the two together and you have a complete account of why a platform gets bought before anyone has diagnosed anything. Diagnosis is standing in the center of the goal. It's the better play and it offers no cover. A signed contract with a vendor's name on it is a dive: visible, conventional, and defensible in the meeting where it goes wrong. Nobody gets fired for approving the platform. That sentence is usually said as a joke about procurement. It's actually a precise description of a measured phenomenon, and it's costing your organization about a quarter of its most important decisions.

None of that makes anyone in the chain foolish. It makes them rational inside the incentives they're standing in, which is a different problem with a different fix.

The number that frightened you into it was probably never measured

Action bias needs a trigger, and in this market the trigger is almost always a statistic.

Seventy percent of digital transformations fail. You've seen it in a vendor deck, a board pre-read, a LinkedIn post this week. Mark Hughes went looking for the empirical basis of the broader claim that 70% of change efforts fail, traced it back through five published sources, and found nothing underneath any of them. The figure circulates because it's repeated, not because it was measured.

The adjacent numbers say narrower things than the people quoting them believe. McKinsey's 2018 global survey found 16% of respondents said their digital transformations both improved performance and sustained the improvement. That's a survey of self-reported perceptions, and it isn't a failure rate. BCG published a 70% figure in 2020 whose own breakdown records 30% as fully successful and 44% as creating some value while missing targets. Falling short of an objective is not failing, unless you're selling something.

The 2025 version is the MIT NANDA report, compressed into the headline that 95% of enterprise generative-AI pilots fail. It rests on 150 interviews, 350 survey responses, and 300 public deployments, and critics noted within days that the method never accounted for efficiency gains, cost avoidance, or churn reduction. Most people repeating the number haven't read the method. I'd include myself in that until I went and read it.

The argument here isn't that these programs succeed. It's narrower and more uncomfortable: the firm that sells you the frightening number is usually the firm that sells you the cure. That's not a conspiracy, it's a business model, and it works precisely because it lands on a decision-maker already primed to dive. A leadership team frightened by a statistic nobody validated will approve a solution nobody validated, because the two documents arrive stapled together.

If that's happened in your organization, it isn't evidence that your leadership is careless. It's evidence that the mechanism works as designed on careful people.

What a diagnosis finds that a deck doesn't

Here's where I can tell you what this looks like from inside the work rather than from the literature.

In 2019 I spent six months in the global finance organization of a large pharmaceutical company, working out why its master-data request process kept breaking. Dozens of data domains, five international sites, four operating sectors, and the friction landed in the most sensitive moments of the fiscal calendar. By the time I arrived, the answer had already been bought. A Big Four firm had been engaged to architect new tools and infrastructure, sold on failure statistics, approved by the board.

Everyone complained about the request form, so the form looked like the problem, and much of the proposed build pointed straight at it. The research said otherwise. Requesters could complete 5 of the 30 required fields unaided. The other 25 needed the master data management (MDM) team, every time. Submission error rates ran between 30% and 50% across every site and every domain, and roughly 30% of North American requests were rejected and returned, then resubmitted, often through an entirely different channel, because five or more parallel submission channels coexisted with no shared owner.

Formal form. Email. Attached spreadsheet. Photocopy. Screenshot.

The form wasn't the bottleneck. The form was the receipt at the end of a process nobody owned end to end. A better form, or a new platform underneath it, would have printed a nicer receipt.

Then we went to Manila.

The Southeast Asia and North America MDM team worked out of an office where one local mapper had built her own workaround using SAP IDocs, the standard records that carry data between enterprise systems. It was cutting errors inside her team by roughly 70%. Nobody had asked her to build it. Nobody outside her team knew it existed. And the operating model around her had no mechanism by which a good local fix could become a global pattern, so it stayed where it was.

The fix was already in the building. The organization had no way to notice, and it was about to spend a great deal of money on infrastructure that wouldn't have found her either.

A tool can work perfectly and still not help anyone

There's a reason the platform would have missed her, and it's older than any of this.

Design thinking taught a generation to ask about desirability, feasibility, and viability. Useful, and incomplete. The two questions that predict whether a tool survives contact with an organization trace back to Fred Davis's technology acceptance work in 1989, and they're simpler.

Efficacy: does this thing work, under conditions we can verify?

Utility: does it help this specific person, inside the actual shape of their Tuesday?

A tool can pass the first and fail the second. Almost every failed rollout lives in that gap. The vendor demonstrates efficacy, because efficacy is demonstrable in a controlled room. Utility can only be observed where the work happens, which is the one place the evaluation never goes.

Nielsen Norman Group put a fine point on this in 2024. Maria Rosala and Kate Moran took the Synthetic Users platform, ran it against three studies NN/g had already run with real humans, and compared. Real and synthetic participants gave markedly different answers, and the synthetic responses were flat next to what actual people produce, because human behavior is context-dependent in ways a model has no access to.

Their conclusion is the most useful sentence written about enterprise AI adoption, and it isn't really about AI: synthetic users don't replace research, they replace the guilt of not doing research.

That generalizes further than they took it. The dashboard replaces the guilt of not knowing the numbers. The platform replaces the guilt of not owning the process. The pilot replaces the guilt of not having a strategy. In each case an organization buys an artifact that produces the feeling of the thing, and the feeling is cheaper, faster, and much easier to present at a steering committee than the thing.

Rosala and Moran are careful about where the tool earns its keep, which is the honest version of this argument. Rosala used synthetic users to pressure-test a discussion guide in a domain she didn't know well, before going to talk to real clients. That's scaffolding for the work rather than a substitute for it.

What shipped before anyone bought anything

The engagement produced a future-state operating model: five parallel channels consolidated into one workflow, a named accountable owner at every step, and the Manila mapper's workaround promoted from a local fix to a global pattern. That went into a conceptual design reviewed at CFO-designee level.

The part worth pointing a board at is smaller. Alongside the future state we shipped a quick-wins set that relieved operational pressure inside the first quarter, before any technology procurement was decided. No platform, no license, no integration timeline. Relief in Q1, from understanding the process well enough to name where it was leaking.

The parallel enterprise performance management workstream found the same shape from the other end. About 90% of the long-term planning workload was going into manual data cleaning rather than planning, with finance professionals operating in practice as extract-transform-load engineers, hand-preparing data no system was moving for them. We set a target of taking the gross-to-net close cycle from roughly three days to roughly half a day, and the provenance of that number is the argument: it wasn't a forecast and it wasn't a vendor projection. It was a threshold drawn from a peer finance function already running at that speed inside the same company. An internal baseline beats an external benchmark every time the work has to defend itself in front of an executive committee.

None of it required a new tool. All of it required going to Manila.

Before you sign

New tools aren't the enemy. I'd rather work with a well-chosen platform than without one, and the capabilities available this year are real in a way the 2019 versions weren't. But a tool is a bet that you already understand the problem, and buying one before the diagnostic work isn't a shortcut to understanding. It's a decision to stop looking, and it arrives with a timeline attached so the stopping feels like progress.

There are smaller versions of the alternative, and any of them beats none. The full version is a diagnosis before the procurement decision. A shorter one is a week spent finding out who has already built a workaround. The smallest is a single sentence, written down before the vendor arrives, naming what you want the tool to fix. If you can't write that sentence, the purchase isn't ready, and you've learned that for free.

Somewhere in your organization there's a version of the Manila mapper. She has already solved a piece of this. She isn't in the vendor evaluation, she isn't on the steering committee, and nobody has asked her.

Find her before you sign anything.


Key takeaways

Buying a tool before the diagnostic work is a decision to stop looking. The research explains why it feels safe. Find the fix that already exists in the building first.

Concepts to name

  • Action bias (Bar-Eli et al., 2007). When the norm is to act, inaction feels worse than an identical bad outcome that followed action. So people act, even where standing still is the better play.
  • Defensive decision-making (Artinger & Gigerenzer, 2018). Knowingly choosing the option that is second-best for the organization because it is safest for the decision-maker. Driven by how the organization treats failure.
  • The fear-number business model. The firm that sells you the frightening statistic usually sells you the cure. The two documents arrive stapled together.
  • Efficacy vs. utility (Davis, 1989). Efficacy is "does it work under conditions we can verify." Utility is "does it help this person inside the shape of their actual day." Tools pass the first and fail the second.
  • Replaces the guilt, not the work. An artifact that produces the feeling of the thing is cheaper to present than the thing itself.

The numbers

  • 80%. Share of 950 managers who reported that at least one of their ten most important decisions in the past year was defensive, meaning not the best option for the organization (Artinger & Gigerenzer, Business Research, 2018). About a quarter of top decisions were.
  • 286 penalty kicks. The sample in which staying in the goal's center was the better strategy and goalkeepers dived anyway (Bar-Eli et al., Journal of Economic Psychology, 2007).
  • 70%. The "70% of transformations fail" figure has no validated empirical basis. Hughes traced it through five published sources and found nothing underneath any of them.
  • 16%. Share of respondents in McKinsey's 2018 survey who said their transformation both improved performance and sustained the gain. A far narrower claim, and a self-reported one.

Techniques

  • Before you buy, ask both questions: does it work (efficacy), and will it help this specific person (utility)?
  • Go find the local workaround already solving part of the problem. Someone on the front line usually built one.
  • Write down what you want the tool to fix before the vendor arrives. If you can't, you aren't ready to buy.
  • Check whether the number driving the decision is a measured rate, a self-reported survey, or a figure that only circulates.

Further reading

  • Bar-Eli, M., Azar, O., Ritov, I., Keidar-Levin, Y., & Schein, G. (2007). "Action Bias Among Elite Soccer Goalkeepers: The Case of Penalty Kicks." Journal of Economic Psychology, 28(5).
  • Artinger, F. & Gigerenzer, G. (2018). "C.Y.A.: frequency and causes of defensive decisions in public administration." Business Research.
  • Davis, F. (1989). Technology acceptance model. MIS Quarterly.
  • Rosala, M. & Moran, K. (2024). "Synthetic Users." Nielsen Norman Group.

Sources

  • Bar-Eli, M., Azar, O. H., Ritov, I., Keidar-Levin, Y., & Schein, G. (2007). "Action bias among elite soccer goalkeepers: The case of penalty kicks." Journal of Economic Psychology, 28(5), 606 to 621. 286 kicks analyzed; a supporting survey of 32 professional goalkeepers.
  • Kahneman, D. & Miller, D. (1986). Norm theory, the account Bar-Eli and colleagues draw on for why inaction feels worse.
  • Artinger, F. M. & Gigerenzer, G. (2018). "C.Y.A.: frequency and causes of defensive decisions in public administration." Business Research. Max Planck Institute for Human Development; 950 managers across all hierarchy levels of a public-sector organization.
  • Hughes, M. Tracing of the "70% of change fails" claim through five published sources, finding no empirical basis.
  • McKinsey (2018), global digital transformation survey. The 16% figure. Self-reported perceptions, not a measured failure rate.
  • Boston Consulting Group (2020), digital transformation success-rate breakdown: 30% fully successful, 44% creating some value while missing targets.
  • MIT NANDA report (2025). The 95% headline; 150 interviews, 350 survey responses, 300 public deployments.
  • Davis, F. (1989). Technology acceptance model. MIS Quarterly.
  • Rosala, M. & Moran, K. (2024). "Synthetic Users." Nielsen Norman Group.
  • Figures from the author's engagement with a global pharmaceutical company's finance organization (2019 and 2020): 5 of 30 fields completable unaided, 30% to 50% submission error rate, ~30% North American rejection rate, five or more parallel channels, the Manila mapper's ~70% local error reduction, ~90% of long-term planning workload spent on manual data cleaning, and a gross-to-net close-cycle threshold of roughly 3 days to 0.5 days drawn from an internal peer baseline.