Field note No. 07 · Evidence and measurement

A Green Dashboard Is Not a Working Service

If you've sat in a review where every light was green and known something was wrong, you were right. Accounting researchers have measured the mechanism, and they found the standard fix makes it worse.

Ben Siegel5 min readStrategy & Experience

A dashboard reporting green is a claim, not a fact.

If you've sat through a review where every indicator was green and left with the feeling that something underneath was wrong, that instinct was worth more than the dashboard. You probably didn't say it out loud, because disagreeing with a green board means arguing against the only evidence in the room. That's a reasonable thing to have kept to yourself, and it's also the exact moment the organization loses its last chance to catch the problem cheaply.

What you were sensing has a name, and it's been measured under controlled conditions.

The measured version of what you suspected

Start with the general principle. Charles Goodhart, a British economist, observed in 1975 that when a measure becomes a target it stops being a good measure. Once people are managed against a number, they optimize the number, and the number drifts away from the thing it was meant to represent.

Accounting researchers took that from aphorism to evidence. Jongwoon Choi, Gary Hecht, and William Tayler named the mechanism surrogation: managers lose sight of the strategic construct a measure was built to represent, and begin acting as though the measure is the thing. Not through laziness or bad faith. The metric is concrete, visible, and rewarded, while the goal it stands for is abstract and further from the day's work. Of the two, only the proxy is something a person can act on directly.

Then they found the part that should change how you run a review. In The Accounting Review in 2012, in a paper titled "Lost in Translation," they showed experimentally that tying incentive compensation to a single measure intensifies surrogation. The harder you reward the proxy, the more completely people substitute it for the strategy it was serving. Which means the standard response to a metric that isn't moving, a bigger bonus attached to it, is the intervention that destroys its meaning fastest.

There's a companion finding, and it's the one worth acting on, because it points at a fix rather than a hazard. In the Journal of Accounting Research, the same researchers found that involving people in selecting the strategy reduces surrogation. People who helped choose what the organization was trying to achieve stayed connected to the goal behind the number. People handed a metric to hit did not.

Read those two together and you get an uncomfortable but actionable conclusion: a measure handed down and heavily incentivized is the configuration most likely to hollow itself out. A measure people helped choose, rewarded alongside others, holds its meaning longest.

What the green was covering

Here's what that looks like in a real finance organization.

In 2019 I led the human-centered design research on a finance transformation at a large pharmaceutical company. The finding that mattered came early. Roughly 90% of the long-term planning workload was manual data cleaning. Skilled finance people, hired to plan, were spending most of their week working as extract-transform-load engineers: pulling non-standard data, reconciling it by hand, reshaping it into something a model could accept before any planning could begin.

The dashboards above that work looked fine. Green status. Reports delivered. Cycles closed on time.

Both things were true simultaneously. The service the executives watched was healthy by every measure it tracked. The 90% cost sat one layer below those measures, so it never reached them. Nobody was hiding it, and no one had gamed anything. The measurement system had simply never been designed to capture that layer of work, and the proxy was fully satisfied while the goal, a finance function that spends its time on finance, was not.

Why the green board is the dangerous one

A red dashboard gets fixed. Someone notices, someone investigates, the problem surfaces.

A green dashboard sitting on top of a broken service is the harder case, because the green actively suppresses the investigation. Everything the organization built to detect trouble is reporting no trouble. The hidden cost doesn't merely stay hidden. It stays hidden behind a wall of evidence saying there's nothing to look for, and the better instrumented your dashboard, the more complete the suppression.

This is why "we have great metrics" isn't the reassurance people take it for. Great metrics tell you the proxies are well instrumented. They tell you nothing about whether the proxies still track the goals. Those are different questions, and the second is the one that ends careers quietly, eighteen months after the board went green and stayed there.

Say what kind of number you're holding

The research on that engagement set a target: take the gross-to-net close cycle from roughly three days to roughly half a day.

The kind of number that is matters more than its size. It wasn't a forecast, and it wasn't a vendor's projection of what a platform might deliver. It was a threshold, drawn from a peer finance function already operating at that speed inside the same company. The half-day figure existed, in the same organization, on real cycles, before anyone proposed anything.

That distinction is most of the discipline. An external benchmark tells you what someone else claims to have achieved under conditions you can't inspect. An internal baseline tells you what's already possible in your own building, with your own constraints, which means nobody in the room can dismiss it as someone else's special case. It has already survived the objection before the objection is raised.

If you can't say where a number came from and what kind of number it is, it isn't evidence. It's a guess formatted to look like one, and the people you're presenting to can usually tell.

Go and look at the work

There are a few versions of the check and any of them beats none.

The smallest: take one green metric that matters and ask what work it can't see. The larger one: go and watch. Not the dashboard, the work. Sit with the team whose output rolls up into that green number and watch what they actually do all day. The gap between what the metric measures and what the people are doing is where your next failure is already forming, funded and invisible, protected by the very instrument you built to catch it.

The structural version, if you have the authority: stop handing measures down. Bring the people who will be measured into the choice of what the organization is trying to achieve, and reward more than one thing. The research says that's what keeps a number attached to its meaning.

The green will hold right up until it doesn't. By the time the dashboard finally turns red, the thing it stood for has usually been broken for a year.


Key takeaways

A green dashboard is a claim, not a fact. Optimize the claim long enough and you lose the thing it stood for.

Concepts to name

  • Goodhart's Law (1975). When a measure becomes a target, it stops being a good measure.
  • Surrogation (Choi, Hecht & Tayler). People lose sight of the construct a measure represents and act as though the measure is the thing itself, because only the proxy can be acted on directly.
  • Proxy vs. goal. Every metric stands in for something you can't watch directly, and the proxy always costs less to satisfy than the goal.
  • The green board suppresses the investigation. A red dashboard gets examined. A green one is a wall of evidence saying there's nothing to look for.
  • Internal baseline beats external benchmark. What's already possible in your own building survives the executive-committee objection before it's raised.

The numbers

  • Incentives intensify surrogation. Tying compensation to a single measure measurably increases how completely people substitute the metric for the strategy (Choi, Hecht & Tayler, The Accounting Review, 2012). The usual fix for a stalling metric is the fastest way to destroy its meaning.
  • Participation reduces it. Involving people in selecting the strategy lowers surrogation (Choi, Hecht & Tayler, Journal of Accounting Research). People who helped choose the goal stay connected to it. People handed a target do not.

Techniques

  • Take any green metric that matters and ask what work it hides.
  • Go and watch the team whose output rolls up into the number. The gap between the metric and the work is where the next failure is forming.
  • Reward more than one measure, and bring the measured into choosing what's measured.
  • Prefer internal baselines to external benchmarks, and always say what kind of number you're citing.

Further reading

  • Choi, W., Hecht, G., & Tayler, B. (2012). "Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation." The Accounting Review.
  • Choi, W., Hecht, G., & Tayler, B. "Strategy Selection, Surrogation, and Strategic Performance Measurement Systems." Journal of Accounting Research.
  • Harris, M. & Tayler, B. (2019). "Don't Let Metrics Undermine Your Business." Harvard Business Review.

Sources

  • Goodhart, C. (1975). Goodhart's Law, as later formalized: "When a measure becomes a target, it ceases to be a good measure."
  • Choi, W., Hecht, G., & Tayler, B. (2012). "Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation." The Accounting Review, 87(4), 1135. Experimental evidence that single-measure incentive compensation intensifies surrogation.
  • Choi, W., Hecht, G., & Tayler, B. "Strategy Selection, Surrogation, and Strategic Performance Measurement Systems." Journal of Accounting Research. The finding that involvement in strategy selection reduces surrogation. The term surrogation originates with these authors.
  • Harris, M. & Tayler, B. (2019). "Don't Let Metrics Undermine Your Business." Harvard Business Review, 97(5). The managerial treatment of surrogation.
  • Figures from the author's engagement with a global pharmaceutical company's finance organization, Enterprise Performance Management workstream (2019): roughly 90% of long-term planning workload spent on manual data cleaning, and a gross-to-net close-cycle threshold of roughly 3 days to 0.5 days drawn from an internal peer-function baseline rather than a forecast.