Actionable Metrics Start With a Decision

A product dashboard can look reassuring and still leave a team stuck. Signups rise, the cumulative user total reaches another milestone, and a feature chart points upward. In the meeting, the practical question remains: should the team improve onboarding, bring in more visitors, change the feature, or leave it alone? A number that cannot help answer that question is doing little work, however attractive it looks.

actionable metrics: decision compass, abstract chart screen, and action lever progressing left to right, outcome tokens, balance scale, blank notebook, potted fern

An actionable metric is a measure defined closely enough to inform a particular choice. The definition must say whose behavior counts, what event counts, over what period, and what comparison would make the result meaningful. It should also be close enough to a customer outcome that improving the number would be worth the effort. Mixpanel describes a vanity metric as one that can look impressive while failing to guide a decision, and notes that a cumulative total such as registered users needs more context to say anything useful about current performance in its discussion of vanity metrics.

That distinction does not make signups or page views useless. A signup count can answer whether a campaign is bringing people to a product. It cannot, by itself, answer whether those people accomplish anything after arriving. The same measure can be useful for one decision and inadequate for another. The way forward is to start with the choice facing the team, then give the metric a definition that can actually separate the available options.

Start with the decision the team could make

Suppose a product team is considering a simpler first-run experience. Its decision is whether to keep the new flow, revise it, or return to the existing one. “More activity” is too vague to settle that choice. The team needs to identify the first meaningful task a new user is trying to complete, then decide whether the flow makes that task easier without damaging another part of the experience.

For an illustrative document-sharing product, imagine the meaningful task is sending a document that a recipient can open. A count of account creations would tell the team how many accounts were created. A count of clicks on “Share” would tell it whether that control attracted attention. Neither number says whether a recipient could open the document. A more useful primary measure might be the share of new account holders who send a document that a recipient opens within seven days of signup. The seven-day window is an example choice for this fictional product, not a universal rule; a different product might require a shorter or longer window to reflect its actual use.

Writing the question first also limits unnecessary data collection. If the decision concerns new users’ first successful share, the team does not need every measure on the company dashboard in the room. It needs the success measure, the few steps that explain where people stop, and measures that would reveal a harmful trade-off. This is the practical meaning of “actionable”: a result can lead to a different, named action, rather than merely prompt another report.

A useful test is to finish two sentences before building a chart: “If this measure improves while the other conditions hold, we will ___.” “If it does not improve, we will ___.” If both blanks contain the same plan, the number may be interesting, but it is not yet helping with this decision. If the blanks cannot be filled, the team may need to sharpen the question or choose a different measure.

Give every rate a population and a clock

The denominator often carries the decision. Imagine, still illustratively, that 1,000 people sign up in one week and 200 complete the first successful share within seven days. That cohort’s seven-day completion rate is 200 of 1,000, or 20%. If 1,500 people sign up the following week and 240 complete the task within their own seven-day windows, the number of successful people rose, but the completion rate fell to 16%. The team gained 40 successful users while a smaller share of new users reached the task. Those are different facts with different implications.

The choice between a count and a rate depends on the question. If the immediate concern is how many successful users the product served, the count matters. If the concern is whether onboarding works for each new arrival, the rate is more revealing. Neither figure should quietly stand in for the other. Mixpanel’s point that raw totals need context is especially relevant when the number of people entering the product changes from one period to the next.

The population needs equal care. “New users” might mean every account created, only people who verified an email address, or people who actually opened the product. Those definitions produce different denominators. A team should select the one that matches the decision and keep it fixed across the comparison. Removing users who never reached the share screen might make the completion rate look better, but it would hide a failure in the path to that screen. Conversely, including accounts that were never eligible to share would make the flow appear worse for a reason it cannot address.

Time matters for the same reason. Comparing yesterday’s signups with last week’s completed tasks gives the earlier group more time to succeed. In the document example, each signup cohort should receive the same seven-day observation window before its completion rate is compared. If the decision cannot wait that long, a shorter measure can help the team monitor early behavior, but it should be named as an early signal rather than presented as the seven-day outcome.

Segmentation should follow a plausible difference in the user journey. A desktop flow and a mobile flow may expose different obstacles; a first-time visitor and a returning account holder may have different tasks. Breaking the rate down by such groups can show where a change deserves attention. The danger is treating the most favorable slice as the product’s overall result. A segment is useful when its definition was chosen for the decision and its size and comparison period remain visible.

Connect a customer outcome to a controllable input

A team still needs a way to link daily work to a larger product goal. Amplitude defines a North Star metric as a central measure of the value customers derive from a product in its North Star resources. For the illustrative sharing product, the company might consider successful document exchanges a candidate for that role. The word “successful” would need an operational meaning, such as a sent document that a recipient can open. Counting documents uploaded would represent a different event and could grow even when no recipient receives value.

That central measure is a direction, not a complete work plan. A product team can change the wording of an invitation, the steps in the share flow, or the way access settings are shown. It cannot directly order recipients to open documents. The team’s input measures should track the parts of the journey it can influence: for example, the proportion of new users who reach the share screen or the proportion of attempted shares that result in a usable link. Such measures help explain why successful exchanges might rise or fall. They also expose a poor tactic: increasing button clicks by making the button prominent while leaving the actual exchange broken.

This creates a useful chain for a decision. The broader outcome says what customer value the team wants to increase. An input measure describes the behavior the proposed work is expected to change. A task-level measure shows whether users complete the relevant journey. The chain is a hypothesis, not a guarantee. If the share-screen rate rises but successful exchanges do not, the team should investigate what happens after the screen rather than declare the new flow a success.

The choice of central measure should fit the product’s value exchange. “More time in product” might be useful for one service and a sign of friction for another. “More messages sent” might indicate valuable conversation, or it might mean that users have to ask repeatedly for help. The name of the metric cannot settle that ambiguity. The team has to explain what the customer receives when the counted event occurs and what would make an apparent increase hollow.

Put a possible downside beside the hoped-for gain

One attractive number is a poor basis for releasing a product change. Microsoft Research describes experiment results using an overall outcome, local measures that help explain it, data-quality measures, and guardrails that show unwanted changes elsewhere in its account of trustworthy experimentation. The point is practical: an improvement in the intended task can come with a cost that changes the release decision.

In the document example, a shorter share flow might increase completed shares while accidentally making access settings harder to understand. A guardrail could track failed recipient opens or access-related support requests, if those events are available and consistently defined. If those measures worsen, the team may need to revise the flow even though its primary completion rate improves. The preferred result is not merely a larger numerator; it is a better customer outcome at an acceptable cost.

The guardrail should reflect a credible way the specific change could hurt users. A new file upload method suggests watching failed uploads; a change that adds processing suggests watching delay; an invitation redesign suggests watching whether invited people can complete the next step. Collecting a long list of unrelated measures makes it harder to see the relevant trade-off. A small set tied to the proposed change makes the release call clearer.

That call also requires judgment. The supplied research does not provide a universal amount of improvement worth accepting or a universal amount of guardrail deterioration that is safe. A team should decide what downside it is unwilling to accept for this change, using its own product conditions, before looking at the result. If the downside is unacceptable, an increase in the primary measure does not erase it. If the costs are modest and the customer outcome improves, the team can weigh the gain against the cost openly.

Check what changed before saying why it changed

Metrics are observations, and an observation can move for reasons other than the product behavior a team intended to improve. Microsoft’s research on interpreting online experiments describes how changes in the measured population, lost event data, and short-lived novelty effects can alter what a metric appears to say about a product change. Those problems matter in ordinary reporting as well: a revised event definition or a new mix of incoming users can make this month’s rate unlike last month’s even when the experience itself has not improved.

Start with the event. If the team changes what counts as a successful share while redesigning the flow, it cannot interpret a rise in successful shares as a clean improvement in user behavior. It must know when the measurement changed and whether the old and new events can be compared. A useful metric definition therefore includes the exact counted event, the data source, exclusions, and the date on which any definition changed. That record lets the team distinguish an operational change from a measurement change.

Then look at who entered the denominator. Suppose the illustrative product begins attracting more users who only want to store documents privately. The share completion rate may fall because those users do not intend to share, even if the share flow works just as well for people who do. The team should examine the composition of the signup cohort and decide whether sharing is still the right first-value measure for the new audience. It should not quietly exclude the private users after seeing the result merely to restore a favorable chart.

Finally, consider whether the change lasts. A newly visible control can attract initial exploration that fades after people learn what it does. The Microsoft paper discusses novelty effects in which an early movement did not persist across later visits in an experiment it reports. For the sharing product, a spike in first-day clicks would not by itself justify a permanent redesign if seven-day successful exchanges stay flat. Waiting for the outcome window has a cost: the decision takes longer. It also prevents the team from mistaking curiosity for durable value.

If the team is comparing two product versions, it should not infer that one caused a change merely because two charts moved together. A controlled comparison can support a stronger conclusion, but even experiment results require careful interpretation when the counted events or populations differ; Microsoft’s paper documents that problem in its discussion of metric interpretation. When a clean comparison is unavailable, the honest conclusion may be that the measure changed and the cause remains uncertain. That is still useful: it tells the team what it must learn before committing more resources.

Turn the result into a specific next move

An actionable metric reaches its purpose at the decision, not at the dashboard. Return to the proposed first-run flow. If the seven-day successful-share rate rises for comparable signup cohorts, the event definition is stable, and the relevant guardrails remain acceptable, keeping the new flow is a defensible choice. If share-screen visits rise but successful exchanges do not, work on the step between preparing a share and opening it. If successful exchanges rise while failed recipient opens also rise, revise the access experience before expanding the flow. If the comparison is compromised by missing events or a changed population, repair the comparison before claiming a product gain.

These calls are deliberately conditional. The source material supplies no standard target rate for a sharing product and no proof that a particular product change caused a hypothetical increase. The illustrative figures above show how to read a denominator, not what any real product should achieve. A team should set its own threshold around the value of the outcome, the cost of the change, and the downside it can accept. Those choices are part of the decision; they cannot be imported from a generic list of “best metrics.”

The most useful artifact is a short written definition beside the chart: the decision it serves, the counted event, the eligible population, the observation window, the comparison, the likely input, the relevant guardrails, and the actions attached to the plausible results. A colleague should be able to read it and understand why a rise would matter and what would make that rise misleading. When the team cannot fill in those fields, the metric is not ready to carry the decision.

Actionable metrics do not require a more elaborate dashboard. They require a closer fit between a customer outcome and a choice the team can make. Start with the decision, define the population and clock, connect a controllable input to the outcome, and read any gain alongside its costs. The result is a number with a job: helping people decide what to do next.

Run your growth team from one screen.

Invite only