In this guide
Choose the next commitment, not a guaranteed winner
  1. 01Name the decision
  2. 02Separate constraints and unknowns
  3. 03Inspect the support
  4. 04Choose a bounded test
An organizing sequence for this playbook, not a validated scoring model. Return to the decision when new evidence changes the picture.

Evaluating an early idea means deciding what it deserves next: a little more investigation, a bounded test, a revision, or a clear stop. It does not mean predicting its eventual success from a presentation and a spreadsheet. Keep three questions separate: is there a genuine constraint, what evidence supports the benefit, and what would the organization have to commit?

Imagine the equipment-hire business from our brainwriting guide. Its team has proposed an app showing collection readiness, a manual check of every equipment kit, and a message sent before customers travel. All examples in this article are fictional. The team has two days available for investigation, not permission to launch a new service. That distinction should shape the entire review.

Decide what this review can authorize

Write the decision at the top of the working document. “Choose one uncertainty to investigate this week” invites different reasoning from “approve a six-month investment.” Without that boundary, someone evaluates a rough sketch as a finished business, while someone else approves it as if the next step cost nothing.

Name the decision owner, the available time, the people affected and the evidence expected at the next review. A workshop can recommend an option without controlling the budget. Say so before asking people to vote. If the sponsor has already ruled out replacing the inventory system this quarter, put that constraint on the table rather than revealing it after the team has spent an afternoon designing one.

GOV.UK’s prioritization guidance connects choices to research, performance information and stakeholder input, and recommends revisiting them. For an early idea review, the useful implication is to make the basis of the decision visible. A priority is a judgment made under particular conditions, not an enduring property of an idea.

Put the proposals on comparable footing

Ask for a short description of each idea: who encounters the problem, what changes for them, how the proposed response would work and which belief matters most. A polished slide deck should not be the price of admission. Equally, a memorable title is not enough to evaluate.

In our hire business, “readiness app” becomes “customers see whether their reserved kit has been checked and set aside before travelling.” “Checklist” becomes “a named staff member checks the kit and its accessories before confirming collection.” These descriptions reveal a dependency: a digital readiness message still needs a dependable readiness decision.

Keep proposals distinct long enough to understand them. Combining everything into one attractive super-idea can conceal extra work. An app, a new storage routine and proactive customer calls are three commitments, even if they eventually belong together. Record who proposed an option so that questions can go back to them, but do not make status or presentation confidence a scoring criterion.

Three colleagues compare idea sketches at a table, examining a customer notebook beside a polished presentation board.
Figure 1. An attractive presentation and a useful evidence record do different jobs. In this imagined review, the team looks beyond the polished proposal to examine what supports it.Illustration credits

Separate a hard constraint from an unanswered question

A hard constraint is a condition the proposed next step must satisfy. Examples might include an agreed spending limit, an unavailable system interface or a required approval. Check that the constraint is real, current and relevant to this decision. A limit on a live launch may not prevent a paper prototype or a non-operational walkthrough.

Use three possible findings: the condition is met, it is not met, or it is unknown. Unknown is not a polite version of no. If the team does not know whether inventory data can be exported, assign that check rather than giving technical feasibility a low score based on a guess.

When an idea fails a genuine condition, ask whether a different version could satisfy it without losing its purpose. The app may be out of scope, while a manual readiness call could still be investigated. Conversely, do not disguise a rejected implementation as a “small experiment” if it creates the same operational exposure. Required specialist review still belongs with the people qualified to provide it.

Some preferences deserve explicit debate rather than the status of a gate. “It must look innovative,” “our competitors already do this” and “the director likes apps” are not explanations of customer benefit. If novelty matters strategically, describe why and how much weight it should carry.

Judge evidence separately from attractiveness

For each proposal, distinguish the size of the hoped-for benefit from the strength of the support. An ambitious idea can have little evidence; a modest improvement can be well supported. Combining those judgments into a single traffic light makes it hard to tell which problem the team needs to solve.

For the hire business, a customer describing an unnecessary journey is relevant to the problem. It does not establish that an app is the preferred response. A prototype that people can understand supports a usability judgment, not a claim that inventory information will remain accurate during a busy shift. Use customer interviews for accounts of experience and assumption mapping to identify the consequential beliefs still in play.

There is a reason to be careful about familiar-looking options. In the studies summarized by Rietzschel, Nijstad and Stroebe, participants favored feasible and desirable ideas at the expense of originality. Explicit originality instructions changed that balance, with trade-offs in other ratings. This is evidence about those selection tasks, not proof that unusual proposals make better investments. It does suggest making your selection criteria explicit instead of assuming that “best” means the same thing to everyone.

Ask reviewers to write their initial reasoning before discussing a shared score. Then examine disagreements. One person may understand the maintenance workload, while another knows the customer’s time pressure. Averaging their ratings too early would hide information the decision needs.

Use a small decision record, not a magic number

The following comparison is a proposed working record for our fictional team. It contains questions and judgments, not measured results.

ProposalBenefit to investigateCrucial uncertaintyProportionate next step
Readiness appFewer unnecessary collection journeysCan the displayed status be trusted?Trace how readiness is established today
Complete-kit checkMissing accessories found before collectionCan staff perform and maintain the check?Observe a bounded preparation trial
Pre-collection callCustomers hear about a problem before travellingCan contact happen early enough to help?Explore timing with customers and staff

If you use weighted scoring, define the scale and weights before applying them. A score of four should mean something describable, not simply “I like it.” Keep the supporting note beside the number and mark missing information explicitly. Do not compensate for a failed essential condition by adding points elsewhere.

Try a simple sensitivity check: would a modest, plausible change in the weights reverse the ranking? If so, the apparent winner depends on a preference the group should discuss. Decimal precision does not remove that uncertainty. Nor should one scoring table force a small service adjustment and a new business model into the same investment logic.

Choose the smallest informative commitment

The Design Council’s framework places testing and refinement alongside developing alternatives. That is useful here: evaluation can lead to learning rather than an immediate build-or-kill decision. The next step should address the uncertainty that could change the choice.

Suppose our fictional team decides to explore the complete-kit check first. It appoints a preparation lead, chooses a limited set of ordinary collections and defines what to record: which accessories were checked, whether the kit stayed together, how much work the check added and what happened when something was missing. The team keeps existing safety and service controls in place. It does not promise that the check will solve the problem.

An equipment-hire worker checks a drill case and its accessories while a colleague records observations on a clipboard.
Figure 2. In the fictional hire business, a bounded kit-preparation trial can reveal missing accessories and extra work. It does not establish the value of a complete booking app.Illustration credits

Agree how the observations will be interpreted before beginning. If the check cannot survive normal handoffs, the storage or ownership arrangement needs revision. If it works only because an extra observer does the job, that is not evidence that normal staffing can sustain it. If the problem is mainly late changes to reservations, checking earlier may address the wrong point in the process.

For a communication or interaction question, a prototype may be sufficient. GOV.UK’s prototyping guidance describes using representations of different fidelity before committing to a build. Choose fidelity around the question: a sketched notification can support a conversation about meaning, but it cannot demonstrate that a live data feed is reliable.

Return a useful answer to the people who contributed

Close the review with one of four plain-language outcomes: investigate, revise, defer with a reason, or stop. Attach an owner and a next review point to anything that remains active. “Interesting—put it in the backlog” is not a learning commitment if no one knows what would bring it back.

Explain the reason at the same level of specificity as the proposal. “Not strategic” gives the contributor little to work with. “We are investigating collection readiness first; payment changes are outside this quarter’s scope” makes the boundary understandable. Preserve valuable observations even when the proposed response is not selected.

After the test, compare the original reasoning with what was learned. Did the team answer the intended question? Was the evidence weaker than expected? Did a dependency become visible? That record helps the next review begin with knowledge rather than another round of pitches.

Evidence and limits

This playbook combines research on idea selection with institutional design and prioritization guidance. Its review record, worked example and suggested decisions are practical adaptations, not a validated scoring instrument. The cited idea-selection research is used within the limits of its published abstract; it does not establish the effectiveness of this particular workflow.

Use the approach when an idea is specific enough to examine and the next commitment is bounded. It is not a substitute for technical, financial or other specialist assessment where consequences require it. Sometimes the defensible decision is to stop. Sometimes it is to keep an unfamiliar option alive long enough to ask a better question.

Sources: