How to Run an Experiment Assumption Sprint

6 min read Guides

A practical guide to surfacing hidden product assumptions, ranking them by risk, and designing cheap decisive tests.

An experiment assumption sprint is a short, structured pass over a product idea or plan to find the assumptions most likely to break it. The output is a ranked experiment plan that says which risk matters most, what cheap test would change your mind, and what signal counts before the test starts.

The reason to run one is simple: teams often build around the assumptions they least want to inspect. A sprint creates a ritual for finding the load-bearing beliefs before implementation turns them into sunk cost.

The real risk sits beneath the feature debate

Product discussions often stay at the feature level. The team debates scope, design, and timing while the real risk sits underneath: whether the customer cares, whether the workflow is usable, whether the business can support it, whether the team can build it, or whether the go-to-market path exists.

The cost appears later as rework. A product that should have been rejected becomes a prototype. A demand risk gets tested with a technical spike. A technical risk gets tested with a landing page. A polished demo wins internal approval but never measures the harsh truth.

What the manual process looks like

Done well, the sprint is a five-step ritual:

  1. Write the product idea or plan in plain language, including the target user and intended outcome.
  2. Surface assumptions from PM, designer, and engineer perspectives across eight categories: Value, Usability, Viability, Feasibility, Ethics, Go-to-market, Strategy, and Team.
  3. State each assumption as a falsifiable claim with confidence and what breaks if it is false.
  4. Prioritize the list on an Impact x Risk matrix, where Risk is (1 - Confidence) x Effort.
  5. Design tests only for High-Impact and High-Risk assumptions, with a pre-set pass, fail, or learn threshold.

That last constraint is what keeps the sprint useful. Testing everything is wasteful. Testing only the easy assumptions is theater.

What an agent can automate

The sprint has a fixed method and a lot of structured thinking, which makes it suitable for a PM agent:

  • Surface assumptions broadly. The agent checks all eight categories rather than stopping at Value, Usability, Viability, and Feasibility. Ethics, go-to-market, strategy, and team risks are often where early ideas fail.
  • Make assumptions falsifiable. It rewrites vague worries into claims that can be tested, with confidence ratings and failure consequences attached.
  • Prioritize with a matrix. It sorts assumptions by impact and risk, then applies the quadrant verdicts: Proceed, Experiment, Defer, or Reject.
  • Choose the right test family. Demand and willingness-to-pay risks become pretotypes with an XYZ hypothesis and skin-in-the-game signal. Technical, task, alignment, edge-case, or workflow risks become Proof-of-Life probes.
  • Set the threshold first. Every test gets its metric and decision rule before it runs. Probes also get a disposal plan so throwaway artifacts do not become accidental product.

The agent proposes the tests. It does not run them until the human approves which ones are worth executing.

The guardrails that make it safe

The danger in experiment work is false confidence. A nice prototype, a favorable opinion survey, or a broad market report can feel like evidence while leaving the core risk untouched.

The safe shape is a workflow that writes the experiment log, self-critiques the plan, and then waits. The reviewer checks whether all eight categories were covered, whether only High-Impact and High-Risk assumptions earned experiments, whether the tests measure behavior rather than opinion, and whether the thresholds were set before the work starts.

Set it up in Task Machine

The Assumption testing & experiment design playbook provides a starting point for the method above. You need an active Task Machine workspace with Chat, workspace-management and Playbook-installation access (workspace owners have it). No connected services are required.

1. Find the playbook

Open Search in your workspace and enter "Assumption testing & experiment design". The command center lists Set up Assumption testing & experiment design under Playbook setup.

The command center offering Set up Assumption testing & experiment design

2. Start the conversation

Choose Set up Assumption testing & experiment design. Task Machine opens a dedicated Chat with the Playbook card and an editable, unsent request. Read the intended job and outcome. Add your situation and send it when ready. Opening the draft does not install anything or start work. This walkthrough uses settings that require approval of the proposed Playbook.

Chat with an editable unsent request based on Assumption testing & experiment design

3. Agree the working brief

Use Chat to agree the inputs, expected output and limits before asking for a proposal. The Agent needs the riskiest assumption, the success metric, experiment ideas, and the sprint timebox. Use concrete language. For example, name the customer behavior you need to see, the threshold that would change the decision, and the timebox for the first probe.

Chat recording the working brief and review boundaries for Assumption testing & experiment design

4. Review the proposed Playbook

Ask the Agent to generate the Playbook from the agreed brief. Open its proposal in Chat and check the instructions and resources it will install, which carry more detail than the conversational summary. Review the plan before installing. Confirm the workflow proposes tests, waits for approval, and records pass, fail, or learn criteria before any test runs. Ask for a revised proposal if anything is missing or changes the job.

The Assumption testing & experiment design proposal reviewed inside Chat before approval

5. Approve and prepare the first work

Choose Approve on the proposal in Chat when the configuration matches your brief. Task Machine installs that reviewed configuration. The approved item retains its review details. If your autonomy settings allow direct installation, this approval may not be required. Check the resulting configuration in that case too.

Complete any remaining secure service setup from the installation details in Chat. Inbox keeps those setup items available if you return later. Prepare the source documents and inputs before starting the first Task or Workflow. Installation does not authorize sending, publishing or changing an external service beyond the boundaries you agreed. Check the reviewed tracker or results document before the first cycle so evidence and human decisions have a durable home.

The approved Assumption testing & experiment design configuration in Chat

What good looks like

A useful sprint leaves you with three artifacts:

  • A ranked assumption list. The highest item is both high-impact and high-risk, whether or not it is easy to test.
  • A decisive test design. The method matches the risk type, and the threshold is set before execution.
  • A disposal plan. Any throwaway probe has a planned end date so it does not become unowned production work.

Common questions

Should every assumption get an experiment? No. Only High-Impact and High-Risk assumptions should earn experiments. Low-risk assumptions can proceed or wait, and low-impact high-risk assumptions are often rejected.

What is the difference between a pretotype and a Proof-of-Life probe? A pretotype tests demand or willingness to pay. A Proof-of-Life probe tests whether a technical, task, alignment, edge-case, or workflow risk can survive contact with reality.

Can the agent run the experiment automatically? No. It proposes the experiment plan and waits for approval. The human decides which tests run.

What should go in the experiment log before the first run? Add current bets, prior tests, known assumptions, available channels, success thresholds, and decision rules.