How to Red-Team a Plan With an Agent
A practical guide to pressure-testing plans, PRDs, and strategies with assumption grilling, pre-mortems, and approval.
Founder, Task Machine
Plan red-teaming is the practice of pressure-testing a plan, PRD, launch strategy, or product bet before the team commits to building. It extracts the assumptions that would break the plan if false, attacks those assumptions fairly, and turns the risk review into tests, thresholds, owners, and decisions.
The point is to find the cheapest useful evidence while the plan can still change. A good red-team memo helps a team decide what to test this week, what to harden before launch, and what is already well-reasoned.
Polite reviews leave the riskiest claim untouched
Most plans get reviewed politely. People comment on scope, language, and timeline, but the load-bearing claim survives untouched: the user has this problem, the proposed mechanism will change behavior, the dependency will land, the channel will work, or the market will tolerate the tradeoff.
When that claim is wrong, the team finds out after building. The expense comes later, in the sprint, launch, campaign, or implementation that followed a plan nobody challenged hard enough.
What the manual process looks like
Done by hand, a serious red-team review has a clear sequence:
- Read the plan and attached materials before asking the team for more context.
- Grill the plan one question at a time until the core bet can be stated in a single sentence.
- Extract every claim and separate load-bearing assumptions from cosmetic details.
- Steelman each important claim, then attack the strongest version instead of a strawman.
- Write each risk as a falsifiable "Fails if" statement.
- Run a pre-mortem by imagining the plan already failed and sorting risks into real risks, overblown concerns, and unspoken worries.
- Rank the surviving kill-assumptions by impact, likelihood, and cheapness to test.
- Write a memo with evidence to get this week, kill criteria, cheapest tests, mitigations, owners, and what could not be assessed.
This is a disciplined review. A generic risk list is not enough.
What an agent can automate
The Plan review & challenge playbook gives an agent that exact review pattern:
- Grill to understanding. The agent walks the plan branch by branch, resolves dependent questions in order, reads what is already attached, and finishes with the plan's load-bearing assumption in one line.
- Attack the steelman. It extracts claims, separates load-bearing from cosmetic, states the strongest case for each claim, then attacks that version.
- Make risks falsifiable. Each important weakness is written as "Fails if ___" and tied to evidence, a kill criterion, and the cheapest test.
- Run the pre-mortem. Risks are categorized as Tigers, Paper Tigers, and Elephants, with launch-blocking Tigers getting mitigation, owner, and date.
- Self-critique the memo. The agent checks for strawmen, fabricated risks, generic items, missing "what could not be assessed" notes, and whether the cheapest high-impact test is surfaced at the top.
The agent improves the review, but it does not approve the plan or decide what to build.
The guardrails that make it safe
Red-team work needs a fair adversary, not a chaos engine. The playbook requires the agent to attack the strongest version of each claim, state what is well-reasoned, and avoid inventing weaknesses when the plan is sound.
The final guardrail is human approval. The risk memo waits for a reviewer, who decides which tests to run, which mitigations to assign, and whether the plan should proceed, change, or stop.
Set it up in Task Machine
The Plan review & challenge playbook provides a starting point for the method above. You need an active Task Machine workspace with Chat, workspace-management and Playbook-installation access (workspace owners have it). No outside service authorization is required for the install.
1. Find the playbook
Open Search in your workspace and enter "Plan review & challenge". The command center lists Set up Plan review & challenge under Playbook setup.

2. Start the conversation
Choose Set up Plan review & challenge. Task Machine opens a dedicated Chat with the Playbook card and an editable, unsent request. Read the intended job and outcome. Add your situation and send it when ready. Opening the draft does not install anything or start work. This walkthrough uses settings that require approval of the proposed Playbook.

3. Agree the working brief
Use Chat to agree the inputs, expected output and limits before asking for a proposal. The Agent needs the plan summary, decision stakes, known failure modes, and risk tolerance. Give the agent enough context to attack the real bet: what the plan proposes, what decision depends on it, what already worries the team, and how much risk is acceptable.

4. Review the proposed Playbook
Ask the Agent to generate the Playbook from the agreed brief. Open its proposal in Chat and check the instructions and resources it will install, which carry more detail than the conversational summary. Check that the memo flow includes grilling, assumption attack, pre-mortem, ranked risks, self-critique, and human approval. Ask for a revised proposal if anything is missing or changes the job.

5. Approve and prepare the first work
Choose Approve on the proposal in Chat when the configuration matches your brief. Task Machine installs that reviewed configuration. The approved item retains its review details. If your autonomy settings allow direct installation, this approval may not be required. Check the resulting configuration in that case too.
Complete any remaining secure service setup from the installation details in Chat. Inbox keeps those setup items available if you return later. Prepare the source documents and inputs before starting the first Task or Workflow. Installation does not authorize sending, publishing or changing an external service beyond the boundaries you agreed.

What good looks like
A useful red-team memo is narrow and actionable:
- The core bet is explicit. A reviewer can see the one-line load-bearing assumption.
- Risks are falsifiable. "Fails if" statements name the condition that would break the plan.
- The top tests are cheap. The memo points to evidence the team can get this week.
- Sound claims are allowed to survive. A fair red-team review says what holds up as well as what does not.
Common questions
Is red-teaming the same as a pre-mortem? No. A pre-mortem imagines the plan already failed and works backward. Red-teaming attacks the assumptions and logic now. This playbook uses both methods.
Will the agent just produce a long risk list? It should not. The workflow ranks the top kill-assumptions and requires the cheapest test, kill criterion, and evidence for each. Generic risks fail the quality bar.
Can this replace product leadership judgment? No. It sharpens the decision. The human reviewer still decides what to test, harden, defer, or stop.
What if the plan is sound? The agent should say which claims are well-reasoned and why. The instruction explicitly forbids fabricating weaknesses just to fill the memo.