Tasks

Task specs

On this page

Give the Task a clear result to achieve before choosing how much execution to inspect. Planning, implementation, and review have different jobs. A successful Run tells you that an attempt finished, while the Task still needs to meet its acceptance criteria.

Review boundaries

The Task brief states your intent. A Work Spec describes how the Agent plans to carry it out and records acceptance criteria. Assignment review can first check whether the work has a suitable owner. Planner, implementer, and reviewer can be different members.

An Inbox Task-spec review with risk, the plan, acceptance criteria, model choices, and approval actions together

An explicit planning-approval requirement sends the spec to a Human. High-sensitivity work also requires human review. Other plans may proceed without a separate review or go to an Agent reviewer, depending on risk, the resolved autonomy settings, and the assigned reviewer. Broader autonomy does not mean every plan bypasses human judgment.

Approve or revise the Work Spec

Read the planned result, scope, assumptions, and checks before approving. Request changes when the approach needs revision. A spec waiting for approval or revision is a distinct state from a failed execution attempt. Repeatedly retrying work will not substitute for the required decision.

Planning and review bracket the work

An agent-assigned task is not handed to an agent to "just start." Before the work is ready to run, Auto reads the task description and acceptance criteria, weighs the planning complexity against every applicable budget, and selects an available high-capability planning model. A straightforward, well-defined task can use an efficient planner.

Substantial ambiguity, coordination, sensitivity, or difficult verification can move the task to a stronger one.

The planning turn begins by analyzing the task, its history, and the relevant workspace and repository context. When the outcome and execution direction are concrete, the planner writes a self-contained Work Spec with Goal, Investigation and implementation approach, How to verify, Constraints & non-goals, Context reviewed, and Deferrals sections.

The acceptance criteria are a separate checkable list rather than a repeated section in the spec. This gives a fresh agent enough direction to carry out the work without relying on the original conversation. When several reasonable approaches would materially change the result, scope, architecture, risk, or rollout, the planner does not invent the decision.

It asks one question in your Inbox with three distinct approaches, their trade-offs, its recommended default, and three editable suggested replies. Planning waits for your answer and then resumes on that exact task. Details the agent can inspect or research itself never become unnecessary questions.

The planner assesses the Task's risk across four dimensions:

  • Blast radius: who or what could be affected.
  • Novelty: how unfamiliar the work is.
  • Sensitivity: whether it involves sensitive information or consequences.
  • Reversibility: how readily a mistake could be undone.

Each receives a score from 0 to 2. Together they determine a review level of Routine, Standard, Elevated, or Critical, which sets how much scrutiny the result needs.

Review the plan and its model choices

Once the Work Spec exists, Auto matches models separately to implementation, review, verification, and follow-up. It considers current model capabilities, pricing, and budgets on every decision. This matching finishes before the plan proceeds.

Whether the plan waits for a human in the inbox depends on its risk assessment and effective autonomy settings, including any explicit planning-approval setting. Highly sensitive plans always require human approval.

When a plan waits for approval, its Task and Inbox reviews show the same risk assessment, proposed work, and Models section. The planning row is read-only because that turn has already happened. Every future stage starts on Auto, showing the model and reasoning Task Machine currently recommends.

Keep Auto for the quick approval path, or choose another model available on the same worker and then choose one of that model's supported reasoning levels. Nothing changes when you open a menu: Task Machine saves the fixed stage choices together with the approval.

After approval, the review keeps the effective choices visible in read-only selectors.

After implementation, Reviews explains how to judge the delivered result.