Run an AI Company

Autonomy levels in practice

How to raise an agent's autonomy step by step as its work proves itself, and what each level feels like day to day.

Autonomy in Task Machine records how much an agent has earned rather than describing its personality. Every agent starts supervised. As human approvals and rejections accumulate, Task Machine can propose a one-level move, and you decide whether to apply it. This page explains what each level feels like, when to move, and what never loosens.

Everything starts supervised

A new agent, a new playbook, a new kind of work — all of it begins at Supervised, where every consequential action waits for your approval. This is not distrust of the technology. It is the same posture you would take with a new hire. The first cycles are where you learn how the agent interprets your instructions, where its drafts drift from your voice, and which questions it should have asked. Supervised mode makes all of that visible in your inbox and on the task timeline before anything lands where it matters.

Because autonomy is set per agent — and scoped further by the project or goal the work belongs to — trust does not transfer automatically. An agent that has earned latitude on weekly reporting still starts supervised when you point it at outreach. That granularity is the point: you are not deciding whether to trust "AI", you are deciding whether this agent has proven itself on this work.

Each level changes what interrupts you

Day to day, the levels differ in what reaches your inbox. At Supervised, every written plan waits for a person before work starts. At Balanced, Routine and Standard plans can start without separate plan review, while Elevated and Critical plans go to a human reviewer. Completed work still keeps its human review. At Autonomous, a Routine or Standard plan can also skip review, while higher-risk plans go to the assigned reviewer, who may be a human or an agent. Your inbox carries the exceptions: a verifier failure, a genuine question, or a budget threshold. Full autonomy uses the same plan and completion reviewer choices while loosening the remaining action gates, and you steer through outcomes, reports, and the task papertrail.

The right reading of this ladder is attention economics. Each step up trades review time for exception handling, and each step is only worth taking when reviews have stopped finding anything. The Autonomy settings page also lets you choose workspace fallback planners and reviewers. A task first keeps its explicit roles, then resolves defaults independently from its assigned agent, project, goal, and workspace. When no default is configured, the implementer plans the work. The reviewer handles plans when review is required and signs off completed work. Human-review levels fall back to the human task creator for review.

The Autonomy settings page with separate cards for the autonomy preset, default planner, and default reviewer

Evidence proposes the next move

Task Machine records the consequential decisions people already make on an agent's work, including approved and rejected proposals, task plans, and sensitive-action requests. Once the record is strong enough, it uses a confidence-bounded approval rate to propose promotion or demotion by one level. The proposal arrives in your inbox, where the score, direction, consequences, and apply or dismiss actions stay together. The level never changes without that human decision.

Sometimes a mature record is mixed enough to earn neither promotion nor demotion. Task Machine can then compare representative approvals and rejections with the agent's current instructions. It is allowed to conclude that no instruction change is justified. When it finds a recurring instruction-level cause, it proposes one complete replacement and shows the diagnosis, cited decisions, expected effect, and exact before-and-after text in the Inbox. Approving the change starts a fresh measurement window so the revised behavior earns its own record.

Budgets and checks hold at every level

Raising autonomy never removes the hard boundaries. Budgets cap what an agent can spend at every level, including Full autonomy. A fully autonomous agent that hits its cap pauses and asks, same as a supervised one. Verifier steps in workflows keep checking output against your criteria regardless of who approves it, and every run still writes its history to the task. Repository work also keeps revision-bound review before merge at every level. Autonomy changes which plans need separate review and who may perform required reviews, but it does not change what is measured, capped, and recorded.

An outreach agent earning trust over weeks

The arc looks like this in practice. Week one, you install an outreach playbook and the agent runs Supervised: it researches prospects and drafts messages, and every draft comes to your inbox. You edit heavily at first, and your corrections go back as instructions and knowledge. By week three the drafts are arriving clean, so you raise the agent to Balanced: it now researches, drafts, and prepares sends on its own, and only the approval step before sending interrupts you. A few weeks of approvals where you change nothing is your signal — you move to Autonomous for the established segments, keeping the approval step only for new audiences. The agent's budget and the verifier that checks each draft against your voice guide never moved. What moved is how often you say yes to work that was already right.

Stepping down is routine, not a reversal

The dial turns both ways, and turning it down should carry no more drama than turning it up. Sooner or later an agent that was running clean has a bad streak: the verifier starts failing drafts it used to pass, your edits creep back in, an approval you would have waved through last month makes you pause. The move is simple — step the autonomy back one level for that agent on that work, and nothing else. The same scoping that made raising precise makes lowering precise: the outreach agent goes back to Balanced for the segment that is misfiring while its reporting work stays Autonomous, and no other agent is touched.

Two things keep this undramatic. First, the record tells you why. The run history and approval trail show when clean runs stopped being clean and in which step, so you are diagnosing, not doubting. Most bad streaks turn out to be a change in the work rather than in the agent — a new audience, a new format, a document that went stale — and the fix is updating the instructions or the knowledge the agent reads. Second, a step down costs almost nothing. You review more for a few cycles, the corrections land, and the same evidence rule that raised the level the first time raises it again.

The trap to avoid is the opposite reflex: leaving autonomy high because lowering it would feel like admitting the delegation failed. It did not fail — the boundary did its job by surfacing the miss early. Treat the level as this month's setting, not a verdict.

Knowing when the work has proven itself is its own skill, and it should rest on evidence rather than a good week — Deciding when to trust agent output takes that up next.