Botsitting Eats the Hours Agents Were Meant to Save

6 min read Problems Agents

Workers spend hours a week botsitting AI agents: feeding context, catching mistakes, cleaning up. Autonomy levels and one inbox fix why.

Sol Rashidi runs strategy for a data security company and teaches at Harvard Kennedy School. She also ran four AI agents at once, until she fired two of them.

"I just fired half my agents because they were unreliable," she told Business Insider in July 2026. The agents were supposed to free up her time. Instead they demanded constant supervision: context that drifted and needed correcting, mistakes that needed catching, and output that needed cleaning up before it was usable. "I don't have the time to babysit agents and keep course correcting the context," she said. Her fix was to hire human virtual assistants for some of the work instead.

Rashidi is one of a growing number of people a Glean report calls "botsitters," workers who spend hours every week feeding AI context, debugging its mistakes, and cleaning up after it. The report puts the average at 6.4 hours a week, nearly a full working day, for white-collar workers using AI at their jobs.

Better training and better models will not end botsitting on their own. It happens when an agent's output has no defined path to being trusted, so a person becomes that path by default, on every run.

Why babysitting is the default

An agent told to "keep going until it's done" has no natural sense of when to stop and ask. A person doing the same recurring work eventually notices they are repeating themselves, or that a task has grown bigger than it should be, and pauses. An agent left alone has no such instinct unless the workflow gives it one.

Without that instinct, the only remaining safeguard is a person checking in. So people check in constantly: re-reading output before it goes out, re-supplying context an agent lost between runs, and watching a terminal to see whether the last hour of work is usable. Each check is small and reasonable on its own. Added up across four agents and a full week, it comes close to a full day of work that looks nothing like the work the agent was hired to do.

Rashidi's fix, fewer agents and more humans, is a rational response to that math, and it teaches the wrong lesson. The number of agents was never the problem. Nothing in the setup separated "the agent is doing something I don't need to see" from "the agent just did something that needs my judgment before it counts as done."

The fix is structural

Botsitting disappears when supervision stops being a habit and becomes a property of the workflow itself. Two things have to be true for that to work.

First, every agent needs an explicit answer to "how much can this agent do before someone has to look?" Task Machine calls this an autonomy level, and it is a ladder instead of a single on/off switch:

Level What the agent can do alone
Supervised Only works the tasks it is directly assigned. No delegation, no new agents, no workflow starts.
Balanced May attempt any consequential action, but each one waits for approval before it takes effect.
Autonomous Delegation, task assignment, and workflow starts happen directly. Creating new agents or workflows still waits for approval.
Full Every action applies directly, including creating new agents and workflows.

An agent starts at the level you set, and you raise it a step at a time as it earns the room. That one setting replaces the habit of watching everything with a decision made once per agent, which you can revisit as the agent proves it deserves more.

Second, the review the agent's work still needs has to land somewhere specific instead of wherever you happen to be looking. Task Machine organizes work around three surfaces: Chat to direct agents and fan work out into tasks and workflows, Tasks for the detailed back-and-forth on one piece of work, and the Inbox for everything that needs a decision, such as approvals, questions, proposed work, and exceptions. Without that third surface, review work leaks into every terminal, every chat scrollback, and every "let me just double check this" moment during the day. An inbox turns the leak into a queue you clear on your own schedule.

What replaces each botsitting habit

What botsitting looks like Why it happens What replaces it
Re-reading every output before sending it No verifier or approval step decided in advance what "good enough" means A verifier check or an approval node the workflow enforces before anything ships
Re-supplying context an agent lost between runs The agent has no durable memory of the work it already did Memory and task history attached to the work instead of the chat session
Watching a terminal to see if a long job is still on track No inbox item exists for the moment the job needs a decision An inbox item that only appears when the workflow reaches a real decision point
Deciding case by case whether to let an agent keep going No autonomy level is set, so every action is judged individually A named autonomy level per agent, raised deliberately over time
Fixing the same class of mistake repeatedly Nothing records what went wrong last time Step-level workflow history you can read back and correct from

Read down the list and the pattern is the one Rashidi ran into: every row is a decision that had no home, so it defaulted to whoever was watching at the time. Giving each one a home turns babysitting into review.

What this does not remove

None of this makes review work disappear. An agent at a higher autonomy level still needs judgment applied to its work, but that judgment happens at a decision point you chose instead of continuously. Raising an agent from Supervised to Autonomous also takes time, because it happens after you have watched it make a run of good calls at the level below.

The realistic goal is attention that is legible, bounded, and routed to one place, instead of spread across every terminal and chat window you have open.

Where this fits in Task Machine

Task Machine is built around this operating model. Every agent carries an autonomy level, from Supervised up to Full, and you move it as the agent shows it can be trusted with more. Work runs across the three surfaces, Chat to decide, Inbox to approve, Tasks to steer, so the moments that need your judgment arrive as a specific inbox item with the context attached instead of a vague feeling that you should go check on something.

Supervision stops being a background hum you can never turn off and becomes a queue you clear with intent, on a schedule you control.

If you are running agents and spending more of your week botsitting them than benefiting from them, start a 7-day trial.