Running AI Agent Client Work Without Losing the Margin
Real agent-task data shows agencies how model cost, review, approvals, and reusable delivery workflows affect capacity and margin.
Founder, Task Machine
Agency growth is constrained by skilled people and the number of billable hours they can deliver. Hiring adds capacity slowly and raises fixed costs. Asking the existing team for more output eventually reduces quality or burns people out.
Agents create another source of production capacity. The commercial opportunity is more client work at the same team size, while senior people keep control of scope, quality, and approval. That only improves margin when the agency can see model spend, review effort, requested changes, and accepted delivery.
We analyzed an anonymized, platform-wide 30-day window of work run through Task Machine from August 6 through September 4, 2026. It included 534 tasks with active time, 533 tasks with recorded cost, and 68 completed tasks across product, operational, and real agency delivery work.
Tasks with active time
534
Average recorded cost
$21.43
Tasks completed
68
The interface screenshots use fictional examples. The published figures come from the anonymized 30-day dataset.
The dataset suggests six lessons for agencies that want more delivery capacity without sacrificing client quality.
1. Record revisions as part of delivery
Average active execution time was 1 hour 43 minutes per task, and average recorded cost was $21.43. The stage history shows where the money went in the order work moved through delivery.
| Delivery stage | Share of recorded cost | Agency question |
|---|---|---|
| Planning | 6.1% | Did the brief define an executable outcome before production started? |
| Implementation | 43.0% | Did the agent receive the approved sources and client constraints? |
| Review | 7.3% | Did the reviewer check the result against agreed criteria? |
| Rework | 40.5% | Was the requested change a correction, refinement, or new client scope? |
| Follow-up | 3.2% | Did accepted work create a clear next action? |
Rework covers agent execution after review returns a task for changes. It can include fixing a missed requirement, applying an internal or client review comment, or refining a deliverable after the first version made a tradeoff visible. The 40.5% figure measures cost share only. Task failure rate requires a separate count.
The agency needs to record why each revision happened. A misunderstood fixed requirement should improve the next brief. A new client request should become changed scope. A failed automated check should be resolved before account-lead review.
Task comments and mentions keep each revision reason beside the run, evidence, and client outcome it changed.

The activity history keeps revision reasons attached to the delivery so the agency can distinguish correction from changed scope.
2. Price from average cost, model choice, and an outlier reserve
The median task cost $0.20, the average $21.43, the 90th percentile $81.42, and the highest task $468.04.
Average recorded task cost
Across 533 tasks with recorded cost
$21.43
The median may describe a tiny review or classification task. A fixed delivery price based on that number would underfund substantial work. The average preserves total spend across a portfolio, and service-specific history improves it once enough engagements exist.
Model selection can materially change the cost of similar token usage. Every estimate should name the expected model, human review allowance, and amount reserved for requested changes.
| Estimate field | What to decide |
|---|---|
| Client price | The commercial ceiling for the accepted outcome |
| Expected agent cost | Historical cost for this service type and model |
| Human review allowance | Senior delivery and approval time included in the price |
| Revision reserve | The correction cycle included before a scope decision |
Task-level usage and duration keep active time, elapsed time, total spend, and execution stages attached to the client delivery, giving an engagement owner evidence for the next estimate.

The task detail exposes the time and stage spend that the next client estimate needs to cover.
3. Convert the client brief into an executable contract
A client brief often aligns people around goals and tone. Agent execution also needs explicit sources, permissions, acceptance criteria, and clear conditions for when another run requires a decision.
Before execution, write six items:
- Outcome: Name the artifact, change, or decision the task must produce.
- Source boundary: List approved documents, repositories, systems, and dates.
- Client constraints: Preserve terminology, claims, brand rules, data boundaries, and excluded actions.
- Acceptance checks: Define the commands, checklist, evidence, or review criteria that determine readiness.
- Approval boundary: State which external action, publication, spend, or irreversible change waits for a person.
- Change rule: Explain what becomes a new task or scope decision before another attempt.
This contract makes review traceable. A reviewer can point to a missed condition, and the agency can separate correction from changed scope. Task Machine records the contract as a Work Spec during the agent loop.

The review surface exposes client risk and the delivery boundary before the account lead approves the Work Spec.
4. Keep client decisions attached to the deliverable
People approved 297 requests and rejected 98 during the reporting period. Agents also asked 237 questions, with 210 recorded answers.
Approval decisions
395
Agent questions
237
An approval record should identify the exact action or artifact, the version and evidence covered, the person with authority, and the next step after rejection. A later change should also show whether the earlier approval still applies.
Email and chat can carry a decision, but the message often drifts away from the run and artifact it authorized. Keeping the decision with the task makes support, handoffs, and later client questions easier to resolve. The Inbox presents the exact request, evidence, and approve-or-revise actions together.

The approval request identifies the client delivery, version, checks, and known limit before a person authorizes the action.
5. Package one complete client workflow
A reusable prompt reproduces words. A reusable delivery workflow preserves the brief, source boundaries, production steps, checks, review owner, approval, and response to failure.
Consider a monthly client performance report:
- The client supplies the approved data sources and reporting period.
- The agent prepares the analysis and draft in the agreed format.
- Automated checks catch missing sections, invalid totals, and unsupported links.
- The account lead reviews conclusions, claims, and client-sensitive wording.
- Requested changes return to the same task with a named reason.
- The approved version is delivered and the next reporting task reuses the updated instructions.
The first run may include setup and discovery. Later runs can reuse the intake, checks, and approval path. The agency can then compare cost and revision reasons across deliveries and see whether repetition improves margin.
The workflow builder makes each production, check, review, and approval step visible before the recurring client delivery runs.

The workflow graph exposes the delivery path and human handoff that every recurring run will follow.
6. Keep each client’s work and reporting separate
Each client's tasks, people, agents, documents, credentials, and delivery history belong in that client's workspace. Execution should stay inside that boundary. Agency-wide pricing analysis can use anonymized totals later.
Apply the separation in reporting:
- Share task outcomes and delivery evidence inside the client workspace.
- Use anonymized totals for agency pricing and process improvement.
- Keep raw prompts, transcripts, credentials, and client content out of central cost reports.
- Enforce access in every operation as well as the interface.
The public dataset in this article follows the same rule. It contains aggregate counts, averages, and stage shares while customer content and workspace names remain private. Task Machine workspaces keep each client's members, agents, tasks, documents, and credentials inside its own operating boundary.

Separate workspaces keep client delivery records and access boundaries distinct while the agency moves between accounts.
The confidential agency work in the dataset involved real delivery with external deadlines and review expectations. That experience reinforces the commercial test: additional agent capacity matters when accepted client work increases without hiding model spend or senior review time.
For the next recurring deliverable, write the delivery contract, select the model, name the reviewer, and set the revision reserve. Compare the result with the current averages on Task Machine Pulse.