Running AI Agent Client Work Without Losing the Margin

8 min read Agencies Agents

Real agent-task data shows agencies how model cost, review, approvals, and reusable delivery workflows affect capacity and margin.

Agency growth is constrained by skilled people and the number of billable hours they can deliver. Hiring adds capacity slowly and raises fixed costs. Asking the existing team for more output eventually reduces quality or burns people out.

Agents create another source of production capacity. The commercial opportunity is more client work at the same team size, while senior people keep control of scope, quality, and approval. That only improves margin when the agency can see model spend, review effort, requested changes, and accepted delivery.

We analyzed an anonymized, platform-wide 30-day window of work run through Task Machine from August 6 through September 4, 2026. It included 534 tasks with active time, 533 tasks with recorded cost, and 68 completed tasks across product, operational, and real agency delivery work.

Tasks with active time

534

Average recorded cost

$21.43

Tasks completed

68

The interface screenshots use fictional examples. The published figures come from the anonymized 30-day dataset.

The dataset suggests six lessons for agencies that want more delivery capacity without sacrificing client quality.

1. Record revisions as part of delivery

Average active execution time was 1 hour 43 minutes per task, and average recorded cost was $21.43. The stage history shows where the money went in the order work moved through delivery.

Delivery stage Share of recorded cost Agency question
Planning 6.1% Did the brief define an executable outcome before production started?
Implementation 43.0% Did the agent receive the approved sources and client constraints?
Review 7.3% Did the reviewer check the result against agreed criteria?
Rework 40.5% Was the requested change a correction, refinement, or new client scope?
Follow-up 3.2% Did accepted work create a clear next action?

Rework covers agent execution after review returns a task for changes. It can include fixing a missed requirement, applying an internal or client review comment, or refining a deliverable after the first version made a tradeoff visible. The 40.5% figure measures cost share only. Task failure rate requires a separate count.

The agency needs to record why each revision happened. A misunderstood fixed requirement should improve the next brief. A new client request should become changed scope. A failed automated check should be resolved before account-lead review.

Task comments and mentions keep each revision reason beside the run, evidence, and client outcome it changed.

A fictional Task Machine client-delivery task activity showing an agent question, the agency's answer, an implementation update, and requested changes in one history

The activity history keeps revision reasons attached to the delivery so the agency can distinguish correction from changed scope.

Takeaway: Add one revision reason to every returned task: missed requirement, failed check, agency refinement, or changed client scope. Review the totals before pricing the next engagement.

2. Price from average cost, model choice, and an outlier reserve

The median task cost $0.20, the average $21.43, the 90th percentile $81.42, and the highest task $468.04.

Average recorded task cost

Across 533 tasks with recorded cost

$21.43

Median $0.20 Average $21.43 90th percentile $81.42 Highest $468.04

The median may describe a tiny review or classification task. A fixed delivery price based on that number would underfund substantial work. The average preserves total spend across a portfolio, and service-specific history improves it once enough engagements exist.

Model selection can materially change the cost of similar token usage. Every estimate should name the expected model, human review allowance, and amount reserved for requested changes.

Estimate field What to decide
Client price The commercial ceiling for the accepted outcome
Expected agent cost Historical cost for this service type and model
Human review allowance Senior delivery and approval time included in the price
Revision reserve The correction cycle included before a scope decision

Task-level usage and duration keep active time, elapsed time, total spend, and execution stages attached to the client delivery, giving an engagement owner evidence for the next estimate.

A fictional Task Machine client-report task showing review history and its pending delivery approval beside active time, total elapsed time, total usage, and usage by stage

The task detail exposes the time and stage spend that the next client estimate needs to cover.

Takeaway: For the next proposal, price the expected model cost, named reviewer hours, and one bounded revision cycle separately. Require an engagement-owner decision before spending beyond that reserve.

3. Convert the client brief into an executable contract

A client brief often aligns people around goals and tone. Agent execution also needs explicit sources, permissions, acceptance criteria, and clear conditions for when another run requires a decision.

Before execution, write six items:

  1. Outcome: Name the artifact, change, or decision the task must produce.
  2. Source boundary: List approved documents, repositories, systems, and dates.
  3. Client constraints: Preserve terminology, claims, brand rules, data boundaries, and excluded actions.
  4. Acceptance checks: Define the commands, checklist, evidence, or review criteria that determine readiness.
  5. Approval boundary: State which external action, publication, spend, or irreversible change waits for a person.
  6. Change rule: Explain what becomes a new task or scope decision before another attempt.

This contract makes review traceable. A reviewer can point to a missed condition, and the agency can separate correction from changed scope. Task Machine records the contract as a Work Spec during the agent loop.

A fictional Task Machine Planning & review modal showing high blast radius, high sensitivity, the agent's client-risk assessment, and the delivery plan

The review surface exposes client risk and the delivery boundary before the account lead approves the Work Spec.

Takeaway: Attach the six-item delivery contract to the next client task and require the account lead to approve it before billable production begins.

4. Keep client decisions attached to the deliverable

People approved 297 requests and rejected 98 during the reporting period. Agents also asked 237 questions, with 210 recorded answers.

Approval decisions

395

Agent questions

237

An approval record should identify the exact action or artifact, the version and evidence covered, the person with authority, and the next step after rejection. A later change should also show whether the earlier approval still applies.

Email and chat can carry a decision, but the message often drifts away from the run and artifact it authorized. Keeping the decision with the task makes support, handoffs, and later client questions easier to resolve. The Inbox presents the exact request, evidence, and approve-or-revise actions together.

A fictional Task Machine Inbox approval request for client delivery showing the report version, completed checks, known limit, approval role, and approve or reject actions

The approval request identifies the client delivery, version, checks, and known limit before a person authorizes the action.

Takeaway: Before sending a client approval request, attach the exact version, completed checks, known limits, and approve-or-revise action to the same delivery task.

5. Package one complete client workflow

A reusable prompt reproduces words. A reusable delivery workflow preserves the brief, source boundaries, production steps, checks, review owner, approval, and response to failure.

Consider a monthly client performance report:

  1. The client supplies the approved data sources and reporting period.
  2. The agent prepares the analysis and draft in the agreed format.
  3. Automated checks catch missing sections, invalid totals, and unsupported links.
  4. The account lead reviews conclusions, claims, and client-sensitive wording.
  5. Requested changes return to the same task with a named reason.
  6. The approved version is delivered and the next reporting task reuses the updated instructions.

The first run may include setup and discovery. Later runs can reuse the intake, checks, and approval path. The agency can then compare cost and revision reasons across deliveries and see whether repetition improves margin.

The workflow builder makes each production, check, review, and approval step visible before the recurring client delivery runs.

The Task Machine workflow builder showing a top-to-bottom delivery graph with assignment details, branches, status, and version history

The workflow graph exposes the delivery path and human handoff that every recurring run will follow.

Takeaway: Choose one recurring client deliverable and document its intake, production steps, required checks, reviewer, approval, and revision rule before the next run.

6. Keep each client’s work and reporting separate

Each client's tasks, people, agents, documents, credentials, and delivery history belong in that client's workspace. Execution should stay inside that boundary. Agency-wide pricing analysis can use anonymized totals later.

Apply the separation in reporting:

  • Share task outcomes and delivery evidence inside the client workspace.
  • Use anonymized totals for agency pricing and process improvement.
  • Keep raw prompts, transcripts, credentials, and client content out of central cost reports.
  • Enforce access in every operation as well as the interface.

The public dataset in this article follows the same rule. It contains aggregate counts, averages, and stage shares while customer content and workspace names remain private. Task Machine workspaces keep each client's members, agents, tasks, documents, and credentials inside its own operating boundary.

A fictional Task Machine workspace switcher showing separate Northwind Studio, Harbor Health, and Alder Legal client workspaces

Separate workspaces keep client delivery records and access boundaries distinct while the agency moves between accounts.

Takeaway: Before onboarding the next client, create a separate workspace, assign its members, and decide which anonymized cost and revision totals may appear in agency-wide reporting.

The confidential agency work in the dataset involved real delivery with external deadlines and review expectations. That experience reinforces the commercial test: additional agent capacity matters when accepted client work increases without hiding model spend or senior review time.

For the next recurring deliverable, write the delivery contract, select the model, name the reviewer, and set the revision reserve. Compare the result with the current averages on Task Machine Pulse.