A zero-based bar shows a 15% estimated average increase in customer-support issues resolved per hour with AI assistance in a 2025 study. The overall study included 5,172 agents at one firm; results varied and do not predict business savings.A zero-based bar shows a 15% estimated average increase in customer-support issues resolved per hour with AI assistance in a 2025 study. The overall study included 5,172 agents at one firm; results varied and do not predict business savings.
Sourced research chart. A 2025 QJE study estimated an average 15% increase in issues resolved per hour following AI-assistance rollout at one software firm. Overall sample: 5,172 support agents; this outcome used the subset with quality data. Main rollout: fall 2020–winter 2021. Rounded abstract estimate; uncertainty is not drawn here. Results varied across workers and are not an InstallAI result or a savings guarantee. Sources: Brynjolfsson, Li & Raymond (2025), Generative AI at Work.
Long description

The horizontal axis runs from 0% to 20% change. One bar begins at zero and ends at +15%. It summarizes the published abstract, not observed before-and-after means. The study used a staggered-rollout difference-in-differences design; human agents could edit or ignore AI suggestions. One firm, a specific support tool, worker differences, and incomplete coverage of quality metrics limit generalization.

An AI pilot is useful when it helps you decide what to do next. That decision should reflect the complete work, including preparation, review, errors, and operating responsibilities. A polished demonstration or an enthusiastic comment cannot establish whether the workflow is ready for everyday use.

This guide is for a business owner reviewing a limited pilot. It helps organize the evidence and choose between expanding, revising, or stopping. The approach is original practical guidance, supported by primary sources checked on October 10, 2026.

Restate the question the pilot was meant to answer

Begin with the original task and its success criteria. What work was included? Which users, inputs, actions, and service connections were allowed? What would have counted as a critical failure? Preserve those boundaries in the review.

If the scope changed during the pilot, describe when and why. Separate observations before and after the change rather than combining them into a single headline. A workflow that moved from drafting to updating records was tested under materially different conditions.

Check whether the acceptance criteria were written before the results were known. If they were created afterward, say so and avoid presenting the exercise as a fully predetermined test. The next round can use clearer criteria.

Account for every eligible task

List the tasks that entered the pilot, including failed, abandoned, manually completed, and excluded items. Explain exclusions. A successful-output folder alone does not show how much work the system attempted or which inputs it could not handle.

Separate ordinary cases from unusual ones. The sample may contain mostly easy work because a supervisor screened difficult items out. That can be a reasonable pilot boundary, but it limits what you can conclude about the wider workload.

NIST's Generative AI Profile cautions against extending performance claims from narrow or anecdotal assessments. Keep your conclusion within the conditions you actually observed. NIST AI 600-1, action MS-2.5-001.

Compare complete effort with a fair baseline

Include finding inputs, preparing them, starting runs, waiting where it affects work, reviewing output, making corrections, and resolving failures. Distinguish staff attention from elapsed time. A job can run for a long time without needing attention, or finish quickly while demanding substantial checking.

Compare equivalent tasks under equivalent standards. Do not compare an AI draft with a manually approved final document. If an experienced employee created the manual baseline while a new operator handled the pilot, record that difference.

Keep one-time setup and training effort separate from recurring operation. Both matter, but they answer different questions. Report observed effort without automatically turning it into financial savings, reduced headcount, or additional revenue. Those claims require further assumptions and evidence.

Review failures by consequence

Group problems into understandable categories: unsupported facts, missing details, inappropriate actions, privacy or access issues, service interruptions, and outputs requiring extensive rework. Include near misses caught by the reviewer; their detection is part of the workflow's cost and safety story.

For each significant failure, record the input conditions, what occurred, how it was caught, and whether the cause is understood. An average quality score can hide a serious boundary violation. Decide which failures prevent expansion regardless of the average.

Anthropic's evaluation guidance recommends measurable, relevant criteria and multiple dimensions of performance. Use separate measures for the outcomes that matter to this workflow rather than allowing writing quality to stand in for everything. Evaluation guidance.

Check whether the operating model is sustainable

Review actual provider usage, available billing records, subscriptions, hosting, and staff responsibilities. Note costs that are still estimates or not yet reported. Confirm who owns monitoring, source updates, permission reviews, and recovery when something fails.

Ask whether the reviewer can keep up with the proposed volume. A pilot may have worked because its builder watched every task closely. If that level of attention cannot continue, the next phase needs a different scope or operating plan.

Also revisit data handling. Confirm that pilot records are stored in approved locations and that the proposed expansion would not quietly introduce new recipients, more sensitive information, or additional external actions.

Fictional example

Cedar Demo Dispatch is an invented service coordinator testing draft internal job notes. The sponsor discovers that the initial review counted only accepted drafts and omitted tasks sent back for clarification. The team reconstructs the full task list and separates review time from generation time. It then decides what evidence another limited round would need. This example illustrates honest review; it contains no measured customer result or promised benefit.

Make one clear decision

Expand when the evidence supports the intended next scope and the operating responsibilities are covered. Name the specific additional task group or volume, the approver, and the conditions that would pause the expansion.

Revise when a bounded change could answer an important unresolved question. State the change, the examples it must pass, the owner, and when the result will be reviewed. Avoid an open-ended pilot that keeps consuming effort without a decision point.

Stop when the workflow does not justify further work or cannot meet necessary boundaries. Preserve useful lessons, close unused access through the authorized process, and arrange appropriate retention or cleanup of pilot records.

Pilot-review checklist

  1. Restate scope and the original acceptance criteria.
  2. Account for all eligible tasks and explain exclusions.
  3. Compare complete effort against equivalent manual work.
  4. List critical failures and near misses separately.
  5. Review recurring costs, ownership, and data handling.
  6. Choose expand, revise, or stop with a named approver.
  7. Record the next evidence needed and a clear stopping condition.

Bring this decision record and representative failures to InstallAI when considering a next phase. An honest review gives the implementation conversation a useful starting point, even when the right answer is to simplify the workflow or stop.

Sources checked