Choose a job you already understand
A useful first project begins with an ordinary task: organizing a new inquiry, drafting a handoff note or finding an answer in an approved handbook. Pick something a person on your team knows how to do today. That person can explain what a correct result looks like and spot mistakes during testing. Avoid starting with a vague goal such as “automate the office.” It is too broad to price, test or hand over.
Write a one-sentence scope
Use this pattern: “When this input arrives, prepare this output for this person to review.” For example: “When a service inquiry arrives, prepare a summary and a draft reply for the office manager.” This is an illustrative scope, not a claim that a connector is available. Name the source of the input and where the draft should appear. Keep sending, booking and payment outside the first version unless they are deliberately approved parts of the project.
Choose the information it can use
List the documents and fields needed for that one task. A reply draft may need an approved service list and business hours. It may not need a complete customer database. Identify an owner for each source, check that the information is current, and decide what the workflow should do when a fact is missing. “Ask a person” is a valid outcome. Do not test with private customer records when fictional examples can answer the same design question.
Make review part of the workflow
Decide who checks the result and what they are checking: facts, recipient, tone and any commitment. Keep the review close to the proposed action so the reviewer sees the exact content. OpenAI’s agent-safety guidance discusses human approval, structured outputs and evaluation as measures that can reduce risk. They do not remove the need to test your particular workflow. Source: OpenAI safety guidance.
Test the awkward cases too
Create a small set of examples that includes an ordinary request, missing details, contradictory information, a duplicate request and a request outside your services. Include an input that tries to tell the assistant to ignore its rules. Decide in advance what should happen in each case. Record the outcome and whether a reviewer had to repair the draft. A polished demonstration alone cannot establish reliability.
Compare the kind of help you need
Different tasks need different kinds of review. A summary condenses information that already exists. A draft response adds wording and may suggest a next question. A recommendation asks the system to judge alternatives. An external action changes something outside the conversation. Start by naming which of those results you actually want. Moving from a summary to a sent response adds a new consequence, even when the screen looks similar.
Four candidate workflows to discuss
For a small service business, an inquiry summary can be a useful candidate when someone already reviews every inquiry. The main checks are whether the request is represented accurately and whether missing details stay visible. A project handoff draft can help when notes are scattered, but it must preserve the difference between an agreed task and a suggestion. An answer from an approved handbook can be useful when the document has a clear owner; stale or contradictory policies need attention first. A marketing draft can help with wording, but every claim about results, customers or service availability still needs evidence. These are hypothetical candidates, not reports of installed integrations or customer outcomes.
Recognize a task that is not ready
Pause the idea when nobody can identify the source of truth, the reviewer or the consequences of a mistake. A task may also need ordinary process improvement before AI is useful. If three people disagree about which service policy applies, a new model will not settle that business decision. Resolve the policy and ownership question first. Likewise, if the goal depends on a tool connection nobody has verified, begin with a supplied fictional example rather than promising automation.
Bring a useful first brief
For the setup conversation, bring the task description, one ordinary fictional input and one awkward input. List the tools involved without sharing credentials. Explain what your team currently does with the result, who checks it and which steps should remain manual. Note which information is private and which costs you need clarified. This brief is enough to have a concrete discussion about fit without selecting a server, granting broad account access or assuming a particular product can perform every step.
Agree on the handoff
Before expanding the project, check whether the team can operate it, recognize a failure and pause it. Keep the access owner, running costs and recovery steps with the workflow instructions. Compare actual review effort with the old process, including correction time. Expand only when the first task is useful enough to maintain. Your next step is simple: write the one-sentence scope and bring three fictional examples to the setup conversation.
Sources checked
- OpenAI: Safety in building agents Checked 2026-10-10