A workflow is ready to hand over when another authorized person can operate it, recognize a problem, and stop safely without depending on its original builder. A runbook makes those responsibilities explicit. It should describe the actual setup and the team's decisions in language the next operator understands.
This guide is for a team lead preparing an operational handoff. You will build a short front section for routine use, with linked technical detail for the implementer. Supporting sources were checked on October 10, 2026.
Start with purpose and ownership
Open with the workflow's name, business purpose, and permitted scope. State what triggers a run, what information it may use, and what counts as finished. Include a clear sentence about whether the result is a draft, an internal update, or an approved external action.
Name the business owner, daily operator, reviewer, technical support contact, and backup. One person may hold several roles in a small team, but the responsibilities still need names. Explain who can pause work, approve a restart, change spending controls, or authorize broader access.
Add the date the runbook was last tested and the configuration it describes. A document that silently refers to an older setup can be more confusing than no instructions at all.
Write the normal operating path
Describe the process as the operator experiences it: check the approved input, start the workflow, confirm that it ran, inspect the result, and record the outcome. Include the exact places they should look, using verified internal links or screen labels.
For each step, state what success looks like. “Check the output” is vague. “Confirm that every requested item appears, unresolved details are flagged, and the result is still a draft” gives the reviewer something to do.
Explain what should happen with duplicate inputs, missing files, and tasks that have already been handled manually. Give each item a reference that makes it possible to distinguish a new run from a repeat. Include a simple way to record that an item was intentionally skipped.
Document permissions without exposing secrets
List the connected services and the actions each connection permits. Describe who owns the account and where an authorized administrator manages access. Keep passwords, API keys, recovery codes, and private tokens out of the runbook and screenshots.
Separate ordinary operating permission from permission to change the system. Someone allowed to review drafts may not be allowed to add a new provider or grant access to another folder. Explain the approval route rather than telling an operator to work around a restriction.
Also document where inputs, outputs, and activity records are stored, who can see them, and who owns retention decisions. The handoff should cover accidental disclosure as well as a workflow that simply fails to run.
Give alerts an action
For every important alert, record what it means, the likely business impact, and the first safe check. Link to the relevant dashboard or status page. Distinguish missing output, failed access, spending interruption, and suspicious activity; they should not all produce an instruction to “try again.”
Google's Site Reliability Workbook describes playbooks that explain an alert's impact and possible response, and emphasizes keeping them current. Adapt that principle to the scale of your team. On-call guidance.
Set a boundary for repeated attempts. If the output destination may have received the first attempt, the operator should verify the destination before repeating a write or send. Repeatedly clicking a run button can create duplicate work or additional charges without fixing the underlying problem.
Make pause and recovery specific
Write the supported pause procedure and identify what it actually stops. Pausing new intake may leave existing jobs running. Stopping an application may leave scheduled triggers enabled elsewhere. The implementer should verify these details for the real setup.
During an interruption, preserve the minimum evidence needed to investigate: task references, times, error messages, and configuration version. Avoid copying confidential input into broad support channels. Record who is coordinating the response and who communicates any business impact. Google’s incident-response guidance separates coordination, communication, and operational work. Incident response.
Recovery should include checking the destination for partial results, testing a harmless example, obtaining the required restart approval, and resuming a bounded amount of work. State when the operator must stop and escalate rather than attempting recovery alone.
Fictional example
Elm Demo Services is an invented maintenance company. Its AI workflow prepares internal visit summaries from fictional notes. During a handoff drill, the backup operator cannot find the draft folder and is unsure whether a repeated task would create another document. The team updates the runbook with the verified folder location and a check for existing task references. The example illustrates a useful drill, not a completed customer installation.
Handoff checklist
- Confirm the purpose, allowed actions, and definition of completion.
- Name owners, reviewers, backups, and approval responsibilities.
- Write routine steps with visible success checks.
- Link approved access-management and storage locations.
- Give each important alert a first action and escalation point.
- Verify pause, partial-result checks, and restart procedures.
- Have a new operator complete a fictional run and interruption drill.
- Fix every instruction they had to ask the builder to explain.
Keep technical reference material separate from the quick operating path, but link it clearly. Review the runbook whenever the model, tools, provider, schedule, permissions, or owner changes. Bring the draft and any failed handoff steps to InstallAI when discussing operational readiness. A tested document provides a concrete basis for deciding what still needs attention.
Sources checked
- Google Site Reliability Workbook On Call Checked 2026-10-10
- Google Site Reliability Workbook Incident Response Checked 2026-10-10