How-to guide

Who owns an automation after go-live? Failed runs and staff changes

A practical runbook for Power Automate after go-live: accountable owners, failed-run recovery, connection handover and verified business outcomes.

An automation needs an owner after the project team has moved on. Someone must notice when work stops, decide how an incomplete transaction should be handled and make sure the process still works when a member of staff changes role or leaves.

For a Power Automate cloud flow, that responsibility extends beyond the name shown as owner in the product. It includes the business decision, the connected systems, the identities used by each connection and the evidence needed to establish what actually happened. The following runbook is an illustrative starting point for approval by the organisation's process, technical and security owners.

Assign responsibilities that survive a staff change

Record named people and deputies for five responsibilities. One person may hold more than one role where the organisation permits it, but the decisions should remain clear:

  • Process owner: approves the business rules, acceptable interruption and manual fallback.
  • Technical owner: maintains the flow, investigates failures and controls tested changes.
  • System and connection owners: confirm the permitted identities, access and operation of each connected service.
  • Operational reviewer: checks expected work, unresolved exceptions and recovery results.
  • Change approver: authorises production changes and confirms the evidence required before release.

Co-ownership grants substantial control. Microsoft's cloud-flow sharing guidance says owners can view run history, edit the definition and delete the flow. It also distinguishes ownership from connection credentials: an owner cannot modify credentials for a connection created by another owner. Give access according to the agreed role, and record the connection dependencies separately.

Keep a short operating record for each production flow

The record should identify the environment, flow and solution where applicable; trigger; expected workload; business record identifier; source and destination systems; and current approved release. Include the owner and deputy, monitoring route, support contact and escalation route.

For each action that creates or changes something outside the flow, state how an operator will verify the outcome in the receiving system. Record which identity and connection it uses, without putting passwords, tokens or other secrets in the runbook. Document the approved way to pause incoming work, account for queued or in-flight work and resume operation.

Agree how quickly an exception needs attention according to its business consequence. A routine document notification and an action affecting a time-sensitive transaction may need different escalation. The runbook should name the decision-maker rather than assume that every technical owner can authorise every business correction.

Monitor expected work as well as failed runs

Start with the run history and the failed action's details. Microsoft's troubleshooting guidance explains how to identify a run and inspect its error. It also describes optional repair-tip emails. Verify that the intended owners can see the required information and receive the agreed notifications.

Microsoft documents Run after conditions, scopes and action retry policies for handling failures. These can support an agreed error path, but they require design and testing. Check that an error notification reaches its recipient even when the affected connection is unavailable, and that a failure in logging or notification becomes visible through another agreed route.

Also compare eligible source records with completed business outcomes. A process can be incomplete because an expected trigger never produced a run, or because the flow finished without the intended downstream result. Define the expected completion evidence, the review frequency and who investigates missing outcomes. A successful status alone is insufficient evidence for a business action that must be checked elsewhere.

Before retrying, establish whether the action already happened

A timeout or lost response leaves an important question open: did the receiving system complete the request? Microsoft's Retry pattern guidance explains that a service may process a request successfully but fail to return the response; repeating a non-idempotent operation can then produce an unintended second effect.

For each incident, use the source identifier, run details and receiving-system evidence to distinguish three outcomes:

  1. The action is confirmed complete. Preserve its destination identifier and investigate the remaining steps. Do not repeat the completed action merely because the flow failed later.
  2. The action is confirmed not complete. Correct the cause and use only the recovery route approved for that operation, after checking what the original run already did and what the recovery will repeat.
  3. The outcome is uncertain. Hold the affected business item and escalate to the system or process owner. Keep it unresolved until the evidence supports a safe decision.

The runbook should distinguish an automatic retry of an action, resubmission of a run and manually completing part of the process. Verify the actual scope of each recovery method. A reference number or a preliminary duplicate search does not by itself guarantee protection against duplicate processing, especially when other runs or users can act concurrently. Any destination-supported duplicate protection needs its own configuration and failure tests.

An illustrative recovery decision

Consider a fictional purchasing workflow that sends an approved request to an order system and then records the returned order reference. Assume the process owner has approved the workflow and its recovery rules. The receiving system creates the order, but the response is lost before the flow records the reference.

The operator checks the receiving system using the original request identifier and confirms the order exists. Under the approved recovery procedure, the responsible owner determines how to restore the missing link and complete any remaining steps without creating another order. If the operator cannot establish whether the order exists, the request stays in the exception queue and further submission is held for investigation.

This is an illustrative failure scenario, not a claim about a particular connector or client implementation. The recovery route and duplicate controls must be approved and tested for the actual systems before use.

Treat a staff departure as a controlled handover

Start by identifying the person's flows, ownership roles, connections and business responsibilities. Confirm the successor's permissions, licensing and ability to investigate runs before the planned handover. Follow the organisation's offboarding timetable and security decisions; do not retain a departing person's access as a workaround.

Microsoft's ownership-transfer guidance distinguishes solution-aware flows, whose primary owner can be changed, from non-solution flows, whose owner cannot be changed in place. It describes solution conversion or creating a replacement through supported copy or export/import routes, depending on the environment. Ownership can also affect licensing and request limits. Select and test the applicable route rather than assuming that adding a co-owner completes the transfer.

Review every affected connection separately. Confirm the approved replacement identity, destination access and authentication arrangements. Where a replacement flow is created, plan how the old and new triggers, waiting work and duplicate risks will be handled. Test the handover with representative authorised records, including a failure and an uncertain outcome, then reconcile the results.

If the original owner has already left, Microsoft's orphaned-flow guidance describes how an administrator with appropriate privileges can identify flows without a valid owner and assign a co-owner. That restores an ownership route; connections, permissions and business outcomes still need verification.

Retain the evidence and test the next change

Microsoft lists run retention in storage as 30 days for automated, scheduled and instant flows. Do not assume the standard run history will meet a longer business-record requirement. Agree what evidence to retain, where it belongs, who may access it and how its retention will be implemented in the actual environment.

An incident record should contain the business identifier, affected run, observed error, confirmed downstream state, owner decision, recovery action and verified final outcome. Keep unnecessary personal data and secrets out of alerts and logs.

Before releasing a change, test the routine route, an invalid input, a connection failure, a delayed or missing response, duplicate attempts and the manual fallback where relevant. Record expected and observed outcomes, including what reaches the connected system. Obtain the required approval and review the first production results. Support hours, response targets and ongoing maintenance must be agreed explicitly.

Discuss the operating requirements

SCSB's Business Process Automation service includes agreed workflow design, exception testing and handover to process owners. Bring the flow's purpose, systems involved, current owners and a recent failure or handover scenario so that the operating requirements can be scoped.

This article provides general technology and process information about Power Automate cloud flows. It does not guarantee uninterrupted operation, duplicate-free processing or regulatory compliance. Product capabilities, connector behaviour, licensing and tenant configuration must be checked for the proposed design. Security, business decisions and support commitments require the organisation's approval and agreed scope.

Microsoft and Power Automate are trademarks of the Microsoft group of companies.

Related service