A Salesforce project in distress rarely breaks at random. Underneath the complaint sit the same conditions: a data model that drifted from its design, automation layered by several teams with no agreed order, and a deployment path that slid into manual changes in production. Ninety days is enough to regain control if the work runs as a sequence of gates, each with a decision owner and an exit deliverable. Recovery attempts that skip the sequence fix symptoms and leave the causes in place.
The complaint hides the cause
Sponsors describe what they saw: the last release broke production, the integrator went quiet, sales went back to spreadsheets. The useful question is why one event could do that much damage. The usual answer is risk that built up where nobody could see it, because nobody held a current picture of the org.
So a rescue does not start with a fix. A team that starts building in week one builds on a foundation it has not read, and each change it ships adds to the uncertainty it was brought in to reduce. The conditions also have to be handled together: fixing automation while the data model is still broken only moves the failure somewhere else.
Three gates with owners
The phases follow a dependency chain. You cannot stabilise automation before you know the data model, and you cannot restore deployment confidence before automation is stable.
| Gate | Window | Main decision | Decision owner | Exit deliverable |
|---|---|---|---|---|
| 1. Stop the bleeding | Days 1-30 | Freeze new configuration | Sponsor | Metadata and automation inventory |
| 2. Stabilise | Days 31-60 | Which automation survives on each object | Process owner, with the design authority | Conflict-free automation, a working deployment path |
| 3. Restore confidence, hand over | Days 61-90 | Go or no-go on each controlled release | Release owner | Current-state document, automation inventory, data dictionary |
Gate 1: stop the bleeding
Read before touching anything. Extract the full org metadata, map every active Flow, Apex trigger and leftover Process Builder against the objects it touches, then check the debug logs to see which ones actually run in production. Audit the data model for broken references: lookups to record types that no longer exist, validation rules that contradict each other, fields nobody fills.
The containment action is a freeze on new configuration. Nothing reaches production without a written impact note reviewed by a small change board. The freeze is political, so the sponsor owns it and announces it. A recovery team cannot enforce it alone.
Gate 2: stabilise
Resolve conflicts before improving anything. The most dangerous pattern is competing automation on the same object. Two Flows that update Opportunity Stage on the same save, with no defined order, will race: the result depends on which one runs last. Users report an intermittent bug, and more debugging will not end it, because the behaviour is built into the design.
The fix is one automation entry point per object and trigger event. Choosing which logic survives is a business decision, so the process owner makes it with the design authority, and the choice is written down.
Data model repairs run in parallel, in a fixed order: relationships that break reporting first, because reporting is what stakeholders see, then field-level issues, then validation logic. This gate also sets up the deployment path: a sandbox that reflects production, and version control as the reference for what the org should contain.
Gate 3: restore deployment confidence and hand over
Prove the stabilisation held. Take a bounded but real change, run it through the full path (sandbox, validation, production release with a rollback plan) and record the outcome. Do it three times, each a little more complex. By the third cycle the team has shown the path works and has the habit of using it.
Governance stays small: a decision record, a metadata ownership matrix naming who owns which objects and automations, and a monthly review. The gate closes when the documentation is handed to a named owner on the client side.
What recovery attempts get wrong
Scope creep. Once the org works again, the business asks for features. The recovery team wants to show value and starts building, and the structural problems stay. The defence is an explicit scope gate: gates 1 and 2 are recovery only. Gate 3 can include one or two visible changes, picked because they exercise the new deployment path rather than for their business value. The sponsor is the one who says this to the business.
A date instead of a gate. Ninety days is a planning window. If gate 1 shows deeper data model damage than expected, gate 2 extends. Compressing the work to meet a date produces an org that looks stable and fails again under load.
Documentation left for later. Distressed projects are almost always undocumented. A team that leaves without a current-state document and an automation inventory hands the next team the same hidden risks, and the cycle restarts. Put these documents in the plan as exit deliverables from day one.
How to measure progress
“The org feels better” is not a measure. Each gate needs operational markers:
- End of gate 1: a complete metadata inventory exists, all active automation is documented with its execution order, and no unplanned production change happened in the last two weeks.
- End of gate 2: no known automation conflict remains, the deployment path runs end to end without a manual step, and broken references are resolved for every object in scope.
- End of gate 3: three controlled releases are done, at least two decisions are recorded, and a short stakeholder survey shows progress against the baseline taken in week one.
The markers are operational on purpose. Business outcomes take longer than 90 days to show; at this stage you are measuring whether the conditions for reliable delivery are back. For the diagnostic itself, the org review checklist covers what to read, and the note on technical debt covers what to bring under version control first.
What to check
- The change freeze has an owner, and the sponsor announced it.
- Every object has a single automation entry point per trigger event.
- The team can deploy a change end to end without a manual step in production.
- The current-state document and the automation inventory are listed as deliverables, with a named owner after handover.