At the end of an Agentforce proof of concept, the sponsor should hold a decision record: evidence that lets them choose go, fix and retest, or stop. One natural answer in a demo is easy to get. Evidence for how far real requests can be delegated has to be built on purpose. Production opens on the same logic. An agent answers only within the data it can reach and the access it is granted, so when a pilot stalls at the threshold, look below the prompt first.
Four questions the decision record answers
A PoC that ends with “we confirmed the potential” turns the next meeting into opinion against opinion. The platform team says it saw the feature, security asks about the data path, operations asks who handles a failure. The decision record answers those questions inside the PoC:
- Which requests may the agent handle?
- Which requests must go to a person?
- Which access boundaries apply to the data the agent reads and the actions it calls?
- When an error or an unexpected request appears, who approves which change?
In Salesforce terms, subagents split the scope, actions connect the work the agent performs, and instructions set its behaviour inside that boundary. The Agentforce DX testing documentation describes an agent test as sending an utterance and checking that the response behaviour matches what was expected. So the measure is routing, choice of action, handover and refusal, more than the tone of the answer.
The operations owner approves these criteria before the PoC starts. Criteria adjusted after the results leave a record where the design was fitted to the outcome. At Disneyland Paris in 2022, I audited the B2B processes and set a Salesforce target before the integrator was chosen. The same order suits a PoC: the client writes the go/no-go criteria before a partner builds the demo.
Pick one workflow and write the refusal path first
Put customer service, sales support and knowledge search in one agent, and a good result will not tell you what worked while a bad one will not tell you what to fix. Each workflow has its own data, access, exceptions and owner. One workflow is enough. For example, an employee looks up a policy in knowledge articles they can already access, the agent shows the source article, and it hands over to the owning team when it cannot be sure. That is easy to reverse, and it tests search quality, access, handover and logging at once.
Then write the refusal path before the success path, in one document:
- allowed requests: the request types handled and the inputs required;
- forbidden requests: what the agent does not answer or execute;
- handover conditions: who receives the request when information is missing, the wording is ambiguous or access must be checked;
- action boundary: which of read, suggest, record and external change is allowed;
- recovery: who approves the stop, fix and retest after a wrong answer or call.
Showing a query result and changing the state of an external system do not carry the same risk. Keep state-changing actions out of the first PoC; if one is essential, limit its target and design the human approval and the return path in advance. “It’s only the test environment” does not replace access design, and real personal data does not go into a PoC for convenience.
The test table carries two verdicts
Test cases cover normal requests, missing information, ambiguous wording, forbidden requests, requests outside access and action failures. For each case, record the input, the expected subagent, the actions allowed or forbidden, the mandatory elements of the answer and whether handover is expected. Write expected results as checkable conditions, such as “answer only from article X and hand over to the owning team if it has no answer”, rather than “a good answer” that each reviewer reads differently.
The first verdict is agent behaviour: right subagent, allowed tools, handover or refusal. The platform team can lead it. The second is operational fit: whether that behaviour matches the department’s policy, the split of responsibilities and the data access principles. The business owner and the security or data protection lead give it.
A failed case points to a layer. If the agent showed data the user may not see, check data access, user permissions, and action inputs and outputs before writing longer instructions. If it keeps handing over, check whether the knowledge source, the classification or the handover condition is too narrow.
The exit decision and the production conditions
The exit meeting ends in go, fix and retest, or stop. A go candidate has separated allowed from forbidden requests, held the action boundary in tests, shown a working handover and recorded a cause and an owner for every failed case. Hold signals go in the same document: unexpected actions called for a different reason in each test, no named data or business owner, success that depended on exception access or unrealistic test data, no way to confirm that earlier tests still pass after a fix.
A go does not open production by itself. Five conditions do.
| Condition | What the sponsor checks |
|---|---|
| Data | A list of the questions the agent must answer, with where each answer’s data lives and whether the agent can reach it. |
| Access | Which user context the agent runs in, and what that context allows. |
| Stop owner | The person who switches the agent off when quality drops, and the criterion, named before go-live. |
| Test set | The PoC set has become a regression set, rerun at every change. |
| Support path | The team that receives handovers exists, and its headcount and hours match the handover conditions. |
Not reaching the data it needs and reaching data it should not see call for different fixes. The first is permission mapping: a connection without it still leaves the agent blind. The second is the security model: a permission model designed around people did not plan for an agent as a new actor. Neither is solved by polishing the prompt or switching models.
Agree the personal data processing path with the security team at pilot design; a review started just before go-live can require new consent or a new subcontracting agreement, which does not fit in one sprint. If an integrator built the pilot, date the internal team’s takeover of operating rights at the same time. Splitting change rights after go-live is covered in Agentforce operating model.
What to check
- Did the operations owner approve written pass criteria before the PoC started?
- Can the exit document choose between go, fix and retest, and stop?
- Does each test case name the expected subagent, the allowed actions and whether handover is expected?
- Can you say in one sentence which user context the agent runs in?
- Is the person who can switch the agent off named before go-live?