Skip to contentSkip to contact
Sébastien Tang

Note Field notes

What a year of running three autonomous agents taught me

For a year, three AI agents ran my back office. The failures were silences, and the governance rules that caught them outlived the setup.

Author
, Program and Delivery · Seoul
Published
(updated )
Reading time
5 min

For a year, three AI agents ran the back office of my own practice, and none of them ever did anything dangerous. They went quiet. The audit agent was blind for ten consecutive days before I noticed, and the spec-drafting daemon had been off for 47 days with no alarm to say so. The rules that held were about what each agent could touch, how it asked for approval, and how I would notice that it had stopped.

The setup, and its limits

The work was the janitorial side of a one-person practice: research, pipeline hygiene, drafting, internal tooling. Three agents had deliberately unequal privileges, and I stayed the only judgment node. An orchestrator held the business context and drafted everything a person would read, and its output always stopped at a draft. A coding agent worked from written specs in an isolated branch, with no business data, and never committed to the main branch. A scheduled auditor compared the work board with the repositories and could change nothing; its only output was a pull request. They passed work through files, branches and board cards, so every hop left a trail a person could read weeks later.

My work is programme delivery and teaching Salesforce. This was my own back office, not a client deliverable, and its scale was one person and a work board. The lessons are the ones I now bring to the admins I teach about AI in Salesforce: who holds which rights, which data an agent can reach, and how people adopt a tool that acts without them. The same questions, applied to Agentforce, are in the Agentforce operating model note.

In September 2026 I retired the three-agent setup and folded it into a single assistant, with fewer moving parts. The rules came through intact because none of them depended on having three agents. They bind any component that acts while nobody is watching.

Autonomy fails silent

The audit agent’s job ran every morning and opened a pull request with its findings. The pull requests kept arriving, so the system looked alive. Inside each one there was nothing: the agent could no longer reach the data it was meant to audit. The same review found the spec-drafting daemon, switched off by a configuration batch 47 days earlier. Both times the machinery worked. Feeding and watching failed.

Most writing on agent risk worries about harmful actions. What I met was absence, and no model reports its own absence. Every scheduled component needs a heartbeat and an alarm when the heartbeat stops. A report that arrives empty is not a heartbeat, so the check has to read the content. This belongs in the first release.

Privilege shrinks as autonomy grows

The rule under everything: the more freely an agent acts, the less it may touch. Its sharpest form is the exfiltration triangle. An agent that combines attacker-controlled input, autonomous egress and access to sensitive data is one prompt injection away from leaking that data, whichever model runs it. Any two legs are acceptable; all three are not. So each agent had one leg removed structurally, by a missing credential or a blocked network path, never by a line in a prompt asking the model to behave. The coding agent read untrusted repository content and could push branches, so it held no sensitive data.

Two incidents sharpened the rule. An inbox triage prototype first stripped personal data before the agent saw a message, leaving categories and flags. A live test showed that nobody drafts a useful reply from flags. The fix flipped the trade: the agent read full messages, and the egress leg was cut instead (allowlist proxy, no send credential, no fetch tool). Later, a security review found shell tools on a headless composer that summarised untrusted inbound text: a plausible path from a malicious message to code on the host. It was rebuilt as a pure text composer, with a wrapper script doing all input and output. A headless agent over untrusted input gets no shell and no permission overrides.

Approval is a mechanical contract

No model ever sent an outbound message. Anything client-facing ended in a queue, and a model-free script sent it only after a person had set an approved flag. The send step was deliberately too simple to be talked into anything.

Every proposal, a code change or a board correction, arrived as a pull request. Merging meant approved; closing with a comment meant rejected. That one convention gave the audit trail, the diff review and the ability to hold one change while approving the rest, all inherited from mature tooling.

The contract only works when a task can be checked mechanically. A routing rule decided what the coding agent could receive: if acceptance could not be written as tests passing, a green build and a search returning the expected count, the task stayed with a supervised agent. When early tasks came back refused, fixing the specs fixed the refusals. The model did not change.

Arm by blast radius

Every automated leg was classed as sensing (reads only), drafting (produces something a person approves) or executing (acts with no person in the loop). Sensing and drafting ran all day and could not touch the outside world. Executing sat behind one arm file in the repository: present meant armed, deleting it stopped everything, and its state was visible in version control. Kill switches were files rather than config entries because a file is easy to diff, easy to search and hard to misread during an incident.

The gate earned its place once: the auditor proposed moving a work item to ready-to-ship on a repository that deployed on merge. Such repositories were on the forbidden-action list, so the proposal came to me. The card was legitimate; on another day the same move is an unreviewed production deploy.

The widest automation, board-wide actions across every work item, shipped behind its own switch, ran in dry-run and stayed off. The incidents came from components that were armed. Blast radius set the arming order, and some switches deserved to stay off longer than enthusiasm wanted.

What to check

  • One presence file arms every leg that acts without a person, with a kill switch per leg beside it.
  • No model holds a send credential. A model-free step sends on a flag a person sets.
  • Proposals arrive as pull requests: merge approves, close rejects, nothing is approved verbally.
  • Every scheduled component has a heartbeat and an alarm on silence, and the alarm reads content as well as arrival.
  • Each agent has one leg of the exfiltration triangle removed structurally.
  • Instruction files, identity and voice files, and agent memory are writable by no agent.

Contact

Is this on your table?

Tell me where your programme or your team stands, and what you need to decide.

Request a call

WorkshopsTrainer profile