Module 8 — The AI Agent: Governing It · Lesson 8.1
Autonomy Levels and Safe Mode
Off, Assess, Safe and Full — and the two-sided ceiling that decides your effective level
~12 min
What you'll learn
- Define off, assess, safe and full, and what each permits
- Explain how the workspace ceiling and personal opt-in combine
- Describe what safe mode does to the agent's available tools
- Widen autonomy on evidence rather than on hope
The question 'how much should I let it do?' has no universal answer, which is why Kavanah does not pick one. What it provides instead is four clearly-defined levels, a ceiling admins control, and a record you can read to decide whether to move up. This lesson is about using that properly rather than leaving it at the default forever.
The four levels
Off hides the feature. No assessment, no execution. The agent is a conversational participant only.
Assess shows whether Kavanah COULD do a task, with its reasoning, but never acts. This is a genuinely useful setting for a sceptical team, because it produces evidence — you can watch what it would have done for a fortnight with no risk at all.
Safe lets the agent complete tasks autonomously using internal, non-destructive actions only. It creates and updates tasks, writes comments, prepares drafts. It will never send email, post externally, or delete anything.
Full lets it use any tool the task needs — but each risky action follows your action policies, which by default means it is parked for explicit per-action approval before it executes.
The important thing to notice is that Full is not 'no oversight'. Full plus the default policy means the agent can attempt anything and you approve the consequential parts.
What safe mode actually does
This is worth being precise about, because it changes how the agent behaves rather than just what it is allowed to do.
At safe, risky tools are REMOVED from the agent's toolset entirely. It is not offered the ability to send email and instructed not to use it — it does not have it.
That matters for two reasons. It is a much stronger guarantee than an instruction, because there is no phrasing that talks a model into using a tool it does not possess. And it means the agent's own account of its capabilities is accurate at that level: when it says it cannot send that email, it is describing its toolset rather than declining.
The same stripping applies to features nobody has opted into. A workspace that has not enabled the remote browser has an agent with no browser tools, which is why the experience is byte-identical to one where the feature does not exist.
The two-sided ceiling
There are two settings, not one.
An admin sets the WORKSPACE ceiling — the highest level anyone in the workspace may reach.
Each person sets their own opt-in level for themselves.
Your effective level is the LOWER of the two. So a workspace ceiling of Safe means nobody runs at Full regardless of their personal setting; and a personal setting of Assess means you run at Assess even in a workspace whose ceiling is Full.
This is the right shape because the two decisions are genuinely different. The ceiling is a risk decision about the organization. The opt-in is a comfort decision about one person's work. Neither should be able to override the other in the wrong direction.
A widening path
Four steps, each with an exit criterion, rather than a leap of faith.
Start the workspace ceiling at Safe. Live there for at least a fortnight. Read what it did — the ledger, the task activity — rather than asking people how it felt.
When you can look at two weeks of safe-mode actions and find nothing you would have wanted to stop, raise the ceiling to Full and leave the action policies at their default, which parks the risky categories for approval. You have now widened what it can attempt without widening what happens without you.
Work the approval queue for a few weeks. That queue is data: the actions you approve without hesitation, every time, in the same category, are candidates for a looser policy.
Only then, and only per category, move a policy from approve to a budgeted allowance or an outright allow. Never do this in bulk. The next lesson is entirely about that decision.
Where to set it
Both settings live under Settings → AI Agent. The workspace ceiling is admin-only; the personal opt-in is yours.
One piece of advice for admins: tell your team when you change the ceiling. A ceiling change silently alters what everyone's agent will do, and someone discovering that by surprise is a bad first experience of a feature you want adopted.
Set autonomy deliberately
- 1
Set the workspace ceiling to Safe
Settings → AI Agent, admin only. This is the right starting point for almost every workspace.
- 2
Try Assess for a week if your team is sceptical
It shows what it WOULD do with no risk, which produces evidence rather than argument.
- 3
Read two weeks of safe-mode actions before widening
The audit log, not a vibe check. If nothing there would have needed stopping, you have your evidence.
- 4
It silently changes what everyone's agent does. Surprise is a bad first experience of a feature you want adopted.
What to watch
- Effective level distribution
- What level people are actually running at, versus the ceiling.
- Healthy signal: Most people at the ceiling. A big gap means people do not trust it yet, and that is a conversation rather than a settings change.
- Would-have-stopped rate
- How many of the agent's autonomous actions, on review, you would have wanted to prevent.
- Healthy signal: Zero for a sustained period before widening. This is the only honest input to the widening decision.
Key takeaways
- ·Off hides it; Assess shows what it would do; Safe acts with internal tools only; Full acts with policy-governed approval.
- ·Safe mode REMOVES risky tools rather than instructing against them — there is no phrasing around a missing tool.
- ·Effective level is the lower of the workspace ceiling and your personal opt-in.
- ·Full plus the default policy is not 'no oversight' — it is attempt anything, approve the consequential parts.
- ·Widen on evidence from the ledger, one category at a time, never in bulk.
Next: the policies themselves — seven categories, four modes, and the approval queue where it all lands.