essay · october 2026

Agents know how. Nobody told them how far.

By Viktor Berthelius · 6 min read


Ask a capable coding agent to speed up a slow report and you will usually get a faster report. Ask it to make reporting ten times better and you may well get the same report, a little faster, with a few more tests. The work expands to fill the instruction and stops exactly there. The agent is not being lazy. It is being safe: stay inside the lines, break nothing, ship the increment. Nobody told it where the lines end.

This is not a capability problem. The agent that hands you a tidy patch could, in principle, redesign the module behind it: a cleaner interface, three fewer dependencies, tests that would have caught last quarter’s bug. What it lacks is not skill but permission: a signal that the bigger move is in scope, that the team wants it, and that the bar sits somewhere above “does not regress”.

Capability without direction defaults to caution.

The files we already write

The first wave of agent context was operational. AGENTS.md and CLAUDE.md tell an agent how to work in a repository: which commands to run, which gates to pass, which conventions to keep. That was the right first move. An agent without operational context breaks builds.

Others followed. SOUL.md describes who the agent is: its persona, values and tone. DESIGN.md hands it the design tokens and the reasons behind them. A constitution, where a project keeps one, draws the lines it must never cross. Each answers a real question, and each earns its place.

None of them answers the question that decides whether an agent ships a competent patch or an ambitious leap:

Where is this project going, and how far are we willing to push to get there?

Without an answer, the safe reading wins. Not because safe is right, but because safe is the only reading that assumes nothing about ambition. The agent cannot tell whether you are building the best tool in your category or keeping the lights on for one more quarter. So it hedges. It patches. It increments.

Same task, two contexts

Take a product two years in. A status field has grown seven values. Permissions started as a boolean and picked up string checks along the way. The reporting layer queries the same table four times per request. None of it is a crisis. All of it slows the team down.

Ask an agent to “clean up the reporting layer” and a reasonable one will merge the four queries into two, add an index and stop. A good patch, even.

Now give the same agent four lines from a NORTH.md:

North Star. Month-end reporting is the reason customers stay.
The Bar. p95 under 200 ms on every report screen.
Trade-off Defaults. Clean design over backward compatibility on internal APIs; flip when an external consumer exists.
Reversibility. The reporting schema is internal: a two-way door.

Now the agent can see that reporting is load-bearing, that two queries will not clear the Bar, that nothing outside the team depends on the current shape, and that a redesign is cheap to undo. It can propose the bolder move, and it knows the move is allowed. The agent did not get smarter. The question changed, from “what is safe to do?” to “what is right to do here?”


Not bigger. Calibrated.

Timidity is only half the failure. Give an agent more autonomy without direction and the opposite happens too: the sprawling rewrite, the feature nobody asked for, the defensive abstraction nobody will ever need.

Two agents can both satisfy every linter, compiler check and test in your pipeline. One moves the architecture towards its intended shape. The other turns the repository into an accumulation of reasonable-looking detours. CI cannot tell them apart. Tests prevent broken builds; they do little to prevent drift.

So NORTH.md cuts both ways. Asymmetric Bets say where to go 10x and, just as explicitly, where 1.1x is the right call. Anti-Goals decline the tempting detours before anyone proposes them. Trade-off Defaults settle the recurring arguments and name the one condition that reopens each. The point is not ambition everywhere. It is ambition in the right places, and restraint in the rest.

This matters more as agents run longer and in parallel. When there is no one to ask mid-run, every decision nobody wrote down becomes a guess, and parallel sessions guess differently.

Without direction, more agent throughput tends to produce more work in progress, not more progress.

Seven sections, one job

A NORTH.md has seven sections, and each answers a question an agent would otherwise have to guess. The North Star names the destination in one sentence. The Bar says what done means, in terms a gate can check. Asymmetric Bets split effort between 10x and 1.1x. Anti-Goals list what the project refuses, each with a reason. Trade-off Defaults record decisions already made, and when they flip. Ambition Triggers are short questions that escalate scope. Reversibility marks which doors are one-way.

Together they precompute the decisions that would otherwise be made implicitly, and inconsistently, in every agent session.

Why plain markdown

AGENTS.md and CLAUDE.md proved something simple: a plain file at a known path is a durable way to put context where humans and agents will both find it. It is versioned by git, reviewed in pull requests and readable without tooling. A change of direction goes through code review like any other change.

A YAML schema would validate better and persuade worse. Rules without reasons do not travel. An agent can follow “safety over speed on fiscal code” to the letter, but only the reason lets it apply the rule to a case the author never imagined. Prose carries reasons. That is the job.

llms.txt applies the same idea to websites: a plain file at a predictable URL. Conventions compound. Once enough projects put a file in the same place, tools know where to look.

What it is not

Not a roadmap. No features, no dates. Those belong in the tracker, where they can change every week without touching a file that is meant to stay put.

Not OKRs. OKRs are quarterly and measured on a cycle. A NORTH.md changes rarely, and deliberately.

Not a vision statement or a manifesto. Those are written to inspire or to persuade. A NORTH.md is written to decide. The test is not “is this inspiring?” but “does this change what an agent does when it has two reasonable options in front of it?”


Two screens, on purpose

The format is short by design. A NORTH.md that runs past two screens has lost the plot: if the ambition cannot be stated in that space, it is not clear enough to steer anyone. The discipline of the format forces the discipline of the thinking.

That is also the real cost. Writing a good NORTH.md means deciding what you actually believe about your destination, your bar and your trade-offs, instead of leaving them as shared, unspoken assumptions that agents, and new team members, have no way to reach.

The spec reads in five minutes. A first NORTH.md takes ten. A good one takes longer. The aim is simple: every session after it starts from a better default. Not because the agent changed, but because someone finally wrote down how far.