Le secrétaire de Fernand

Enforce Architecture, Don't Trust Intent

Enforce Architecture, Don't Trust Intent

Why load-bearing architecture rules must become executable checks rather than trusted intent.

Central idea: An architecture is not what a team says it respects. It is what the system makes visible and testable. The rules that carry the system must be enforced as executable checks, because intent, however sincere, cannot refuse a commit.

The short version: 2 min Full article: 17 min

The short version

Every team has an architecture it believes in: the diagram, the ADR, the folder layout, the onboarding talk. And almost every team discovers, months later, that the codebase quietly stopped following it. Nobody decided to abandon the architecture. It eroded, one reasonable-looking commit at a time.

The problem is not discipline. The problem is that a diagram cannot refuse a commit. Documentation, conventions, and collective memory are declarations of intent, and intent has no enforcement power. An architecture becomes real only when deviation becomes visible, and it becomes durable only when deviation becomes costly.

The way out is to express the load-bearing rules of the architecture as executable checks: import constraints, module boundaries, dependency directions, contract presence, wiring completeness. Each check is small. Together they turn the architecture from a story the team tells into a property the system verifies.

In three ideas

  1. Intent decays silently; checks fail loudly. A boundary that lives in a document erodes without a trace. A boundary expressed as an import rule leaves a red mark the moment it is crossed.

  2. The dangerous violations are absences, not presences. A forbidden import is easy to spot. A handler that is written but never wired, a port without an adapter, a validation schema that exists but is not called at the boundary: these pass typecheck and tests. Enforcement must check for missing required things, not only for present forbidden things.

  3. A good rule protects a concept, not a convention. "No file over 120 lines" enforces taste. "The domain must not depend on the infrastructure" enforces the system. Rules that protect real concepts survive; rules that encode preferences teach people to ignore the checker.

Takeaway

Do not ask whether the team trusts the architecture. Ask what would turn red if the architecture were betrayed today. Whatever would not turn red is not architecture yet. It is hope.


Full article

The previous article ended on a promise. It argued that feedback loops become powerful when the code gives them explicit structure to observe, and it showed a small : three states, three events, nine cells, each cell classified. Then it said that the interesting case is the large one, the protocol with dozens of states and events, where the matrix stops being a local check and becomes a genuine audit surface.

This article keeps that promise. But to get there, it has to start with a more uncomfortable observation.

Most architecture is not enforced at all.

Every codebase has two architectures

There is the architecture the team declares: the diagram in the wiki, the decision records, the layered drawing from the last onboarding session, the sentence "the domain never touches the database directly" that everyone nods at.

And there is the architecture the code actually has: the real import graph, the real dependency directions, the real places where validation happens or does not, the real owners of each mutation.

On day one, the two coincide. Then the project lives. A deadline pushes a database call into a component "just for now". A helper migrates from a leaf module into shared utilities and drags its dependencies along. A new feature imitates the wrong example, and its shape becomes the next feature's example. Each step is locally reasonable. None is flagged, because nothing is watching.

This is , and the essential thing about it is that it is silent. Nobody writes an ADR titled "we hereby abandon the layering". The declared architecture stays pristine in the wiki precisely because it lives in the wiki. Documents do not compile, so they cannot break.

Software reflexion models, described by Murphy, Notkin, and Sullivan in the 1990s, were built around exactly this gap: take the high-level model the engineers believe, extract the actual structure from the source, and show the differences. What they found, over and over, is that the two views diverge in every long-lived system, and that engineers are consistently surprised by where.

The lesson is not that teams are careless. The lesson is that alignment between intent and code is not a stable state. It is a controlled process, and without a controller it does not happen.

Trust is not a maintenance strategy. Feedback is.

Every codebase has two architectures

The declared architecture lives in documents and cannot refuse a commit; the actual architecture lives in the code and drifts one reasonable change at a time. An executable rule is the only thing that compares the two on every change.

Loading diagram…

What an executable architecture rule looks like

The alternative to trusted intent is the enforced rule: a check that runs against the codebase and fails when a load-bearing property no longer holds. The evolutionary-architecture literature calls these . Whatever the name, the useful ones share an anatomy.

A rule has three parts.

First, the concept it protects. Not "imports from that folder are forbidden" as a free-floating prohibition, but the system concept behind it: this boundary exists so that the domain stays independent of the delivery mechanism. The concept is the why, and it belongs in the rule, because a rule whose reason is lost becomes ritual, and ritual gets deleted.

Second, the signal it detects. Something mechanically checkable: an edge in the import graph, a pattern in the AST, a name that does not match a required shape, a schema that is declared but never referenced. The signal is where the rule touches the code.

Third, the diagnostic it produces. And this is where most tooling underdelivers. A useful diagnostic does not only point at a line. It points back at the concept: not "import of pg forbidden here" but "this module is inside the domain boundary, and the domain must not depend on infrastructure; the dependency belongs behind a port". The previous article called this . A rule that attributes teaches; a rule that only rejects trains people to appease it.

Concretely, the same boundary rule can be enforced at several layers, and the layers differ in strength, not just in tooling:

  • an import-lint rule (domain/** must not import infrastructure/**), cheap and immediate;
  • compiler-level module boundaries or project references, which make the forbidden edge fail the build;
  • a check over the extracted dependency graph, which can see transitive paths that file-level rules miss;
  • packaging, the strongest form, where the domain literally cannot name the infrastructure because it does not depend on it.

The progression matters more than the tool names. Each step moves the rule from detected after the fact toward impossible to express, and that movement is the subject of the next article. Tools like ArchUnit in the JVM world or dependency-cruiser in the JavaScript world exist precisely to make the middle steps cheap.

On the mechanics, the useful default is simple: rules should be static (decidable from the source, no execution needed), atomic (each rule checks one property at one site, so a failure is directly actionable), and triggered (they run on every change, in the editor or the CI gate, not in a quarterly review). Static, atomic, triggered checks are cheap enough to run always, and checks that run always are the only ones that catch drift when it is one commit old instead of one year old.

Forbidden presences and missing parts

Here is the distinction that separates real architectural enforcement from a pile of lint rules. Architecture rules come in two polarities, and they fail in opposite ways.

A forbids a presence: this edge must not exist, this layer must not be imported here, this global must not be touched from that module. When a prohibition is violated, the evidence is concrete. There is a line to point at. Prohibitions are the easy half, and most existing tooling lives there.

An requires a presence, which means its violation is an absence: every port must have a wired adapter; every externally reachable handler must validate its input; every event this module emits must have a registered consumer; every state machine must handle every event it can receive. When an obligation is violated, there is no line to point at. The evidence is a hole.

Two polarities on one boundary

A prohibition fails on an edge that exists and must not, here core importing web; an obligation fails on an edge that must exist and does not, here a port with no adapter wired to it. The first has a line to point at, the second has a hole.

Loading diagram…

The two polarities matter because ordinary feedback is blind to the second one.

A missing adapter typechecks. A handler that is written but never registered passes its unit tests, because the tests call it directly. A validation schema that exists but is not invoked at the boundary satisfies every local check, because locally there is nothing wrong. The system looks done and is not wired. Nothing turns red, because every tool in the local loop examines what is there, and the defect is what is not there.

This is where architectural enforcement earns its keep over compile-and-test. Checking an obligation requires a model of the required shape: something must say "a port is expected to have an adapter" before the absence of one can be detected. That model is exactly the declared architecture, finally put to work. The reflexion-model authors had a name for this polarity too: divergence for the forbidden presence, absence for the missing part.

One more property falls out, and it is practical. A prohibition can be checked at every moment of a change, because a forbidden edge is wrong the instant it appears. An obligation can only be checked at completion, because a required part is legitimately absent while the feature is still being assembled. Enforce obligations too early and you drown builders in false alarms; skip them and you ship the unwired handler. Prohibitions are checked at every step; obligations are checked when the work declares itself done.

The disposition matrix, at protocol scale

Now the promise from the previous article.

Recall the shape: a state machine where every (state, event) pair is classified with a disposition. Not just the transitions you allow, but an explicit verdict for every cell: Handled, Ignored, Stale, Rejected, Unexpected. The type system makes the matrix total, so an unclassified pair does not compile. At three states by three events, this is a nice local trick. The type checker verifies it while you type.

Take the smallest real protocol on this site: the lifecycle of an article. Four states, a handful of events, and already the question the matrix forces: what does publish mean while the article is still a draft under review, and what does a new edit mean once it is published?

The lifecycle of an article

An article moves from draft to in review to published, and from published to archived; edits return it to draft. The matrix classifies every other state and event pair too, not only these transitions.

Loading diagram…

The diagram shows the transitions someone bothered to draw. The matrix is the diagram plus every cell the diagram leaves blank, each with a verdict, and the type makes the blank cells illegal:

TypeScript

type State = "draft" | "inReview" | "published" | "archived";
type Event = "submit" | "reject" | "approve" | "edit" | "archive";
type Disposition = "Handled" | "Ignored" | "Rejected" | "Unexpected";

// Every (state, event) pair must be classified, or this does not compile.
export const dispositions = {
  draft:     { submit: "Handled", reject: "Rejected", approve: "Rejected", edit: "Handled", archive: "Ignored" },
  inReview:  { submit: "Ignored", reject: "Handled", approve: "Handled", edit: "Rejected", archive: "Rejected" },
  published: { submit: "Rejected", reject: "Unexpected", approve: "Ignored", edit: "Handled", archive: "Handled" },
  archived:  { submit: "Rejected", reject: "Unexpected", approve: "Unexpected", edit: "Rejected", archive: "Ignored" },
} satisfies Record<State, Record<Event, Disposition>>;

Twenty cells, and the reducer that implements them is a separate artifact. Nothing above proves that the reducer agrees with the matrix; that is the oracle's job, and it can be generated, one case per cell:

TypeScript

for (const [state, events] of Object.entries(dispositions)) {
  for (const [event, disposition] of Object.entries(events)) {
    test(`${state} × ${event} → ${disposition}`, () => {
      expect(classify(reduce(state, event))).toBe(disposition);
    });
  }
}

Scale it up. A real protocol (a payment flow, a document lifecycle, a session manager, a sync engine) easily reaches dozens of states and dozens of events. Say thirty states and forty events: twelve hundred cells. At that size, three things change in kind, not just in degree.

First, the matrix stops being readable at a glance and becomes a reviewable artifact of its own. Nobody holds twelve hundred cells in their head. But the matrix can be rendered, diffed, and reviewed like a specification, because that is what it is: the complete behavioral contract of the protocol, with no gaps by construction. A pull request that changes three cells is showing you, precisely, which situations changed meaning.

Second, the matrix becomes an audit surface in the exact sense of this article. It is static: declared in the source, checkable without running anything. It is atomic: each cell is one independently verifiable claim about one situation. It is triggered: totality is re-verified on every change, because it is enforced by the compiler, and the day someone adds a state or an event, every cell that decision creates must be classified before the code compiles again. The matrix is a fitness function over the protocol: not "does this code run" but "does every signal this protocol can receive, in every situation it can be in, have a considered meaning".

Third, the polarity analysis applies inside it. Most cells are prohibitions in spirit ("this event in this state is a protocol violation, page someone"), and the compiler checks them for free. But the correspondence between the matrix and the reducer that implements it is an obligation: the matrix claims saving × segment-ready → Stale, and something must verify the reducer actually treats it as stale. The previous article already conceded this: totality is checked by the compiler, agreement is checked by a test. At protocol scale that test is worth generating from the matrix itself, one case per cell, so the two artifacts cannot drift apart silently. The matrix is the declared architecture of the protocol; the generated tests are its enforcement.

This is the pattern of the whole article in miniature. The team's intent ("we handle late events gracefully") became a structure (the matrix), the structure became checkable claims (cells), and the claims are enforced by two loops (compiler for totality, generated tests for agreement). Nothing about the protocol's correctness is trusted to memory.

Enforce the intent, not the letter

Every enforced rule creates an incentive to satisfy it in the letter and betray it in spirit. This is not cynicism; it is what deadline pressure does to any constraint. An enforcement system that ignores it will be gamed into uselessness.

A concrete example, because the move only makes sense once you see what triggers it. Two rules guard the design system. One says the identity of the product, its colors and spacing tokens, lives only in design-system components: a product component may compose Button and Card, never write bg-brand-600 itself. The other says design-system components are product-agnostic: nothing under ui/ imports product hooks, product types, or a product decision. Now a product component under a deadline grows a hand-written bg-brand-600. The style rule fires. The cheapest repair is not to extract the visual part; it is to move the whole file into ui/, where identity classes are allowed. The style rule goes green. The file is in the right place. The architecture is still violated, because the component still imports product hooks, still knows the product's types, still owns a product decision. Only its address changed. The honest repair was a split: the visual part becomes a design-system component, and the product logic stays where it was and composes it.

The defense is to bind rules to the concept rather than to its cheapest proxy, and to bind them in pairs, so that gaming one lands on the other. Here the second rule is the concept itself, and it is a plain import check: nothing under ui/ may import from a product scope. The moved file trips it on the same commit. A good rule asks who owns the responsibility, not only where the file lives. "Is the file under ui/" is a . "Does this module import product logic" is closer to the concept. "Does anything outside the owning module mutate this state" is the concept itself. Proxies are cheaper to check, and a mature ruleset uses them, but every proxy rule should have the concept written on it, so that when the proxy is gamed, the gap is visible and the rule can be sharpened.

If the code follows the folder rule but violates ownership, the audit should still fail. When it does not, the fix is not to shame the developer who gamed it. The fix is to thank them for the free penetration test and encode what they found.

A diff tells you what changed. An audit tells you what the change did to the system.

Code review, as practiced, is reading deltas. The reviewer sees the lines that changed, plus a few lines of context, and reconstructs the consequences in their head. For local properties this works. For systemic properties it cannot: the information is not in the diff.

A three-line diff can cross a boundary that took a year to establish. Adding one import is a one-line change whose meaning ("the domain now depends on the delivery mechanism") exists only at the level of the whole graph. Deleting a registration line is a one-line change whose meaning ("this handler is now unreachable") is an absence nowhere visible in the diff. The reviewer would need the entire dependency graph, the entire wiring table, and the entire protocol matrix in their head to see what the diff did to the system. No reviewer has that, and in the era of agent-produced changes, no reviewer has the time to fake it either.

This is the division of labor the whole series has been building toward. The diff answers "what changed". The enforced architecture answers "what did the change do to the system": which boundaries it crossed, which obligations it discharged or broke, which cells of which protocol changed meaning. The human review is then freed to do the one thing neither loop can: judge whether the change is a good idea.

What enforcement costs, and where it goes wrong

Honesty section, because this idea fails in known ways.

Old codebases are full of violations, and a rule that fails on all of them will be turned off by Friday. The workable pattern is the ratchet: record existing violations in a , accept them, and fail only on new ones. Old drift is debt to be paid down deliberately; new drift is blocked at the door. Tools across ecosystems converge on this shape because nothing else survives contact with a real codebase.

Rules that encode taste corrode trust in rules that encode structure. Every false positive spends credibility. A ruleset earns the right to block commits by being right about things that matter, and "matters" has a test: does this rule protect a boundary, a contract, an invariant, a responsibility, or a protocol? If it protects a preference, it can be a formatting convention, but it should not fail a build. The checker that cries wolf trains the team to reach for the override, and the override, once habitual, swallows the real violations too.

Every exception must stay visible. Real systems need escape hatches; the architecture that admits no exception is a fantasy. But an invisible exception is architectural debt with the receipts destroyed. The workable form is the annotated, dated, reasoned exemption that the audit reports as an exemption, so the list of places where the architecture is suspended is itself an inspectable artifact.

Enforcement can freeze a wrong architecture. The checks enforce the declared model, and the declared model can be wrong, or can become wrong as the product changes. Enforcement is not a substitute for the judgment that the model deserves enforcing; it assumes it. When the model must change, the rules must change with it, deliberately, in a reviewable commit. That is not a weakness of the approach. A visible, versioned, contestable architecture is precisely what makes deliberate change possible; you cannot renegotiate a rule nobody can see.

And a perfect audit that never ships protects nothing. A useful approximate rule today beats a complete rule system next year. Start with the two or three boundaries whose erosion would hurt most, enforce those, and grow the ruleset the way the architecture grew: incrementally, under feedback.

Agents inherit the architecture you enforced, not the one you meant

The multi-intelligence argument runs through this series, and here it becomes sharp.

A human developer absorbs the declared architecture through channels an agent does not have: the onboarding session, the hallway correction, the memory of the incident that motivated the boundary. An agent working in your codebase has the code, the types, the failing checks, and whatever documents happen to be in context. For an agent, the unenforced part of your architecture might as well not exist, and agents produce plausible code fast, which means they produce plausible violations fast. An agent will happily add the missing import that makes the tests pass, straight across a boundary it has no way to know is sacred.

But turn the same fact around. An agent takes an enforced rule more seriously than most humans do, in one mechanical sense: it cannot argue with the gate, and a red check with a diagnostic that names the concept and the correction path is exactly the feedback an agent converts into a fixed commit. Enforced architecture is not just protection against agents. It is the interface through which agents can be given the architecture at all: the intent, compiled into a form that survives contact with a producer who was not in the room.

The teams that get leverage from coding agents will not be the ones with the best diagrams. They will be the ones whose diagrams are executable.

Conclusion

The first article argued that code is a system, not prose. The second argued that what carries the system must be encoded, not implied. The third argued that explicit structure is what lets feedback attribute failures instead of merely detecting them. This article closes the loop on all three: the architecture itself, the largest structure of all, must be held by the same discipline. Declared, encoded, and then enforced, because a rule that nothing checks is an implication, and the second article already said what happens to those.

An architecture that lives in documents is a belief. An architecture that lives in executable rules is a property. Beliefs erode one reasonable commit at a time; properties fail loudly and get fixed.

Do not ask whether the team still believes in the architecture. Ask what would turn red if it were betrayed today.

And notice what the best rules in this article had in common: the boundary that could not be expressed, the state that could not be left unclassified, the matrix that could not be partial. They did not detect violations. They made violations impossible to write. That is a stronger move than checking, and the series will get there. But there is a step before it, and it is the uncomfortable one. Everything enforced here looks for a forbidden presence, and a check that finds none has said exactly nothing about what should be there. Before making violations unwritable, a standard has to prove presence, and say out loud what it did not look at. That is the next article.

Further reading

The ideas here have long lineages, and the originals are worth your time.

  • Gail Murphy, David Notkin, and Kevin Sullivan. Software Reflexion Models: the foundational work on comparing declared architecture against extracted reality, including the divergence/absence distinction this article leans on.
  • Neal Ford, Rebecca Parsons, and Patrick Kua. Building Evolutionary Architectures: the source of the term architecture fitness function and of the mechanical vocabulary (atomic, triggered, static) used here.
  • ArchUnit (JVM) and dependency-cruiser (JavaScript): mature, practical tools for expressing boundary and dependency rules as tests.
  • The baseline/ratchet pattern, as realized in tools like Betterer, lint baselines, and SonarQube's "Clean as You Code": the adoption model that makes enforcement survivable in an existing codebase.
  • David Harel. Statecharts: A Visual Formalism for Complex Systems: the deep background for treating a protocol's full state-event space as a first-class artifact.

A thought after reading?

If you would like to discuss about this article, you can write to me here. I share because I care and I want to learn. Please teach me with care.