Skip to content

Chesterton's Fence in Software: Ask Why, Then Delete

Chesterton's Fence is quoted as a reason to leave old code alone. Read as a time-boxed debugging question, it says when to remove it and how to record the answer.

Ayhan Sipahi Ayhan Sipahi

Chesterton’s Fence gets quoted as a reason to leave unexplained code alone, yet Chesterton’s own passage allows removal once the purpose is understood. So treat every fence as a time-boxed debugging question with two halves: why was it built, and who depends on it now. Once the budget is spent, default to staged removal, provided the change is reversible and its dependents are observable. Whether the fence stays or goes, leave the answer behind as a comment, commit body, ADR, or flag metadata.

The fence can be a conditional, a retry, a sleep(250), a request header, or a security-group rule in a system you did not build. Nobody on the team can say why it is there. Deleting it risks a hidden dependent. Keeping it forever fills the codebase with weight nobody dares to touch, and a dormant path can be reactivated by accident.

What the Original Passage Allows#

The fence passage comes from G. K. Chesterton’s The Thing (1929), in the chapter “The Drift from Domesticity”. The half that circulates is the reformer who is told to go away and think before clearing the fence. The half that rarely circulates comes a few sentences later, where Chesterton describes what the thinking reformer can conclude: “If he knows how it arose, and what purposes it was supposed to serve, he may really be able to say that they were bad purposes, or that they have since become bad purposes, or that they are purposes which are no longer served.”

That sentence carries the whole argument. The principle demands an investigation, and it explicitly permits removal as one outcome of that investigation. A reviewer who writes “Chesterton’s Fence” under a deletion and stops there has invoked the first half and skipped the second. The code ownership post approaches the same passage from the ownership side.

Two Questions with Different Evidence#

Chesterton’s text asks about purpose. Software adds a second question that his text could not anticipate, because his fence had no API. Hyrum’s Law states it: “With a sufficient number of users of an API, it does not matter what you promise in the contract: all observable behaviors of your system will be depended on by somebody.” The original author’s intent answers only the first question. Whether anyone depends on the behaviour today is a separate matter, and the answer can be yes even when the intent is long obsolete.

QuestionWhere the evidence livesHow to get it
Why was it built?Commits, PR descriptions, linked tickets, ADRs, commentsgit log -S, git log -G, git log -L, git blame -w -C -C -C, .git-blame-ignore-revs
Who depends on it now?Traffic, logs, metrics, call sites, consumers outside the repoCode search, access logs, a metric on the code path, a staged scream test

In practice the two questions rarely share an answer. A retry added for a flaky upstream may be pointless now that the upstream is stable, while a downstream batch job has quietly started to depend on the extra latency it introduces. Archaeology finds the first fact; only observation finds the second.

Recovering the Why from Git History#

Plain git blame is the wrong tool for the first question when the repository has any age. A formatting sweep, a file move or a squash merge makes the blame line point at whoever touched the line last, and the reason looks lost when it is recoverable. Git has specific options for exactly this problem.

# Which commits added or removed this exact string? (pickaxe)
git log -S 'await sleep(250)' --oneline -- src/

# Same idea, but match a regex against added or removed lines
git log -G 'retry(Count|Limit)' --oneline -- src/payments/

# Full history of one function, even as its line numbers moved
git log -L :retryWithBackoff:src/payments/client.ts

# Blame that ignores whitespace and follows code moved or copied across files
git blame -w -C -C -C src/payments/client.ts

# Skip bulk-formatting commits in your local blame (each clone runs this once)
git config blame.ignoreRevsFile .git-blame-ignore-revs

-S finds every commit that changed the number of occurrences of the string, so it surfaces the commit that introduced the line and the one that last removed it. -G matches a regex against added and removed lines, which suits identifiers that changed shape. -L :<funcname>:<file> follows one function through line moves. It relies on git’s default function-name pattern, which finds top-level functions; indented class methods usually need a diff= driver in .gitattributes, or a -L <start>,<end> line range. blame -C given three times also looks for code copied from other files. Commit a .git-blame-ignore-revs file at the repo root. GitHub’s blame view reads it automatically, and each developer opts in locally with the config line above.

Once the introducing commit is found, the commit body and its linked PR or ticket are the highest-value reads. A body that says “add retry for ledger 409” plus a ticket describing a replica lag incident answers the first question in minutes. A body that says “fix” answers nothing.

Sizing the Budget to Reversibility#

The investigation has a cost, so it needs a limit. With no limit, the cleanup stalls; with the same small limit for every fence, a data migration ends up scream-tested. The budget should scale with two properties of the change, and how odd the code looks is not one of them.

The first property is reversibility: can the removal be undone in one deploy or one flag flip, with no data loss? The second is observability of dependents: if something did depend on the fence, would the breakage show up in metrics, logs or support tickets within a window you can wait for? When both hold, a short look at history followed by a staged removal is proportionate. A reasonable anchor is an hour or two with the git history and the linked tickets. When either property fails (data deletion, schema changes, security controls, anything consumed outside your telemetry), the budget expands to a full investigation with owner sign-off and a written record, whatever the size of the diff.

No

Yes

No

Yes

Yes, still valid

Yes, obsolete

No

Fence found: code, flag, rule or resource nobody can explain

Default: time-boxed investigation, then staged removal

Reversible in one deploy or flag flip?

Override: full investigation, owner sign-off, written record

Would a broken dependent show up in metrics or tickets?

Did the time-box find the reason?

Keep it and write the why down

Staged removal, then write the why down

Staged removal with a full observation window and rollback ready

The trade-off of a time-box is that engineering time is spent up front and the budget can be too small for a high-consequence fence. Scaling the budget with irreversibility works because irreversibility is usually visible before the investigation starts: you know whether the change touches stored data or a security boundary before you know why the code exists. A practical form is a rule that any change inside a named set of directories (migrations, IAM, billing) never takes the short path, however trivial the diff looks.

Staged Removal as the Second Half#

A scream test answers the second question directly: switch the thing off, and see who notices. Microsoft’s Inside Track article on server lifecycle describes the crude form of it. A Microsoft engineering team turned off machines nobody claimed and waited to hear from anyone still using them; about 15 percent of the machines turned out to be unused. This is usually presented as the opposite of Chesterton’s Fence. It is closer to the second half of the same investigation, because archaeology cannot find a Hyrum’s-Law dependent that nobody documented, while an observation window can.

The responsible form is staged. Announce the removal where dependents would read it. Disable behind a switch or degrade the behaviour, keeping the old path deployable. Watch the metrics, logs and tickets for a window that matches the slowest plausible consumer. Once the window closes, delete the code, the switch and the tests for the dead path together.

The strongest objection is that staging is ceremony: delete the code, and if something breaks, git revert is the rollback. That holds when the breakage is loud and fast, because a revert the same day is as cheap as a flag flip. It fails when the scream arrives weeks later. By then the revert conflicts with everything built on top of the deletion, and a dependent may have spent the gap writing wrong data that no revert restores. The observation window covers that gap, while the old path is still one switch away.

A scream test also only hears screams. A dependent that runs monthly or quarterly can miss the observation window; a scream that arrives as a customer-facing incident is a real cost; and silent data corruption never screams at all. Because of that last case, the staged test belongs only on the reversible-and-observable branch of the diagram. The observation window has to be set by asking what the slowest dependent’s cycle is, and a comfortable number of days is not an answer to that question.

When to Override the Default#

The default (time-boxed investigation, then staged removal) gives way in two directions.

Investigate fully first, with no staged step, when the change is irreversible or its failure would be silent. Data deletions, schema changes, security controls, and behaviour consumed by systems outside your telemetry all belong here. The SEC’s Knight Capital order is the standing example of the other side of this risk. Per the order, a component that had not been used for years was neither deleted nor deactivated, and a repurposed flag reactivated it on a server that had missed the deployment. Keeping an unexplained fence in place is a decision with its own failure mode, so the full investigation should still end in removal when the purpose is gone.

Shrink the time-box to a brief look at history when the change is trivially reversible, the dependents are all inside your observability, and the archaeology has already failed once. A flag whose introducing commit says “fix” and whose ticket is gone is a reasonable candidate, provided the observation window is set by the slowest consumer and the rollback is a single flip.

Common Failures in Fence Investigations#

The recurring failures are procedural, and each has a direct fix.

  • The fence as a review veto. Fix: the reviewer who invokes it owns part of the investigation or accepts a staged removal with a rollback plan.
  • Plain git blame as the whole archaeology. A formatting commit or file move points at the wrong author. Fix: -w -C -C -C, a committed .git-blame-ignore-revs, and git log -S or -L on the specific line or function.
  • Asking only about original intent. Current dependents may rely on side effects nobody intended. Fix: check call sites and traffic as well as history.
  • Reusing a flag name or bit for new behaviour. Fix: new behaviour gets a new flag; old flags are deleted, never recycled.
  • Investigating and leaving no trace. Fix: every investigation ends in a comment, commit body or ADR, including the outcome “kept because X”.

Leaving the Why Behind#

A fence investigation happens at all because the previous change left no trace of its reasoning. The lasting fix is upstream: every change that adds or keeps a fence records why, in a place the next reader will find without searching. This is where the two classic positions on old code reconcile. Joel Spolsky argues against rewriting whole systems, but his point transfers to a single branch: the hairy parts of old code are bug fixes, and discarding them discards knowledge. The deprecation chapter of Software Engineering at Google is right that code is a liability and the value lies in the functionality. Both hold if the knowledge moves out of the code’s shape and into a record, so the code itself can change.

Comments earn their place by explaining why the code exists. Google’s code review guidance says exactly that: comments are usually useful when they explain why some code exists, and should not be explaining what the code is doing. A why-comment on the fence itself is the cheapest record, and it should name the condition under which the fence can go. For changes with no natural home in a comment, the commit body carries the same information, which is rule seven in Chris Beams’s commit-message guide.

// Why: the upstream ledger API returns 409 briefly after a write while its
// read replica catches up. A 250 ms wait before the single retry avoids a
// false "duplicate" error.
// Remove when the ledger client exposes read-your-writes (see ADR-0031).
await sleep(250);

Decisions that span more than one file belong in an architecture decision record. Michael Nygard’s original proposal keeps each ADR small (title, context, decision, status, consequences) and leaves reversed decisions in place marked as superseded. He frames the problem in terms that match the fence: a new team member who meets a decision without its context can only accept it blindly or change it blindly. The bus factor post covers how ADRs fit into a team’s knowledge base, and the documentation as infrastructure post covers keeping such records alive.

Feature flags are the fence type that most often outlives its purpose, because a rollout ends and nothing forces the cleanup. Pete Hodgson’s article on feature toggles treats them as inventory with a carrying cost. He notes that some teams attach expiration dates, with a test or a startup check that fails once a toggle is past its date. The feature flags in React Native post covers the runtime side; the metadata and test below cover the lifecycle side.

// flags.ts
export const FLAGS = {
  newCheckout: { owner: 'payments', reason: 'ADR-0042 staged rollout', expires: '2027-01-15' },
  legacyTaxRounding: { owner: 'billing', reason: 'ADR-0017 regulator audit', expires: '2027-03-01' },
} as const;
// flags.test.ts
import { describe, it, expect } from 'vitest';
import { FLAGS } from './flags';

describe('feature flag expiry', () => {
  it.each(Object.entries(FLAGS))('%s has not expired', (name, flag) => {
    const expires = new Date(flag.expires).getTime();
    expect(expires, `${name} expired: remove it or extend it with a reason`).toBeGreaterThan(Date.now());
  });
});

An expired flag fails CI with the flag’s name and a message that offers two ways out: delete the flag and its dead path, or extend the date with a reason in the same commit. The same shape applies to firewall and security-group rules, where a description field holding owner, reason and review date is the fence’s plaque. The cost per change is small. The records decay if nobody reviews them, and ADRs decay less when they stay short, which is part of Nygard’s case for small documents.

References#

Related posts