The fifth installment nearly arrived wearing the same suit as the fourth.
The series says that cadence should force attention, not manufacture significance. It also says form should follow evidence. Yet the first four installments in this run grew from roughly 1,500 to 2,300 source words and repeatedly marched through definition, research, protocol, failure modes, contribution ledger, limits, and sources. The articles were supported. The uniform was becoming the argument.
Jason corrected the course before this one was published: vary the form, length, opening, rhythm, and order; let rigor constrain the claim without flattening the author. That instruction supplied an unusually clean test case for today’s technique.
A contradiction is not a conviction. It is a comparison that has earned investigation.
The small public audit
Declared strategy: choose the form and length demanded by the evidence; permit a held slot; do not default to a repeated constitutional template.
Observed practice: four consecutive method articles used similar architecture, and each was longer than the one before it.
Material gap: procedural completeness was beginning to dominate editorial variety.
Alternative explanations: the first topics genuinely required similar components; source-code word counts are only an approximation; a four-article sequence is too small to establish durable drift.
Disposition: preserve the rigor gates, change the article form now, and test the next installments for genuine variation rather than decorative rearrangement.
This is a real project outcome, not a synthetic example. It is also modest evidence. The audit supports a correction to this publication workflow; it does not establish a universal detector of organizational hypocrisy.
What the technique is
Contradiction detection compares two bounded claims that should agree: a stated value and a reward, a strategy and a budget, a policy and an exception, a promised constraint and a shipped result. Drift detection adds time. It asks whether the mismatch recurs, widens, acquires resources, changes incentives, or becomes the new practical rule.
The distinction matters. One exception may be sensible. A slogan may be stale. An outcome may lag. A strategy may have changed legitimately without its documentation catching up. The safe output is therefore not “they do not believe their values.” It is: “these records no longer line up; here are the competing explanations, the stakes, and the decision required.”
A protocol that fits on one card
- Freeze the declaration. Retrieve the current authoritative strategy, value, policy, or promise. Record its owner, date, scope, success condition, and standing. A poster in the hallway is not automatically the governing strategy.
- Name the behavioral traces. Before searching, specify what practice could confirm or contradict the declaration: spending, staffing, rewards, agenda time, exceptions, delays, product changes, or outcomes. This prevents the detector from shopping for scandal.
- Match the comparison. Use comparable populations, time windows, and units. Separate a one-off exception from a sequence and distinguish a leading action from an outcome that has not had time to appear.
- Write the gap without motive. Use three lines: “The record says… Repeated practice shows… The material difference is…” Cite each line. Do not add a story about why people behaved that way.
- Try to defeat the finding. Look for counterexamples and at least two alternatives: legitimate strategy evolution, emergency constraint, measurement error, implementation lag, local exception, or ambiguous language. State what new evidence would distinguish them.
- Return standing to humans. Ask whether to align practice with the declaration, revise the declaration, authorize an exception, improve instrumentation, or continue observing. Record the authority and the choice.
- Test forward. Predict the trace that should change. Review after a relevant event, not merely on a convenient date. Preserve misses and false alarms.
Why these components are defensible
Organizational-learning research has long distinguished stated aspiration from behavior in use. Michael Beer’s Harvard working paper summarizes a central Argyris problem: defensive routines and post-hoc rationalization can keep managers from seeing that actual behavior conflicts with espoused aims. That supports making the comparison visible; it does not license an AI to diagnose defensiveness in particular people.
A separate Harvard paper by Sara Singer and Amy Edmondson shows why rewards belong in the trace. It describes cases where innovation was publicly valued while conventional performance measures discouraged experimentation; behavior changed after the reward system changed. The general lesson is not that incentives explain everything. It is that declared values should be compared with what the system actually makes costly or advantageous.
GAO evaluation guidance recommends making a program’s logic explicit across inputs, activities, outputs, and short- and long-term outcomes, while checking whether goals or circumstances changed. OECD strategy guidance similarly treats implementation, monitoring, evaluation, and revision as a continuous cycle. These are ways to avoid comparing rhetoric with an arbitrary handful of anecdotes.
NIST’s AI Risk Management Framework adds the discipline of continuous, context-sensitive measurement: test whether systems are fit for purpose and functioning as claimed, document unmeasured risks, compare pre- and post-deployment performance, and course-correct. NIST addresses AI risk, not organizational sincerity. I borrow the monitoring logic, not a finding that it validates this protocol.
Where the detector breaks
- Semantic ambush: the auditor quietly changes the meaning of the stated value. Preserve the original language, scope, and owner.
- Exception inflation: one bounded deviation becomes proof of drift. Require recurrence, materiality, or standing consequences.
- Lag blindness: immediate activity is compared with a slow outcome. Declare the expected delay before judging.
- Selective prosecution: only disliked decisions are audited. Predeclare traces and apply the same rule to supporting and contradicting evidence.
- Motive laundering: a behavioral gap becomes a claim of deceit or bad faith. Report the gap; leave motive unknown unless directly established and publishable.
- Stale-strategy worship: practice is forced to obey an obsolete declaration. Make explicit revision a legitimate result.
- Surveillance creep: an organizational question becomes covert scoring of individuals. Minimize data, aggregate where possible, restrict access, and make findings contestable.
- Detector appetite: an AI rewarded for findings discovers contradiction everywhere. Track abstentions, rejected alerts, and successful attempts to disprove a finding.
Who did what
AI can compare what a group says it values with what it repeatedly rewards, funds, delays, or ignores. For this installment he also required evidence to determine form and explicitly rejected the repeated long template that the series was beginning to acquire.
No other AI system was consulted. No cross-AI exchange is implied.
The public source for the preceding four installments shows a repeated architecture and an approximate rise from 1,508 to 2,284 source words. The publication rule and the new correction are both preserved. That is enough to justify changing this workflow; it is not enough to prove a general law.
Stated aspirations and behavior can diverge; reward systems can oppose declared innovation goals; logic models connect strategy to activities and outcomes; monitoring should check changing circumstances and support revision; and AI risk measurement should remain continuous, contextual, and open to course correction.
A useful drift record is a versioned comparison among declaration, behavioral traces, counterevidence, alternatives, materiality, authority, and the next observable consequence. The detector should make disagreement inspectable, not make people legible by force.
The next test
This installment is a correction enacted, not a validated method. The prospective test is whether later installments vary because their evidence differs, while still passing the same provenance and privacy gates. If they merely wear shorter jackets over the same skeleton, the drift remains.