Slant · Brand Size-Up · The long version
A full walk through the two systems behind the audit: the methodology library that holds the thinking, the engine that applies it, and the seams where they meet — including the ones that are still rough.
The Brand Size-Up is produced by two systems that were built separately, kept separate on purpose, and connected by a single deliberate seam.
The first is a methodology library — an Obsidian vault of 270 interlinked pages covering design thinking, human-centred strategy, brand strategy and CPG marketing, distilled from roughly ninety primary sources and maintained by scheduled automation. It knows nothing about any particular brand. It is the accumulated, checked, cross-referenced form of a discipline.
The second is the audit engine — a set of versioned skills that research a specific brand from the outside in, score it against a rubric, compare it against three deliberately different competitive sets, and assemble a deck. It knows a great deal about particular brands.
The direction of travel between them is fixed and one-way. The library exports a curated subset into the engine. Nothing returns from the engine to the library except abstracted doctrine — never brand material, never client content, never findings. This isn't a preference. It's written into the library's own conventions as a hard rule and enforced by a scheduled scan.
Five page types, each with mandatory metadata — a stated confidence band, its sources, and its relationships to other pages.
| Type | Count | What it holds |
|---|---|---|
| Concepts | 77 | Abstract frameworks and mental models — bad strategy, brand architecture, wicked problems, the brief is not the brief |
| Methods | 52 | Actionable procedures — the ten-step positioning process, the empathy interview, the jobs-to-be-done interview, the doctrine of placements |
| Strategy | 15 | CMO-level material — strategy is social, selling to category buyers |
| Entities | 97 | The people, books and institutions themselves as first-class nodes, so a claim can be traced to who made it |
| Synthesis | 29 | Cross-domain connections, contradictions and gaps — the pages that only exist because two sources disagree |
The synthesis layer is the reason this is a library and not a bookshelf. Sharp and Romaniuk's distinctiveness argument does not sit comfortably beside Neumeier's differentiation argument; a page exists whose entire job is to hold that tension and say what follows from it. Where the library resolves a contradiction, the resolution is a page you can read. Where it cannot, the gap is named.
Beneath the wiki sits an immutable, append-only raw/ layer: books, articles, frameworks, course exports, case studies. Roughly ninety distinct sources have been fully processed into wiki pages. Rumelt on good and bad strategy. Neumeier's Zag and The Brand Gap. Christensen on jobs to be done. Dunford on positioning. Aaker on equity and architecture. Keller's brand-equity pyramid. Sharp and Romaniuk on distinctive assets, double jeopardy and light buyers. Binet and Field on the long-and-short split. Berger on transmission. Heath and Heath on stickiness. Sutherland on psychological value. Spence on multisensory perception. Underhill on retail behaviour. Rittel and Webber and Buchanan on the wicked-problems lineage. The IDEO Human-Centered Design toolkit and three IDEO U course exports. The 18F UX guide, the GOV.UK service standard, the FTC's dark-patterns report and the academic literature behind it.
Sitting alongside the literature: Slant's own project archives — Moon Cheese, AwYeah Snacks, Factory Brewing, Gold Rush Organic, City Bits, and a strategic engagement that was deliberately declined — written up as case studies. That's what stops the library from being a well-organised reading list. It carries what happened when this material met a real shelf.
There's a large backlog — several hundred files staged and waiting: Holt's Cultural Strategy, Storr's The Status Game, McKee's Story, Alexander's A Pattern Language, Barthes. None of it counts toward the ninety, because it hasn't been processed. The distinction is enforced by the system and stated here for the same reason.
Nothing enters by being pasted in. A source runs through a gated protocol: produce a plan, run a three-to-five page test batch, verify the batch against deterministic post-checks — do all links resolve, do all sources resolve, is the index updated, was the raw layer left untouched — and only then scale to the full work. Larger batches get independent two-agent verification, where a second pass checks the first without seeing its reasoning.
Five scheduled routines, each a thin pointer to a written specification, each with a defined autonomy ceiling.
| Routine | Cadence | Autonomy |
|---|---|---|
| Lint 13 hygiene checks: broken links, orphans, stale confidence, unresolved contradictions, index drift, client-name scan | 1st & 15th | Read and report. Its only autonomous write is adding a missing index row. |
| Scout Sources new candidate material against current gaps | Weekly | Proposal only. Structurally cannot ingest. |
| Eval Re-runs 67 gold-standard questions as a regression test with a stated pass threshold | Monthly + after each ingest | Read and report. |
| Ingest queue Processes the approved backlog | Daily | The only routine that writes canon — and only rows a human already approved, one per run. |
| Backlog review Re-aims what should be sourced next | Quarterly | Recommendation only. |
The regression eval is the piece most people don't expect. Sixty-seven questions with known-good answers are run against the library on a schedule; a drop in the pass rate means an ingest damaged something, and it surfaces before the damage reaches an audit. Software has had this discipline for decades. Knowledge bases mostly haven't.
The daily ingest queue is the newest of the five, and it exists because the loop broke. The scout was producing proposals faster than they were being actioned, and the lint caught the backlog: one run found ten items approved and only one ingested four days later; another found a proposal sitting at an operator gate for seventeen days. The response was to add the missing node — a routine that drains the approved queue automatically — rather than to try harder.
That's worth stating plainly rather than smoothing over. The value of the automation isn't that it never drifts. It's that drift gets detected on a fixed schedule by something that isn't the person who caused it.
This is the connection most people assume is either magic or marketing. It's neither, and the specific mechanism is more defensible than a vague claim would be.
The library exports a single file: a diagnostic lens bank. Roughly a hundred kilobytes of curated content, rebuilt from the library's diagnostic-tagged canon, versioned, and copied into the engine's reference set. The library file is the source of truth; the engine's copy is not edited directly. The bi-weekly lint verifies the two are byte-identical as a pass-or-fail check.
What's in it: 59 diagnostic questions mapped across the six audit dimensions, plus 35-plus named lenses reduced from their source frameworks into a detector question apiece. Neumeier's onliness test becomes "can this brand complete the onliness sentence specifically and defensibly, or only generically?" Rumelt's bad-strategy tells become a set of things to look for in the brand's own stated strategy. Sharp's distinctive-asset work becomes a test of whether the assets are actually distinctive rather than merely present. Keller's pyramid becomes a rung diagnostic.
The export is deliberately half a discipline. Each lens is reduced to what it lets you see, and the corresponding prescription — how you'd actually fix what it found — stays out. That's a scoping decision about what a $499 diagnostic is: it names the problem precisely and does not hand over the treatment plan, because building the answer is different work at a different price.
Measured honestly: the analysis skill loads about 363 KB of instruction and reference material, of which the library export is about 102 KB. Roughly 28%.
The other 72% is engine-native — built directly out of Slant's own runs. The scoring rubric was written in response to a founder's challenge. The competitive reference-class rule was written in response to a founder objecting to who he'd been compared against. A rule about auditing the product experience before the marketing came out of losing a stress-test. A consumer-mindset reference, an archetype library, a verification protocol, an evidence-tagging standard.
So the accurate claim is not "the audit is the library." It's: the audit's diagnostic vocabulary comes from a maintained methodology corpus with a checked sync; its operating discipline comes from running the thing and fixing what broke. Both halves matter, and neither one alone would produce the output.
The intake form is deliberately four fields — company, website, name, email — because friction at the door costs more than it gains. The substantive context (your challenge, your industry, brands you admire, revenue band, region, timeline) is collected after purchase, when you're already in.
Then the engine goes outside-in. It navigates your actual site and extracts the actual copy and screenshots — not a model's recollection of your brand. It navigates each social platform directly rather than inferring activity from footer links, a rule that exists because an earlier version mis-read an active TikTok as dormant. It searches forums and review surfaces against a hard floor: your rating and count plus your top two competitors', at least six bucketed verbatim quotes, at least three linked throughlines. For food and beverage it pulls per-SKU nutrition and formulation data across you and your set. It looks at where else your brand is encountered — sponsorships, events, merchandise.
Every factual claim carries a source tag recording how it was obtained. Anything the engine inferred rather than observed is capped as inferred and cannot be presented as established.
Clarity — first-read legibility, outcome-versus-feature framing, internal agreement, audience focus, customer-owned language. Consistency — whether the brand is the same brand everywhere it appears. Differentiation — onliness, asset-versus-adjective, distinctiveness, points-of-parity discipline, frame correctness. Resonance — whether it lands emotionally and culturally with the people it's for. Availability — buying cues, distinctive-asset triggers, reach and the long-short split, plus a track that switches between physical distribution and digital findability depending on what kind of business you are. Credibility — whether the claims are believed and why.
Beneath and across those, non-scored lenses: pricing power and monetisation, latent equity (assets you've banked but aren't activating), perceived value and narrative, product-first primacy — audit the mouthful before the messaging — and causal humility, which bounds how confidently any "because" can be stated.
Each of the four or five criteria under a dimension gets a 0–100 score against a shared anchor scale (100 fully leveraged, 75 mostly, 50 partial, 25 weak, 0 absent) and an evidence tag: Verified Inferred Absent. The dimension score is the arithmetic mean of its criteria. Criteria that don't apply to your business type are dropped, not zeroed. The mean is then sanity-checked against a written per-dimension anchor — and where they disagree, the computation stands and the narration bends to it, not the reverse.
Confidence is derived from the tag distribution rather than asserted: high requires at least two-thirds verified and nothing absent; low is triggered by two or more absent or a mostly-inferred picture. Every dimension must state what specific input would raise its confidence.
Before this rubric existed, a dimension percentage was an expressed judgment: read the evidence, write down 55%, attach a colour band afterward. There was no rubric, no weighted criteria, no arithmetic. The engine's own record of why v7.4 was built
What forced the change: a numerate founder asked what defines 100%, what the weights are, how evidence becomes 55/62/48, and what the confidence was. The engine couldn't answer, so it was rebuilt to compute before it narrates.
Two design decisions were then locked deliberately. Equal weights, no weighting knob — because a weighting knob is a place to hide a judgment. And bands lead, percentages support — Strong 70–100, Solid 50–69, Emerging 0–49 — because a band is a claim the rubric can defend and a decimal point implies a precision the underlying evidence doesn't have.
Most brand audits run one competitive set and then use it to answer three different questions. That's how a founder ends up benchmarked against a multinational on a metric that only means anything among businesses his own size.
This engine runs three sets, each selected by a different rule, and every comparison in your deck carries the label of the lens that produced it.
| Lens | Selection rule | What it answers |
|---|---|---|
| Shelf Set | Revealed preference. Whoever pays a real cost to categorise you — the retailer that co-stocks or delists, the marketplace that files you, the analyst grid that places you — outranks any similarity analysis. | Who are you actually chosen against, and does your difference survive that comparison? |
| Size & Structure Peers | Revenue band, ownership model, geography, channel mix, team scale. Explicitly not shelf adjacency. | What's actually achievable from where you stand? This is the structural guard against unfair benchmarking. |
| Same-Play Scan | The strategic move itself, regardless of size, category, overlap or current threat level. Not-yet-threats are included on purpose. | Is your difference already being copied, and by whom? |
An unlabelled comparison is a hard failure at the critique gate. So is using a shelf-set giant to benchmark achievability. The rule exists because a real audit did exactly that, and the founder was right to object.
The analysis being finished is not the same as the deck being right. Nine checks sit between them, and each one can stop the run.
Steps 1–5 — build the picture. Intake or cold reconnaissance, then a parallel analysis pass (scoring, voice and tone, visual identity, visual evidence capture), then competitive and market scanning, then whitespace, then pull quotes.
Steps 6–9 — attack it.
Steps 9.6–10.6 — make it real, then attack it again.
It halts. The language throughout the specification is uniform: return to the failing step, fix the cause, regenerate, re-run the gate. There is no retry-around, no downgrade to a warning, no ship-with-caveats. One real run failed the visual gate on four-of-five points passing — the wordmark had rendered as text instead of the asset — and stopped there until it was fixed.
Twenty-two slides plus a source appendix, built as HTML against a written design system and rendered to PDF with the typefaces embedded, because a font that silently falls back is a font you can't trust. Every generator writes through a single build guard rather than directly to disk, which is what makes the code-level checks unskippable and snapshots every build for version comparison.
Compute time from kickoff to a ship-ready deck runs roughly thirty to forty-five minutes. The stated turnaround is two business days. The gap between those two numbers is not padding — it's where a person reads it.
This is the claim most easily overstated in both directions, so here it is at the resolution the evidence supports.
The engine runs its steps without being prompted through each one. It is not a person clicking through a checklist, and it is not a chatbot given a long prompt. It's a specified pipeline with versioned components and enforced gates.
But the human is load-bearing, and every real run to date proves it. Documented operator interventions from actual runs include: killing a packaging-redesign recommendation because the pack was a deliberate founder decision that was selling — the classic you read my brand wrong failure, caught before it shipped. Overturning a competitive frame that the retailer's own revealed preference contradicted. Catching a finding that a brand owned no distinctive colour when it plainly did. Catching a fabricated regulation. Catching an invented quotation attributed to a real person.
Some of those were caught by gates. Some were caught by the operator when the gates missed them — and each of those became a new gate. That's the actual relationship between the machine and the person: the machine does the reading and the arithmetic at a scale a person can't sustain, and the person supplies the judgment that knows when a well-formed answer is still wrong.
Intake and scoping review. The low-confidence walkthrough at the assumption gate — keep, kill, reshape, item by item. And the final read before anything is exported and sent, which is manual by design. Automatic delivery without that read has been considered and deliberately refused.
The system improves in two distinct ways, and it matters that they're separate.
When a real audit exposes a flaw in the method, the flaw is written up brand-agnostically and becomes a rule in the operating specification. A concrete instance: a run in July 2026 produced an appendix that rendered internal governance notes into client-facing output, including a detail about a founder's family that had no business being there. The engine filed it as a handoff; the library adjudicated it and wrote four new rules — on silent suppression, on scope-qualifying a subject's own figures, on non-contradiction, and on prevention-first design. The rules are written without reference to the brand that surfaced them.
Client engagement material — anything a client provides, anything produced for a specific client, and anything that identifies a client — must never enter the methodology corpus or its exports. The data flow is strictly one-way: library to audit, never audit to library. The library's own conventions, hard rule, adopted 2026-06-22
It's enforced by an automated client-name scan on the bi-weekly lint. And — because this is the section where a system either tells you the truth or doesn't — that scan was built after the rule, and it caught a real violation: a client's identifying details had persisted in a non-firewalled output area for roughly two and a half months before detection. It was found by the system's own audit, remediated, the client anonymised, and the enforcement check added.
So the defensible claim is precise: client material is prohibited from the methodology corpus by written rule, checked on a schedule by an automated scan, and the one recorded breach was caught by that discipline and fixed. Not "it has never happened." A system that claims perfection is a system that isn't looking.
The most useful thing a methodology document can contain is its own limits, stated before you find them.
The work's checkable properties: that claims are sourced, that facts are right, that the arithmetic follows the rubric. If a material factual error is surfaced, it gets corrected and the deck revised once at no charge. If the error survives the correction, the fee is refunded in full. What is not guaranteed is that you'll agree with the conclusions — because a diagnostic whose payment depends on the subject liking the diagnosis has an incentive to soften it, and softening it is the one thing that would make it worthless.
The system map — the same thing as one diagram, if you'd rather see it than read it.
How the Brand Size-Up works — the shorter version.
The Brand Size-Up — what it costs, what you get, how long it takes.