Someone reading a piece we published on entity mapping left a comment that gets said constantly right now: go win the parametric side. Build parametric authority. Four words, sounding like a task you hand to a team with a deadline.
It is not a task. In almost every case, the work behind those four words is already finished, it took years, and none of it was done with a language model in mind. There was no line item for it because there was no category to bill it against. Understanding why changes what you should actually do next.
Parametric authority is not a campaign, it’s a residue
Two terms get used as if they were interchangeable, and they are not. Training data is the raw text a model ingested. Parametric authority — what a model already believes about you before it looks anything up — is whatever survived compression into its weights. The relationship between the two is real, it’s directional, and it’s lossy: a great deal goes in and never comes back out in usable form.
That compression step is the whole story. Publishing more content does not automatically deposit more into a model’s memory. What gets retained is the description of your company that enough independent sources repeated, in enough different phrasing, that the pattern survived. Volume from one source does not substitute for variety across many.
The research draws a hard line under “just publish more”
Three separate studies converge on the same wall. Kandpal and colleagues, presenting at ICML, showed that a model’s accuracy on a given fact tracks the number of relevant documents it encountered during pretraining, and established that relationship causally rather than just correlationally. They also estimated a model would need to scale by many orders of magnitude before it could answer confidently about a subject with thin coverage. A bigger model does not rescue a thin footprint.
Mallen and colleagues arrived at the same ceiling from a different angle: models handle well-covered entities fine and struggle badly with the long tail, and scaling mostly sharpens recall at the popular end while leaving the tail roughly where it started.
The sharpest finding belongs to Allen-Zhu and Li. They showed that a fact only becomes reliably extractable from a model once it appears in sufficiently varied phrasing during training. Without that variation, the fact can be sitting inside the weights and still return zero accuracy under questioning — present, but unusable. Their experiment ran on a controlled dataset aimed at engineers building pretraining pipelines, so it shouldn’t be stretched into a direct claim about brands. As a mechanism, though, it explains a lot: ten mentions from one source behave nothing like ten mentions from ten sources.
None of this is a full rejection of newer arrivals. A company founded in 2023 that a current model describes accurately did not beat the mechanism — enough independent sources happened to describe it at once, the same way something goes viral. The exception runs on the same rule as everything else.
There is no account to send the invoice to
The second reason “build parametric authority” can’t be assigned as a task is that there’s no corpus to target directly. Elazar and colleagues, in a project called What’s In My Big Data, examined ten corpora used to train popular models. In C4 — the Colossal Clean Crawled Corpus, built from a single Common Crawl snapshot — even the single most common domain accounted for less than five hundredths of one percent of documents. Whatever you publish on your own site is a rounding error against that scale.
Common Crawl’s own documentation adds a complication: its crawler respects robots.txt and deliberately avoids overloading servers, so heavily trafficked domains tend to be underrepresented relative to their real-world importance, and plain HTML gets favored over other formats. A separate paper on C4, led by Jesse Dodge, found the sites inside the corpus don’t map cleanly onto the sites people actually use most — and a lot of where companies accumulate description, on large rendered platforms rather than static HTML, is exactly the territory a polite crawler treats gently.
Most of what a model “knows” about you predates the question
Language models have been in general use for roughly four years. The text underneath them is considerably older. The best-documented corpus, C4, was pulled from an April 2019 snapshot; Dodge’s team sampled a million of its URLs, used the earliest Internet Archive index date as a proxy for when each page was written, and estimated 92% of it dated between 2011 and 2019, with a long tail stretching ten to twenty years further back. C4 isn’t what today’s production models train on, but the pattern it documents is instructive.
Cheng and colleagues at Johns Hopkins pushed this further with research on “effective cutoffs.” They found that the date a model reports and the date its knowledge actually concentrates around often differ substantially — partly because new Common Crawl dumps still carry older material, and partly because deduplication struggles with near-duplicate content. The upshot: a model’s picture of your company is probably older than its stated cutoff suggests. Whatever standing you have today was largely deposited before your category existed as something to optimize for, so the PR and review management your team did years ago had to be genuinely good to be paying off now.
Marketing built this without a line item for it
Here is the uncomfortable part for anyone running a marketing organization. Nearly every function responsible for parametric authority already reports into marketing: PR, analyst relations, community management, event presence, local press, crisis response, review operations. None of it sits outside the department’s ordinary remit.
What none of those functions ever owned was the output itself. PR earns coverage a journalist chooses to write. Analyst relations earns an assessment an analyst forms independently. Reviews are the clearest case: you control the response, never the review, and thousands of those accumulate over years in language nobody at the company chose. None of that text entered a model as your words. It changed how outside parties described you, and independent description is the only channel into the weights that exists.
Which makes marketing’s oldest discipline the relevant one after all. Earned media has always meant paying for outcomes you don’t get to author — staff, agencies, and time are the real cost, whether or not the coverage looks “free.” That constraint didn’t arrive with language models; it’s just now the only mechanism that reaches parametric authority, and whoever holds the marketing seat today owns a result produced under a previous team’s budget, often years after those people moved on.
Why none of this comes apart easily, either
Cohen and colleagues, writing in Transactions of the Association for Computational Linguistics, tested prominent knowledge-editing methods and found they fail to produce consistent changes — editing one fact triggers a ripple of related facts that also need updating and usually don’t. Researchers with direct parameter access, no adversary, and full knowledge of the target still can’t cleanly rewrite a single fact.
That cuts both ways. An outside party can’t install a new description just by publishing harder. But standing built from thousands of independent descriptions also doesn’t unravel from one bad quarter or one competitor’s campaign — distribution produces both the difficulty of editing and the durability of what’s already there. It also means a model can keep repeating something that stopped being true a while ago, which is a separate problem worth its own treatment.
What this means for the work you can actually plan
Parametric authority moves on the timescale of model generations, not campaigns. It responds to being described by others, not to what you publish. Most of the functions that build it were never measured against this outcome, and that’s probably fine — they shouldn’t suddenly be managed exclusively for it either.
What is genuinely actionable is the layer sitting on top of this: retrieval. What a model says about you today can be measured, tracked release over release, and compared against what your own team believes is true — which is a very different exercise from trying to author the next training run. We’ve written separately about that measurement gap in what to tell clients when organic traffic drops, and about the practical side of getting into the corpus in the first place in how to get into model training data.
So when someone proposes going out and “winning” the parametric side, the accurate answer is that you’ve already been paying for it — across years of other people’s sentences, through work nobody counted as an investment in machine memory because there was nothing yet to count it against. That work compounds differently now: what you do today shapes tomorrow’s review, press mention, or forum thread, which feeds the next training set. The company-controlled side of this — whether search engines can find and parse those signals at all — is closer to what we cover in entity mapping and whether it reaches ChatGPT and in why training data cutoffs function like a ranking factor.
Frequently asked questions
What is parametric authority, exactly?
It’s what a language model already “knows” and states about a company without doing a live lookup — the description that survived compression into its weights during training, as distinct from anything it retrieves at query time.
Can a brand build parametric authority through content marketing alone?
Not directly. Research shows a model’s accuracy about a subject tracks the volume and variety of independent sources describing it, not the volume of content the subject itself publishes. Self-published content matters for retrieval and for feeding future independent coverage, but it isn’t the same lever.
Why does a model’s knowledge of my company seem outdated?
Training corpora skew older than most people assume, and research on “effective cutoffs” shows a model’s knowledge often concentrates well before its stated training date. What the model says about you may reflect deposits made years ago.
If I can’t edit parametric authority directly, what can I actually work on?
Focus on retrieval — what the model can find and cite right now — and on giving future independent sources (press, reviewers, community members) accurate, varied material to describe. Measure current AI visibility regularly rather than trying to reverse-engineer the next training run.