Many organisations believe they have a volume problem. They publish more, open a blog, add service pages, expand their documentation, write case studies, produce expert content. Then they notice that the site remains hard to understand, that pages compete with one another, that the right content does not carry the right messages, and that both AI systems and people discovering the organisation seem to retain only an impoverished version of the whole.
In the majority of cases, the problem is not the absence of content. It is the absence of a usable corpus.
What we call a usable corpus
A usable corpus is not simply a stack of URLs. It is a set of content that can be browsed, linked, interpreted, and reused without the overall meaning being distorted at each step.
For a services company, this means that commercial pages, proof, articles, FAQs, and expertise profiles tell the same value structure. For a software publisher, it means that product marketing, documentation, support, comparisons, and release announcements do not contradict one another. For a multi-domain group, it means that each asset has a clear role in the overall reading.
A corpus becomes usable when it helps a reader, whether human or machine, answer precise questions:
- what does the organisation do;
- for whom;
- with what offering structure;
- with what proof;
- in what order should things be read;
- which content serves to explain, which serves to convert, which serves to demonstrate.
Signs of a corpus that does not carry
The symptoms are often normalized when they actually reveal a deep problem.
The first sign is duplication without hierarchy. The same ideas recur everywhere with slightly different wording. A service page, a FAQ, a blog post, and a sector page repeat the same promise without specifying each piece’s function. The person discovering the organisation feels like they are going in circles.
The second sign is documentation with no commercial role. The site has content, sometimes a great deal of it, but it does not become a positioning asset. The pieces remain isolated.
The third sign is false depth. There is a lot of text, but little structure. The organisation appears dense; in reality, it is no more readable.
The fourth sign is disordered reading. A person can arrive on any page and encounter a fragment that does not help them understand the whole. Generative AI systems experience exactly the same thing: they pick from a corpus that does not clearly tell them what to retain.
Why this issue becomes more important with AI
AI systems worsen the problem because they read by condensation. They do not naturally respect the implicit hierarchy your team has in mind. They do not spontaneously assign the same importance to the pages that seem central to you. They work with what is available, repetitive, well-linked, or more easily interpretable.
If the corpus is too thin, they fill the gaps. If it is too noisy, they simplify. If it is contradictory, they produce an average. In every case, a poorly structured corpus reduces the quality of the reading.
That is why a “rich” site can remain poor from a comprehension standpoint.
What organisations often do instead
When they sense this problem, many teams react with production. They commission more content. They open new sections. They want new pages. They add a FAQ, a glossary, case studies, a resource centre. The corpus grows. The structural problem remains.
Other teams react with simplification. They cut everything, reduce pages, compress the message, remove nuances. This can improve immediate clarity, but at the cost of lost precision. The site becomes lighter and less accurate.
In both cases, volume or surface is treated, not architecture.
Our approach
We treat the content issue as a problem of editorial architecture.
The work involves distinguishing:
- what pertains to discovery;
- what pertains to proof;
- what pertains to qualification;
- what pertains to conversion;
- what pertains to governance.
Then, these roles must be linked together. A service page should not do the same thing as an article. A FAQ does not need to carry the depth of a diagnostic. A proof piece should not look like a sales page. A diagram does not replace an article. An article does not replace a proof page.
When this architecture is clear, the corpus becomes more useful. Each piece has a function. Each relationship becomes more intelligible. Unnecessary repetitions decrease.
Contexts where this issue becomes critical
It becomes particularly visible:
- when an organisation has published extensively without ever structuring the whole;
- when several teams write with different objectives;
- when a redesign is planned while the current corpus has never been mapped;
- when a marketing site coexists with documentation, support, a blog, and satellite assets;
- when people discovering the organisation read a few fragments and leave with a partial or erroneous understanding.
For a B2B software publisher, the problem often shows up in the artificial separation between marketing and documentation. For a specialized SME, it shows up in the dispersal among service pages, case studies, and expert content. For a personal brand, it shows up when public appearances, the main site, and secondary assets do not add up.
What a usable corpus changes
The most important change is the reusability of meaning. People discovering the organisation understand more quickly. Search engines and AI systems have more coherent anchor points. The right pages stop being buried. Deep content reinforces commercial content instead of duplicating it.
A usable corpus does not necessarily require more content. It requires better-articulated content. That is a decisive difference.
When this issue warrants intervention
This work should begin when:
- the site appears rich without being readable;
- teams sense there is “a lot of material” without knowing what to put forward;
- documentation, commercial pages, and expert content move in silos;
- a redesign or editorial expansion is planned;
- external responses remain too thin despite genuine investment in content.
At that point, the question is no longer “what should we publish?” The real question is: what corpus must be built so that comprehension holds over time? Semantic content architecture provides a structural answer to this question.