AI's Incomplete Knowledge Base
Knowledge Gaps to Knowledge Collapse
(n.b., This article is my ruminations on Deepak Varuvel Dennison’s ”What AI doesn’t know: we could be creating a global ‘knowledge collapse’” [The Guardian, 18 November 2025].)
What I read in a week sometimes has no relation to what is published during a week. This is one of those cases.
I read Dennison’s piece because one of the more spurious claims I read about artificial intelligence is that AIs will somehow allow people to query and interact with “all the world’s knowledge” (ref. Sundar Pichai’s assertion in 2023 that Bard, Google’s then-competitor to ChatGPT, would “combine the breadth of the world’s knowledge with the power, intelligence and creativity of our large language models”).

Dennison, in this article that was originally published over at Aeon.co, does a nice job of laying out why this isn’t (and can’t be) true. He starts with a look at how the language models are designed and created:
“The most popular models privilege dominant ways of knowing (typically western and institutional) while marginalising alternatives, especially those encoded in oral traditions, embodied practice and languages considered “low-resource” in the computing world, such as Hindi or Swahili.”
The ironic thing is that perhaps the most common source of information for the companies, Common Crawl, represents only a slice of the Internet:
“Data from Common Crawl, one of the largest public sources of training data, reveals stark inequalities. It contains more than 300 billion webpages spanning 18 years, but English, which is spoken by approximately 19% of the global population, dominates, with 45% of the content. However, there can be an alarming imbalance between a language’s demographic size and how well that language is represented in online data. Take Hindi, the third most popular language globally, spoken by about 7.5% of the world’s population. It accounts for only 0.2% of Common Crawl’s data.”
The Mozilla Foundation back in 2024 put out a nice look at Common Crawl as it relates to generative AI (“Training Data for the Price of a Sandwich”) that echoes Dennison’s concerns. In it, Stefan Baack and Mozilla Insights wrote:
“Common Crawl does not contain the “entire web,” nor a representative sample of it. Despite its size, there are important limitations on how much of the web is covered. The crawling process is almost entirely automated to prioritize pages on domains that are frequently linked to, which makes domains related to digitally marginalized communities less likely to be included. The language and regional coverage is strongly skewed toward English content. Moreover, a growing number of relevant domains like Facebook and the New York Times block Common Crawl from crawling most (or all) of their pages.”
(note: Common Crawl has been caught in the battle between publisher’s and an AI companies. See Kate Knibbs’s 2024 piece in Wired, “Publishers Target Common Crawl In Fight Over AI Training Data.”)
The passage that really caught my attention in Dennison’s piece was this one:
“The AI researcher Andrew Peterson describes this phenomenon as “knowledge collapse”: a gradual narrowing of the information humans can access, along with a declining awareness of alternative or obscure viewpoints. As LLMs are trained on data shaped by previous AI outputs, underrepresented knowledge can become less visible - not because it lacks merit, but because it is less frequently retrieved or cited.”
This idea of “underrepresented knowledge” sticks with me because it seems to me to be more than the “low-resource” information cited by Dennison: it’s the growing number of information deserts as local media outlets collapse (or are consumed by conglomerates) or as local information and knowledge fails to be captured. Human knowledge expands and contracts over time, the question is what knowledge is (or should be) significant and meaningful enough for preservation…and to whom does that knowledge matter, when, and in what context?
In my talks at local libraries about artificial intelligence, I stress “(Maybe) trust but (always) verify” the answers given by AIs. As we turn to AI as an alternative to search, it’s (still) not a replacement: for issues that matter, that requires knowing (or researching) trusted sources of information. Data literacy (still) matters…and system literacy (particularly in terms of understanding the major mechanics of AI) is becoming every bit as important.

There will soon be a sister institution to the venerable original on 80th and Central Park West. An “American Museum of Artificial History.”
On display, the evolution of Dry Life from earliest stirrings of spell-check to auto-fill, and auto-complete. Highlights of the halting steps to primitive predict analytics. Then on to current LLMs.
If Dry Life recapitulates the Wet variety, there will be a long pause until proto-sentience achieves sapience.
On Earth, that step was just 500,000 years ago. while Wet Life goes back 3.5 billion. One part in 7,000.
Call Babbage’s 1834 engine the “single-cell.” Turing’s machine came a hundred years later; representing the “multicellular.” Yet another hundred years further on, and we are stopped at the soaring cliff of complexity.
It’s not about packing more transistors onto chips. The gap is orders of magnitude more potential connections, as well as the interactive growth and continuous rearrangements of same.
Restate that in personal terms. What dynamically programmable powers are required for a machine to feel the sun on its face, experience the wetness of rain, or hear the thunder?