Something interesting has been happening beneath the surface of ChatGPT.
On August 8th, the way ChatGPT searches the web changed overnight. According to data from Promptwatch, the proportion of searches that ChatGPT deliberately restricted to a specific website using the site: operator jumped from just 0.37% to 16.8% in a single day.
Then, less than a week later, something else happened.
Reddit has recently been accounting for around 3.8% of the sources cited by ChatGPT. On August 14th that fell below 1%, with Promptwatch recording an average of just 0.52% over the following four days - a drop of 86%. And this isn’t the first time this has happened.
In September last year, Reddit accounted for somewhere between 10% and 14% of ChatGPT citations before collapsing to around 3% overnight. There are plenty of interesting implications here, but I think there’s something much bigger going on too.
When consumers use ChatGPT, it doesn’t just search the internet for answers. It decides whether it needs to search at all, what to search for, and increasingly, where on the internet to search.
And we have almost no visibility into how or why these decisions are being made.
The hidden layer
This is quite different to how we’ve traditionally experienced finding information online.
When you search on Google, you can see the results it returns. You can see what ranks first, what ranks tenth, what’s sponsored, and you can decide which sources you trust enough to click on. Google also provides guidance on how its search systems work and regularly communicates significant changes to its algorithms. An entire SEO industry has developed around understanding, testing and optimising for these systems.
AI platforms are dramatically changing this dynamic.
When ChatGPT searches the internet, there are now a series of decisions happening before we ever see the answer. Does it need to search? What should it search for? Should it search the whole web or specific sites? Which information should it retrieve? Which sources should it use? And ultimately, what should make it into the answer?
Consumers only really sees the final stages of this process. Everything else is happening beneath the surface, and this is creating what I think of as a hidden layer of the internet. An increasingly important layer sitting between consumers and the open web that decides which parts of the internet an AI platform sees before it decides what to tell us.
Citations give us some visibility into the end of this process, but they don’t tell us anything about what happened before it. And as more of our interactions with the internet are facilitated by AI platforms, that distinction is going to become increasingly important.
An invisible information marketplace
There’s another dimension to this hidden layer that I’ve written about previously.
Two weeks ago I wrote about TIME and Mobian experimenting with a new form of advertising designed specifically for AI platforms. Sponsored articles for brands like Ally Bank were being published on TIME.com but kept separate from the main website and made available for AI platforms like ChatGPT to discover and use as sources.
I called this Paid Organic: paid content designed to behave like organic information inside AI platforms.
But this isn’t just about advertising. It demonstrates how this hidden layer can potentially be influenced from both sides.
AI platforms are making decisions about what information they search, retrieve and ultimately use. At the same time, publishers and brands are starting to think about how they can deliberately put information into the places those AI platforms are looking.
Together, this creates the foundations of an entirely new information marketplace sitting between consumers and the open web.
There’s nothing inherently wrong with either side of this. AI platforms need to make decisions about which sources they trust, and brands have always tried to make their information discoverable. But as this new hidden layer develops, the decisions being made within it could become enormously valuable.
And unlike the search results, sponsored listings and webpages that built the commercial internet, almost none of it is visible to consumers.
What could this become?
To be clear, there’s currently no evidence that OpenAI is selling preferential access to ChatGPT’s search results or deliberately favouring commercial partners. That’s not what concerns me. What concerns me is what this hidden layer could eventually become.
Today an AI platform might decide to search particular websites because it believes they’re more relevant or trustworthy. Over time, it might develop preferred sources for particular types of information. There are perfectly legitimate reasons for doing this - if I’m asking a medical question, I’d probably prefer ChatGPT to prioritise the NHS over a random health blog. The concern is what happens if commercial relationships eventually start to influence these decisions.
OpenAI already has content partnerships with publishers around the world, including News Corp, Axel Springer, the Financial Times and TIME. There’s no suggestion that these partnerships currently influence which sources ChatGPT searches, but as AI platforms become an increasingly important way for consumers to access information, being inside their consideration set will become enormously valuable. Could publishers eventually pay for preferential access? Could commercial partners be prioritised over competitors? Could some sources effectively become more visible to AI platforms than others? And most importantly, would we even know?
This is where the lack of transparency starts to become more than an information problem and potentially becomes a competition problem. Google has enormous power because it decides what ranks, but AI platforms could have even greater power because they can increasingly decide what gets considered before anything is ranked at all.
A publisher, retailer, comparison service or brand doesn’t necessarily need to fall from position one to position ten. It could simply disappear from the information environment being used to construct the answer. As AI platforms increasingly become the interface between consumers and the internet, that gives a relatively small number of companies significant power over which information, products and businesses consumers discover.
That’s a very different kind of gatekeeping.
Transparency needs to go deeper
None of this means that AI platforms shouldn’t make decisions about which information they use. They have to. The quality of their answers depends on being able to identify relevant, trustworthy sources and filter out the enormous amount of unreliable information that exists online.
But as these decisions become more important, I think we need much greater transparency around how they’re being made. Citations are a good start because they tell us where the information in an answer came from, but they only give us visibility into the very end of the process. They don’t tell us why an AI platform chose to search one source rather than another, whether some sources are given preferential treatment, or whether commercial relationships have any influence over those decisions.
This matters because AI platforms are quickly becoming an important gateway to the internet. The more we rely on them to find information, recommend products and ultimately take actions on our behalf, the more power this hidden layer will have over what we discover and the decisions we make.
I’m not particularly worried about what the hidden layer means today. I’m concerned about the potential of what it could become tomorrow.
If AI platforms are going to become the gatekeepers to the internet, we need some visibility into how the gate works.
“The future is already here, it’s just not evenly distributed.“
William Gibson






