The Invisible Visitors

Why the future of AI may begin with the bots no one ever sees.
When website owners open their visitor statistics, they usually expect to learn more about their audience. How many people visited? Which articles attracted attention? Where did readers come from? Increasingly, however, those statistics reveal something unexpected. Many of the visitors are no longer people. They are machines.
Search engine crawlers. AI crawlers. Monitoring systems. Security scanners. Automated agents quietly moving across the internet, discovering, indexing and collecting information long before most human readers ever arrive.
Public debate focuses on foundation models. Policymakers increasingly focus on regulating artificial intelligence itself.
Yet another layer receives remarkably little attention. Before a model can learn… Before an algorithm can reason… Before artificial intelligence becomes intelligent… Someone—or something—must first read the world.
🟦 The Hidden Internet
Most people experience the internet through websites, social media platforms and search engines. Artificial intelligence experiences it very differently.
Before a foundation model can generate a single answer, enormous quantities of publicly available information must first be discovered, indexed, classified and organised. That work is performed not by people, but by billions of automated requests made every day by crawlers, indexing systems and software agents operating continuously across the web.
This invisible activity forms an infrastructure beneath the visible internet. Without it, today’s foundation models could neither be trained nor continuously updated. The chatbot may be what users see. The crawler is where the process begins.
🟦 From Readers to Crawlers
For decades, website analytics primarily reflected human behaviour. Today, many publishers discover something different. Alongside genuine readers appears a growing population of automated visitors. Some belong to search engines, ensuring websites remain discoverable. Others monitor cybersecurity, verify network integrity or archive digital content.
Increasingly, however, another category is becoming visible. AI crawlers systematically collect publicly available information that may later support indexing systems, retrieval services or the development of future foundation models.
In other words, websites are increasingly being read by machines before they are read by people. That represents a profound shift in the way knowledge circulates across the internet.
🟦 The Questions Europe Has Yet to Answer
The European Union has established one of the world’s most comprehensive regulatory frameworks for the digital economy through the Digital Services Act, the Digital Markets Act and the AI Act. Together these initiatives strengthen transparency, accountability and safety across platforms, markets and artificial intelligence systems.
Yet these frameworks understandably concentrate on what happens once digital services and AI systems already exist.
Far less attention has been devoted to the invisible infrastructure that continuously gathers the information upon which those systems ultimately depend.
As artificial intelligence becomes increasingly important for Europe’s competitiveness and digital sovereignty, several broader questions begin to emerge.
- Who governs the automated systems that continuously collect the information from which future AI models learn?
- Should AI crawlers become more transparent about what they collect, when they collect it and for what purpose?
- If foundation models increasingly depend upon publicly available knowledge, how should publishers, researchers and creators participate in that value chain?
- Can Europe meaningfully shape artificial intelligence without understanding the infrastructure that feeds it?
- And if large-scale data collection becomes a strategic capability, should it also become part of Europe’s discussion on digital sovereignty?
These are not simply technical questions. They are increasingly questions about governance.
🟦 The First Operational Layer
Artificial intelligence is often described through algorithms, data centres and ever more capable models. That description remains incomplete.
Long before a foundation model is trained, information must first be located, filtered, indexed and organised. Collection may therefore be the least visible layer of artificial intelligence—and perhaps one of its least understood.
Before algorithms learn, bots collect. Before models reason, crawlers organise. Before intelligence emerges, knowledge must first be discovered.
Collection may well be the first operational layer upon which every other layer of artificial intelligence ultimately depends.
Signal
Artificial intelligence does not begin with a chatbot. Nor does it begin with a foundation model. It begins with billions of automated visitors quietly traversing the internet every day, collecting the information from which future intelligence is built.
Europe has taken important steps towards governing artificial intelligence. The next challenge may be understanding—and eventually governing—the invisible infrastructure that allows artificial intelligence to learn in the first place.
Because before algorithms become intelligent… Someone—or something—must first read the world.
Editor’s Note
This Signal expands the discussion introduced in Who Builds Europe’s AI? by exploring a largely invisible layer of the AI ecosystem: the automated systems that collect and organise the information upon which foundation models depend.
Credit
Illustration: Altair Media (conceptual editorial artwork visualising automated web crawlers as the invisible collectors of the digital knowledge that powers artificial intelligence).
Caption
The invisible visitors. Every day, automated crawlers, scrapers and indexing bots quietly move across millions of websites, collecting and organising information long before humans encounter it. Before algorithms learn, before foundation models reason and before AI applications appear, this hidden layer silently builds the knowledge infrastructure upon which artificial intelligence depends.
