Part I: Before Artificial Intelligence Learned to Read
By Francis Martinson, Contributing Author, The August Dispatch
Every technological revolution begins with an invisible agreement.
The Industrial Revolution depended upon an agreement that factories could transform raw materials into manufactured goods. The Information Revolution relied upon an agreement that digital networks could exchange information freely across geographical boundaries. The Artificial Intelligence Revolution, however, depends upon a far more complicated agreement, one that has remained largely unexamined until recently.
Who owns knowledge once it becomes publicly accessible?
For decades, the internet operated under an implicit understanding between content creators and technology companies. Publishers produced information. Search engines indexed that information. Users searched for information. In return, publishers received traffic, visibility, and opportunities for monetization. Although imperfect, the relationship was largely symbiotic. Search engines did not seek to replace publishers. They sought to direct users toward them.
Generative Artificial Intelligence has fundamentally altered that relationship.
Instead of merely locating information, AI systems increasingly consume it, learn from it, synthesize it, and generate new content that may reduce the need for users to return to the original source. This transformation has created one of the most significant governance challenges in the history of the internet.
The legal disputes now unfolding between AI developers and content creators are often portrayed as copyright battles. While copyright undoubtedly occupies the legal foreground, reducing these disputes to intellectual property alone overlooks a deeper transformation occurring within the digital economy. The central issue is governance.
- Who controls data?
- Who authorizes its use?
- Who benefits economically from its transformation into artificial intelligence?
- Who bears responsibility when those boundaries become blurred?
These questions did not emerge overnight. They have been developing for decades, hidden beneath the rapid evolution of search technologies, cloud computing, machine learning, and increasingly sophisticated language models.
In my earlier article, The Ethics and Implications of Data Scraping and Manipulation on Social Media: A Case Study of Twitter (X), I examined how large-scale data collection raised concerns extending well beyond technical capability. The article argued that the ethical implications of data scraping could not be separated from broader questions surrounding consent, transparency, governance, and societal trust. At the time, those concerns were frequently regarded as theoretical discussions surrounding emerging technologies.
Today, they have become central legal, commercial, and geopolitical questions.
This article revisits those earlier governance concerns through the lens of recent developments within the artificial intelligence industry. Rather than examining one company or one lawsuit in isolation, it explores a broader transformation in how society understands digital information, intellectual property, and machine learning.
The objective is not to determine whether artificial intelligence should continue advancing. That debate has largely been settled.
The objective is to examine whether the governance structures developed for the internet remain adequate for an era in which machines no longer simply retrieve information. They learn from it.
The Internet Was Never Designed for Artificial Intelligence

To understand the present conflict, it is necessary to understand the original architecture of the World Wide Web.
When Sir Tim Berners-Lee introduced the World Wide Web in 1989, his objective was remarkably straightforward. Researchers required a standardized method of sharing documents across interconnected computer networks. The web was conceived as an information-sharing platform rather than a commercial ecosystem.
As internet adoption accelerated throughout the 1990s, millions of webpages emerged across universities, governments, businesses, and personal websites. Information became increasingly decentralized. This abundance created a new problem.
Finding information became almost as difficult as creating it.
The earliest search technologies attempted to solve this challenge through relatively simple methods. Services such as Archie, Gopher, Lycos, AltaVista, Yahoo!, and later Google developed automated programs, commonly known as web crawlers or spiders, to navigate websites, collect publicly available information, and organize it into searchable indexes.
These crawlers did not exist to appropriate knowledge.
Their purpose was discovery.
Every webpage identified by a crawler increased the probability that a user would eventually visit the originating website. Publishers generally benefited because improved discoverability translated into greater readership, advertising revenue, subscriptions, or commercial transactions.
An unwritten agreement emerged:
- Website owners permitted search engines to crawl publicly available pages.
- Search engines indexed those pages.
- Users clicked search results.
- Traffic returned to the publisher.
Although disagreements occasionally arose concerning ranking algorithms, advertising, and search dominance, the underlying relationship remained mutually beneficial.
The internet’s information economy was therefore built upon reciprocity rather than replacement.
The Social Contract of Search
This informal relationship was reinforced by technical conventions rather than legislation.
One of the most influential examples was the Robots Exclusion Protocol, commonly known through the robots.txt file.
Website administrators could voluntarily indicate which portions of their websites automated crawlers should or should not access. While not legally binding, responsible search engines generally respected these directives.
The existence of robots.txt reflected something important.
Governance on the early internet often emerged through cooperation rather than regulation. Technology companies, publishers, software developers, and standards organizations collectively established expectations that enabled innovation while minimizing conflict.
Google’s rise during the late 1990s further strengthened this ecosystem.
Rather than merely cataloging webpages alphabetically, Google’s PageRank algorithm evaluated hyperlinks as indicators of authority and relevance. Search quality improved dramatically. Users discovered information more efficiently. Publishers gained unprecedented global visibility.
For nearly two decades, this model transformed the internet into the world’s largest publicly searchable repository of human knowledge.
Few questioned its underlying assumptions because every participant generally derived value from the exchange.
Search engines gained users. Publishers gained audiences. Users gained information.

The relationship functioned because each participant continued depending upon the others.
Artificial intelligence would eventually disrupt that balance.
From Indexing Information to Learning From Information
The distinction between search engines and modern artificial intelligence appears subtle to many users.
Technically, however, the distinction is profound.
Traditional search engines primarily function as navigational systems. They identify relevant information and direct users toward its original source.
Large language models function differently. Rather than pointing toward knowledge, they statistically learn patterns embedded within enormous collections of text, images, software code, academic publications, books, news articles, discussion forums, and countless other digital artifacts.
Information ceases to be merely indexed.
It becomes training data.
That distinction changes everything.
When an AI model processes billions of documents during training, it is no longer acting solely as an intermediary between publisher and reader. It becomes an interpreter capable of generating responses that synthesize information drawn from countless sources simultaneously.
The economic relationship therefore changes. Publishers may receive little or no traffic from information that once required direct engagement with their websites. Artificial intelligence creates value by transforming data into predictive capability.
Consequently, the question is no longer whether information is publicly accessible. The question becomes whether public accessibility alone provides sufficient justification for commercial AI training.
That question represents the beginning of the modern AI governance debate.
This is Part I of a series. Part II continues the conversation.
Join The August Dispatch Newsletter
For curated news, practical prompts, and how‑to guides—plus first dibs on workshops.





Leave a Reply