Before Google, the web was indexed entirely by hand
Today, automated crawlers map billions of pages in seconds. But in the early 1990s, finding information required manual effort and curated lists. Discover how we moved from human-edited directories to the complex, distributed computing systems that power our modern digital discovery.
The concept of automated retrieval predates the modern web. In 1945, Vannevar Bush envisioned the 'memex,' a system designed to navigate expanding scientific indices using connected annotations similar to today's hyperlinks. By the time the web emerged, early tools like Archie (1990) focused on searching FTP file names rather than web content. As the web grew, the manual process of maintaining a central list of servers, once managed by Tim Berners-Lee at CERN, became impossible to sustain.
The transition to automation arrived through 'web robots' and crawlers. In 1993, Matthew Gray produced the World Wide Web Wanderer to measure the web's size, while JumpStation became the first tool to combine crawling, indexing, and searching. Innovation accelerated with WebCrawler in 1994, which introduced the ability to search any word within a page. This era also saw the rise of link analysis, where algorithms like Robin Li's RankDex (1996) used hyperlinks to measure website quality—a technique that heavily influenced Google's PageRank.
Modern search engines are massive, distributed computing systems. They rely on automated 'spiders' that follow directives in robots.txt files to crawl sites, extracting metadata, headings, and text. This process is continuous, as engines must constantly update their indexes to reflect the changing web. While Google has dominated the market since the 2000s, holding roughly 89–90% of the global share as of May 2025, the landscape is shifting toward AI-powered assistants that integrate large language models to provide conversational, contextual responses.
Source: Search engine