John Mueller and Martin Splitt from Google discussed LLMs.txt and markdown, with Mueller revealing an unexpected detail about the initial intention of LLMs.txt and also clarifying the significant deficiencies in the proposed standards.
What Discovery Entails and Its Significance
In the realm of information retrieval, discovery within search involves a search engine identifying the presence of a particular webpage, constituting a component of the broader search engine structure.
Search Engine Structure:
- Adding the URL to the crawl to facilitate discovery.
- Crawling involves downloading and analyzing the content.
- Analyzing the data and storing it in an organized database for easy retrieval is known as indexing.
- Ranking is the section that captures everyone’s attention.
- Serving is the final step involving displaying the ranked web pages in search results.
Discovery is the initial stage in the search process that leads to ranking and providing website links.
Discovery is essential for a webpage to be noticed and included in search results.
Discovery is not included in the proposed LLMs.txt standard, which highlights its significance.
The initial purpose of LLMs.
John Mueller mentioned that he had encountered one of the individuals involved in developing the LLMs.txt proposal, who clarified that the intention behind LLMs.txt was not to enhance a site’s discoverability.
Many website owners invest resources in creating LLMs.txt to improve their visibility and ranking on LLMs. However, this contradicts the original purpose of LLMs.txt, which is not related to discovery.
Mueller provided an explanation:
The proposal was not intended to facilitate search engines or LLM systems in discovering all content. Instead, it aimed to help LLM systems already familiar with a site to explore additional content.
Using this method to improve visibility for AI or search systems does not seem logical.
Mueller further clarified that numerous individuals are utilizing LLMs.txt with the intention of assisting the Discovery process, even though that is not its intended purpose.
LLMs.txt can’t be relied upon as they are based on the subjective information provided by site owners, which may not accurately reflect the content in the HTML.
He went on to say:
You are essentially informing these systems that you have the greatest website and listing all the necessary pages for visitors to access, along with encouraging them to purchase your products or services.
In an LLM system, it is inherently unreliable to rely on the content to distinguish between various websites.
Agentic Directions
Mueller suggests that certain standard proposals could be beneficial for assisting an AI agent, possibly referring to the Web Model Context Protocol (WebMCP).
He provided an explanation:
An automated system on your website can assist visitors in finding information on how to purchase a photograph, which would be beneficial for users looking for guidance on making a purchase.
The system will not search multiple websites for automated information when you ask where to buy a photograph; instead, it will seek out the best website.
LLMs.txt Does Not Involve AI Discoverability
Mueller returned to the topic of how individuals are misinterpreting LLMs.txt as a method to attract attention from AI systems.
He pondered this issue.
Optimizing to be found does not seem logical from that perspective.
When an agent visits your website, it appears to be an open topic for discussion currently, particularly regarding the proposal of LLMs.txt. Various JSON files and recognized file formats are being considered.
They offer a programmatic interface for specific URLs or mechanisms, similar to WebMCP.
I believe those are essentially separate conversations.
Discovery and ranking continue to be connected to HTML.
Mueller emphasized that Discovery operates at the HTML level.
He provided an explanation.
The basic SEO aspect of finding a website that offers photographs is mainly connected to HTML and standard web pages.
If a user chooses a particular service, there is more opportunity to assist an agent or a language model system in finding the best approach within that service.
Lots of ideas are being explored, but there is yet to be a definitive solution that everyone will adopt. It may take some time for these different systems to converge around a common standard file type or mechanism.
If AI agents become the main way users interact with websites, WebMCP could be more valuable than LLMs.txt for websites, especially ecommerce sites.
WebMCP is particularly suitable for e-commerce as it emphasizes providing AI agents with practical functions such as product filtering, searching, identifying, comparing products, and assisting in adding items to a shopping cart.
AI agents can use website HTML meant for humans to navigate with the help of WebMCP, unlike LLMs.txt.
LLMs.txt and WebMCP do not aid in a website’s discovery by AI as they were not designed for that purpose. The initial ranking stage, which involves discovery, relies solely on HTML. Given this scenario, what would be your next step?
Google’s Search Off The Record Episode 111 can be heard.
Image provided by Shutterstock/Master1305




