Categories: CMS

NLWeb: The first step toward AI-ready Knowledge Layer

AI is creating a new front door to the web. Every major shift on the web has changed how information is discovered and consumed. HTML gave browsers a way to render pages. Schema.org gave search engines a way to understand them. Now, AI agents need something more: a way to access, understand, and interact with a website’s knowledge directly. 

That’s where NLWeb comes in. 

What Is NLWeb?

Built by R.V. Guha, the engineer behind RDF, RSS, and Schema.org, NLWeb is an open standard that lets a website expose their knowledge and capabilities through a natural language interface so AI agents can query it directly instead of having to reconstruct answers from web pages. 

It creates a common way that allows AI systems to discover, retrieve, and interact with trusted information more efficiently through natural language.  

Functionally, NlWeb sits on top of the structured information many organizations already maintain, such as Schema.org markup, product catalogs, business profiles, knowledge graphs, and other structured content. Each NLWeb instance can be exposed as a Model Context Protocol (MCP) server, which remains technology-agnostic across vector databases and LLM providers.   
 
Rather than forcing AI systems to reconstruct meaning from HTML, websites can make their information directly accessible in a structured, machine-readable form. 

When a user asks, “What’s the best family resort near Yellowstone with pet-friendly cabins?” the AI doesn’t simply search for keywords. 

It needs to understand: which businesses exist; what products and services they offer, relationships between locations, current availability, trustworthiness of information, whether the information is recent. Much of this information already exists on company websites.  
 
The problem is that it was created primarily for people, not for AI systems. 

As a result of which every AI model spends considerable effort reconstructing meaning from HTML, documents, and fragmented data before it can even begin answering the user’s question. 

What if your website could make that knowledge directly accessible to AI? 

How NLWeb is Helping AI Understand Your Content with Less Guesswork

Every AI-generated answer begins with retrieval. Before ChatGPT, Gemini, Claude, or Copilot can answer a question about your business, they first have to build an understanding of it. 
 
Unlike traditional search engine, which retrieves webpages, an AI system retrieves information from multiple sources: your website, product pages, blogs, reviews, business directories, news articles, knowledge bases, and countless third-party sources.  It then resolves entities, connects, evaluates facts and builds an answer. 

Dig deeper: Why entity authority is the foundation of AI search visibility 
 
In other words, AI doesn’t simply retrieve your business. It constructs an understanding of your business.

When business information is fragmented, unstructured, inconsistent, or incomplete, the model is forced to fill in the gaps. Different AI systems may retrieve different sources and apply different confidence thresholds, leading them to construct different understandings of the same business.  
 
This is what we mean by the “inference tax.” 

Every retrieval requires inference. Every inference introduces uncertainty. 

NLWeb can reduce that tax by giving AI systems a direct, natural-language interface to the knowledge a website makes available. Instead of relying entirely on the model to reconstruct meaning from pages and fragmented sources, the website can provide a direct path to that knowledge. 

This matters because uncertainty can lead to hallucination. Hallucinations are often described as models “making things up.” In reality, many hallucinations are a consequence of incomplete grounding. When the underlying facts and relationships aren’t explicitly available, the model has to infer what is most likely to be true. Sometimes those inferences are correct. Sometimes they aren’t. 

Dig deeper: How to give AI the context it needs  

Future Readiness: Build Your Knowledge Once. Publish It Everywhere.

If your entities are accurate, your relationships are complete, and your information is current, NLWeb becomes the interface through which AI systems access that knowledge hence reducing unnecessary inference and opportunity for error.  

NLWeb is one important step in making websites AI-ready but over the past year we’ve seen the emergence of a growing ecosystem of AI standards, including llms.txt, WebMCP, Entity Maps, Agentic Resource Discovery (ARD), Unified Commerce Platform (UCP), Agent Communication Protocol (ACP), agents.md, and many others.  
 
Each serves a different purpose, but together, they point toward the same fundamental shift:  

How do businesses expose their knowledge in a way AI can consume directly? 
 
The mistake businesses make is optimizing for today’s protocol independently. Instead, they should prepare for tomorrow’s AI ecosystem.  
 
We’ve been testing this principle in the real world. Over the past several months, Milestone has been working closely with Microsoft to explore what NLWeb looks like in production. We wanted to understand how AI agents discover, interact with, and consume websites that expose structured knowledge through NLWeb and Model Context Protocol (MCP).   
 
Our early deployment results include: 

  • 27% increase in AI traffic
  • Quick organic discovery by all five major AI ecosystems – Google, OpenAI, Anthropic, Microsoft, and Meta
  • More AI requests reaching structured knowledge endpoints
  • Higher confidence in retrieval through grounded business knowledge
  • Quicker Knowledge Graph updates being consumed

The strongest validation is that AI ecosystems are beginning to discover, query, and interact with structured knowledge endpoints organically, without any special integration or promotion.   
 
These are still early results, but they reinforce an important signal: the web is beginning to evolve toward more standardized, machine-readable interfaces for AI. 

How We’re Building It at Milestone

And this is where our thinking has evolved.  Rather than implementing every emerging standard independently, we believe organizations should build the knowledge foundation once and make it available through whichever interfaces AI systems require. 

At Milestone, we believe the answer is to build a canonical AI Knowledge Layer. It is a scalable, machine-readable repository that becomes the single source of truth for everything AI needs to understand about a business. That knowledge is created once, governed continuously, and published through whichever machine-readable interfaces the ecosystem requires. 
 
Every time a new protocol emerges. The publishing layer evolves but the knowledge layer remains intact.

This philosophy has shaped how we’re building the Milestone AI Platform.  At its core is a canonical AI Knowledge Layer that continuously maintains trusted business knowledge and publishes it across multiple AI-ready formats and protocols enabling.  

  • A single source of truth for business knowledge
  • Consistent representation across AI ecosystems
  • Less repeated inference by AI systems
  • Better grounded responses
  • Faster adoption of emerging standards without rebuilding data models

This is the architectural shift every organization should begin preparing for.  The organizations that succeed won’t be the ones supporting the most protocols. They’ll be the ones building the strongest knowledge layer. 

Dig deeper: Your website isn’t ready for AI agents — here’s what needs to change 
 

Timothy Talreja

Share
Published by
Timothy Talreja

Recent Posts

Webinar: 5 Must-Haves for Hotel Websites in the AI Era and Product Roadmap

AI is changing how travelers discover, evaluate, and book hotels, and the hotel website must…

3 days ago

Milestone CMS Is Now PCI DSS 4.0.1 Certified

We're happy to share that Milestone has earned PCI DSS 4.0.1 certification, confirmed through an independent…

1 week ago

Webinar: The Hotelier’s Guide to Search, AI Visibility, and Budgeting for 2027

Hotel budgeting season is here. Join us on Thursday, August 13, for a practical webinar…

2 weeks ago

Webinar Recap: The AI Visibility Flywheel – How Brands Get Featured in AI Search

In this session, Milestone experts Benu Aggarwal, Bill Hunt, and Ritika Chugh discussed how AI…

3 weeks ago

From Search to Action: What Google I/O 2026 Means for Your Digital Presence

Google I/O 2026 marked a turning point, not an incremental upgrade. AI is no longer…

2 months ago

Webinar Recap: Is Your Historic Hotel Visible in AI Search?

In this session, Milestone experts Mike Supple, Aparna Iyer, and Brandon Ahearn discussed how AI…

2 months ago