Digital Marketing

Why Data Integrity is the New SEO Tech: From Crawling to Trust

Two years ago, Google dropped support for nine ItemTypes in its rich search gallery. This happened not too long after the introduction of ChatGPT and when mass adoption began:

Photo from the author, July 2026

The question remains whether this decline will continue, but the recent removal – FAQ/FAQPage – has since sparked a debate about the role of schema.org in the future of Search.

Schema Is Dead, Right?

While others are doing tests and experiments to understand if the schema is really making a positive impact on the citations between the responses of the field, Gianluca Fiorelli pointed out that it is possible to do these tests on limited datasets. With that in mind, let’s remind ourselves of the wording of the revocation message for the rich results of the FAQ:

“…We will be deprecating the FAQ search feature, rich results report, and support for Rich results testing in June 2026.”

Notice here what they didn’t say – i.e. the use of the FAQ schema is no longer necessary. This is because the cancellation is only for the rich effects – the display feature. The schema itself is a cognitive layer – identifying entities and the relationships between them. Is the schema dead? In my opinion, it is far from it. While some properties are being withdrawn, others, like Product, are being expanded.

That being said, I also realize that adding schema is not a magic bullet that contributes to citation growth. However, that growth goes beyond the metrics we used to rely on like citations, impressions, etc. Suganthan Mohanadasan wrote a great piece about how schema has three “lives”:

  1. Google index pipeline.
  2. LLM pretraining (directly).
  3. LLM to return working time.

SEOs used to focus on the No. 1 as a positive contributor to success metrics. But the schema goes beyond what we are used to, or more accurately, report on. Schema is immortal; the one aspect of the display that it has benefited from is reduced instead.

Biggest SEO Threat: Ambiguity

A schema is an ontology, like a web standard, that can contribute to understanding data integrity. Data integrity risk is ambiguity. Ambiguity leads to illusions. Snowball Hallucinations. Finally a combination of results, which can lead to wrong results or even wrong pre-training for LLM which can have long lasting effects.

If an agent can misread you, at some point they will. LLMs can then run the risk of falling into a “semantic drift” that deviates from facts and favors narrative. This is explored in the piece entitled “Sangue e Grafi: Teaching a Small Model to Read the Bloodline” by Andrea Volpini and Chiara Carrozza where borderline models tend to fall short of telling the facts, while the small model provided by the information graph tools is equivalent to them.

Sangue e Grafi, by WordLift
Photo from the author, July 2026

→ Further reading: Information Retrieval Part 1: Consistency

5 Layers of Data Integrity

All this proves my belief that the role of SEO is to increase the integrity of the data, which is the schema that plays its role. Below, I show five layers of what data integrity can include:

5 layers of data integrity
Photo from the author, July 2026
  1. Businesses: What is there, and what is that thing. Object, Organization, and Person, anchored with @ids and linked to Wikidata, GS1, ISNI, or ORCID so the agent knows your “Apple” to the fruit.
  2. Relationship: How those things come together. @id and sameAs, RDF. Schema integration feature for Yoast SEO and EntityMap by Dixon Jones.
  3. Format: How the property is divided and assigned. JSON-LD, RDFa, and Microdata. Markdown, too (including LLMs.txt, agent.md, OKF) and endpoints (content dialogs, ARD, MCP).
  4. Actions: What can be done, is announced to the agents. Schema.org actions like BuyAction, and the new WebMCP, ACP, and UCP.
  5. To see: Basis, third person perspective, emotions, etc.

Synthesis, Guidance, and Implementation

In a post I wrote in October of last year, I said, “SEOs will have to consider both sides of the web and how to use both.” Emerging agreements (all launched in the last two years) reflect this fact, with the new “agent base stack” typically taking on one of three objectives:

The protocol The goal What it does
sitemap.xml Integration Every canonical URL is a single XML index.
llms.txt Integration A summary of the site’s content with important information and links to further reading.
Yoast Schema Aggregation Integration The JSON-LD page level is a single graph connected to the entire site.
EntityMap Integration The site’s business announcements on a single graphic map.
Information catalog Integration Structured, unstructured data, and SaaS have become the dominant context engine.
OK Integration The site information has been a bunch of bookmarks in /okf/.
ARD · ai-catalog.json Integration Domain tools and agents in the catalog; federal registries in addition to it.
OpenKB Integration The source documents are compiled into a markup wiki.
Schema.org Guidance A shared vocabulary that tells machines what things mean.
agents.md Guidance How agents should represent and interact with you.
Markdown for Agents Usage The same URL is used as a clean tag for content discussions.
Some Markdown output Usage Alternate .md version linked with rel=alternate.
/crawl at the end Usage It provides a page, or an entire site, as a clean search engine.
WebMCP Usage It exposes site actions such as tools that an agent can invoke.
NLWeb Usage Import schema, feeds, and sitemaps to answer natural language queries.
ACP Usage Agent output within ChatGPT against merchant product data.
UCP Usage The common language of agent’s commercial actions in all areas.

These three principles help to reduce the number of requests while increasing the efficiency of tokens. Some of the principles above are detailed within Search Engine Journal, including mine on ACP and UCP and Emina Demiri-Watson’s article on OKF, ARD, and others earlier this month.

But there is something missing in these protocols that have…

There is no Consensus or Standard Agreed upon

Schema.org was born out of a partnership between Google, Microsoft/Bing, and Yahoo! (Yandex joins later) who launched under joint governance. The same thing happened five years earlier with an XML sitemap. When search engines needed a ranking, they just sat down and built one – together.

Nothing like this is happening now, and it comes to the detriment of SEOs who want honest clarification on what should and should not be used on the sites they work on. Even the most basic facts about usage are contested, of which the argument about putting down is a good example of this.

While these discussions are ongoing, there is no place where the platforms meet and agree on a single international standard. The ecosystem has changed so much that these companies are no longer in the business of Search and the beauty of the web, but now they must look at how their businesses affect jobs, the economy, livelihoods, and the future of humanity as a whole. Therefore, I do not believe that the questions asked by SEOs are high on their priorities.

What Can You Do About It Now?

Looking back at the five layers of data integrity, the four you control can be seen below when looking at what a basic agent stack might look like:

Agent base stack options
Photo from the author, July 2026

There’s a lot to consider, and they all have different goals and technical debt. Decide which ones work best for you, and take whatever doesn’t require too much technical credit.

If I had to pick an order, it would be this.

  • Fix your @ids and add sameAs links to Wikidata and other authorities first, because everything else rests on it.
  • Then check how it is actually interpreted, with NLWeb, rather than assuming that the graph reads the way you intended.
  • If you’re in ecommerce, check the product feed before touching anything shiny, and check out BuyAction while you’re at it: only ReadAction and SearchAction are used at any real scale today, so the field is really wide open. Check out the latest news about what’s been added to the product schema.
  • Attempt to use WebMCP. It can be done on any website, and it doesn’t have to be ecommerce.
  • Markdown serving and content discussions can wait until you have the engineering power to spare (unless you can use Cloudflare’s Markdown for Agents).
  • Look at OKF and ARD. When Google introduces new protocols, I always pay attention – especially when it comes to how an agent or LLM understands the site as a whole.

None of this is a bet on a specific protocol. Using any of these minimizes the risk that the LLM will have to go “a long way” to make the answer it wants to answer. By betting is betting to accept any agent that appears anywhere of the platform, this will also help you think more about how your site can be interpreted by them, and how this improves if these conventions are accepted.

Even if you want to do small tests away from large sites, it is worth checking not only to see if there are good results in it, but also to understand how they all work in practice. This is exactly what I did recently, when I rebuilt my personal website, with several “agent-friendly” principles in place, including content negotiation, markdown alternate, llms.txt and WebMCP.

Don’t Rush the Protocol, Manage the Layers Underneath

Currently, it appears that there is no single “winner” who will advance from proposal or agreement to the web standard. There is no consortium to replicate what was done with XML sitemap and Schema.

The stack is now huge, but there is one thing they all share. Whether it aggregates, directs, or consumes, each is provided with the same substrate: Intuitive entities, transparent relationships, and content that a machine can read without guesswork or additional research.

Rankings were the success metrics of the old web. Trust, integrity, accuracy, and authenticity. Achieving this is still the role of SEO.

Additional resources:


Featured image: hmorena/Shutterstock

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button