
Welcome to another episode of The Search Session. I’m Gianluca Fiorelli. Joining me are Martha van Berkel, a semantic SEO and structured data strategist, and Mark van Berkel, a semantic technologist, both based in Canada.
Martha and Mark join me to explore how semantic search and structured data are changing as AI becomes part of search and the wider web. The conversation connects knowledge graphs, data governance, and agent-ready websites to one central question: what does it really take to keep AI from getting confused about your brand?
What you’ll learn in this session:
Why semantic search is moving beyond rich results: connected structured data, semantic data layers, and knowledge graphs are becoming part of how brands manage and communicate their information.
What business leaders need to understand about semantic search: technical explanations should follow the business need and focus on the value of a high-quality, governed data layer for AI.
How connected schema markup creates a governed data layer for AI: @id, sameAs, taxonomies, and shared entity definitions help unify facts across pages, reduce conflicting information, and support consistent brand understanding.
What data governance requires from enterprise schema markup: clear ownership, automation, standards, scalable maintenance, and consistent entity definitions are needed to prevent schema drift and support a reliable data layer for AI.
Why the semantic web’s “alphabet soup” needs a maturity framework: connecting schema markup, entity management, NLWeb, and MCP to business use cases helps SEOs present the semantic data layer as a strategic marketing asset.
What role Open Knowledge Format may play in the emerging data ecosystem: it may support internal knowledge structuring and web-accessible data, while MCP offers a more practical way to package information around specific user needs.
Why Schema.org still matters for AI understanding: search indexes use structured data to organize information, while LLMs recognize its vocabulary from training, helping improve AI visibility and reduce hallucinations.
What agentic interactions still need from semantics: accessibility metadata, Schema.org, and SKOS help structure content and entity relationships, while standards for agent actions—and how those actions lead to conversions—remain the next frontier.
How to balance natural-language creativity with machine clarity: entity disambiguation, consistent definitions, and structured relationships give AI a reliable foundation beneath the brand’s language and tone.
My guests bring years of semantic SEO experience to a clear conversation that goes beyond the buzzwords. Give it a listen.
Topics covered: semantic search · structured data · knowledge graphs · AI hallucinations · Schema.org · taxonomies · ontologies · data governance · schema drift · entity management · agentic AI · NLWeb & MCP · natural language
About the Guests

Martha van Berkel
CEO & Co-founder of Schema App
Martha has more than 25 years of experience in enterprise technology, SEO, and semantic search. She is known for helping brands use structured data, entity linking, and knowledge graphs to improve how search engines and AI systems understand their content.
In 2011, she co-founded Schema App with Mark van Berkel. The company helps enterprise organizations create, deploy, and govern connected schema markup and semantic data layers at scale, providing search engines and AI systems with clearer, more consistent information about their brands.
Martha writes regularly for Schema App about schema markup, knowledge graphs, and AI search. She also contributes to Search Engine Journal and speaks frequently at industry events.

Mark van Berkel
CTO & Co-founder of Schema App
Mark has more than 20 years of experience in software development, enterprise consulting, and semantic technology. He is known for building knowledge graphs, ontologies, and structured data systems that help search engines and AI understand enterprise content.
In 2011, he co-founded the company that became Schema App with Martha van Berkel, launching the platform in 2016.
Mark writes for Schema App about Schema.org, knowledge graphs, ontology design, and semantic infrastructure for AI. He also speaks at industry events and webinars on semantic technology, data governance, and the agentic web.
Video Chapters
Transcript
Full conversation between Gianluca Fiorelli, Martha van Berkel, and Mark van Berkel.
Gianluca Fiorelli: Hi, I'm Gianluca Fiorelli, and welcome back to The Search Session. Today, we are going to have a very special episode. Why? First of all, because I'm not going to have just one guest. I'm going to have two guests, which is something that I was doing in the really, really beginning but haven't done for a long time.
Not only this, these two guests work in the same company, SchemaHub. One is the CEO and one is the CTO. And not just that, this couple is a 100% true, literal couple. They have been happily married for a long time, and maybe I will ask at the very end of the episode how this coexistence between professional and private life is.
What is your secret? I will ask that. So please welcome, with really warm cheers, Martha van Berkel and Mark van Berkel. Hey, how are you doing?
Martha van Berkel: Yes, great, thanks. Super excited to be here and excited to bring both of our perspectives. So thanks for being open to the conversation. We're pumped.
Gianluca Fiorelli: And are you ready? Mark, how are you doing?
Mark van Berkel: I'm good, and I'm excited for the conversation as well. I appreciate making one-third space here for me, and I think it'll be a good conversation. I've been eager to have this opportunity to talk about things that I'm really interested in, and also, on the personal side, how we handle being a couple and professionals. I think that's actually interesting too.
Gianluca Fiorelli: Sure, sure. And are you ready for vacation? Off the record, you were saying, Martha, that you are planning at least a few days off. Also, happy birthday. Next week is going to be your birthday.
Martha van Berkel: We're excited to spend some time on the water. In Canada, it's beautiful summertime, and so we're just going to take a break, step out of the life of semantics, enjoy real life, and spend some time with our kids.
Gianluca Fiorelli: Yes, yes, discovering other types of entities.
Martha van Berkel: Yes. I love that.
Gianluca Fiorelli: Okay. Let's start with a classic question I ask all my guests. How is SEO treating you lately, after all these already quite many months of craziness we are living in our industry?
Martha van Berkel: All right. Shall I go first?
Gianluca Fiorelli: Sure.
Martha van Berkel: So I'm loving it. We've been thinking about semantic search and thinking about how the role of schema markup plays a much bigger strategic role in building a knowledge graph for over a decade now. And so it's sort of like I've been watching my watch, being like, “When is the time going to come?”
While I would say some people are seeing lots of change, as an entrepreneur, I like change. I like the fact that we have to stay on our toes, pivot, and continuously learn. I do find there's a lot to learn. My brain gets tired sometimes trying to keep up with all the things. But I'm bullish, excited, and optimistic, just because I feel like the time of knowledge graphs is here. So I think that's my perspective.
Gianluca Fiorelli: And what about you, Mark?
Mark van Berkel: I would say very similar. I think of where we came from. We used to talk about how we help explain your content so it can be understood by machines, and then we veered from years just to say, "Google," and simplify the language around, “We're helping Google,” because they were pioneers in understanding your website through structured data and schema markup.
We've always been working with these applications and systems through knowledge graphs. And so it was probably overbuilt early in our journey, but at this point, it seems more relevant than ever because now we go back to trying to explain your content to machines, and what we've been doing is probably more relevant now in how we have done things.
SEO is treating me nicely, but it is a state of chaos, I think. There is a lot of change, and it is a lot of work for all of us professionals in the industry to keep up and know what's relevant, what's noise, and what's signal. But, as Martha said, as an entrepreneur, I lean into that research and try to keep on top of what's relevant and what's new. So I'm enjoying it, but to be totally honest, the environment in which we work is challenging.
A New Business Case for Structured Data
Gianluca Fiorelli: Yes, I think you're very right. Somehow, I had the same sensation, because, even if not with your depth, I also preached a lot about the benefits of good semantic search in the context of organic search in general. But yes, 10 years ago, or even more, it was very complicated to sell this.
What has changed in terms of selling correctly- not overselling, but selling correctly- the concept of semantic search as one of the pillars of this new era of search, compared with before? What kind of misconceptions do you see from business owners or even from other SEOs who may be contacting you to start a collaboration or use your tools and services? What kind of misconceptions are you still seeing in relation to semantic search? What kind of myths? What kind of wrong assumptions do you see?
Martha van Berkel: I think it's interesting. Ten years ago, people were primarily buying for visibility, like rich results, right? And I would say the measures were really clear, but people were still sort of hesitant. The thing that was easy about selling over the last decade is that Google had it documented, saying, “Do structured data to get these rich results.”
And so, to me, that was like, if you do this, this is the measurable result. I think what's harder now is that there are so many things, like, “Do these things for AI and for understanding.” And I would say, especially for the executives who are buying, it's sort of like, what do I listen to, right?
We were talking about how, as entrepreneurs or specialists in this area, there's a lot to consume and understand. Well, imagine if this is only a little part of your job because you're the CMO and you own all of marketing. So I think the challenge now is being able to really hit the need.
What I'm hearing resonate right now is that—and we have some strong case studies, and there's one with Wells Fargo around—when we connected the schema markup for a location, we were able to solve some hallucinations in AIO.
And when people think about, “How do I reduce risk around AI and around my brand being understood by AI?” and those kinds of stories. Someone said to me just this week, “So we could be doing nothing wrong, right? We could have no schema markup. We could just be doing nothing wrong, and there could be answers wrong out on the website just because AI has done the wrong inferencing.”
And so I think now my whole thing is that structured data, doing it semantically, and building a semantic data layer is something you can take action on and control to reduce risk. Again, that's very macro and sort of what I see as different now.
And then I would say all the foundational work that we've been doing over the years, whether that be connected schema markup, defining entities, or thinking about it as a knowledge graph, we've been talking about that for a decade. There's also tons of research, and maybe, Mark, you want to speak about the research showing that knowledge graphs help AI give more accurate answers. So I think that combination of things is really helping.
Discover what’s driving visibility.
Advanced Web Ranking helps you explore the impact of structured data, semantic SEO, and AI optimization across both traditional search and AI search experiences. Try AWR free.
The Growing Sophistication of Semantic Search
Gianluca Fiorelli: Yes. In fact, I was going to ask you, Mark, from a technical standpoint and in terms of the technical knowledge of what semantic search is and what its purpose is, did you see an evolution in that knowledge, or do you think that people still have to study a lot about it?
Mark van Berkel: There's a lot of prior art, and there's been a lot of work in what I'm going to call the semantic web community. It really has a long history. It started around 2000, so 25 or 26 years ago, when there was a popularized article in Scientific American from a couple of the pioneers in the space that created some interesting technologies that are still really useful today.
But there was another winter in AI during that time period. In the last 10 years, there's certainly been more. It's been interesting. Ten years ago, we couldn't really talk about these things or how we operate or how we do things. And now, I don't think we should still lead with how we solve things. We're going to lose people in the definitions of technology, which isn't really helpful in moving a conversation forward.
But there is a lot more knowledge, for sure, about the terminology, like talking about a knowledge graph or talking about an ontology, or really thinking about semantic search in different ways from how it has been used historically. And I think there's a lot more sophistication among buyers. I'm not the salesperson here. Martha answered it great.
I think it's hard to really lead with how we solve things, but certainly, once they start to ask questions, it becomes great to be able to share and help evangelize the technology set using the semantic web principles that have been around for a while.
And I think giving a high-quality data layer to AI is really the main benefit. Using things like GraphRAG, you can help achieve some of these outcomes. But also, just having a governed data layer is another part of it.
How are you really understanding the content that's out there? Maybe you just have an incoherent mess, to be honest. Some pages might talk about the same thing in very different ways and have incorrect facts.
So there are ways in which you can use the tech stack. But again, the sophistication is coming. I think it's more fun to be able to nerd out once they get past the early sales part of it.
Martha van Berkel: Can I add something, Mark? I think the piece that—and this is maybe where SEOs are stuck today—SEOs are still thinking of schema markup as an SEO thing. And I would argue it's not an SEO thing anymore.
Gianluca Fiorelli: Yes, I agree. In fact, I was going to say something along these lines because I think that, especially when we are talking with clients, so we are the SEO consultant, for instance, or the SEO agency, maybe the mistake, also in order to sell it, to sell the semantic web—let's call it semantic in general, without putting web, search, or any other words along with it—is that structured data, as a thing which is not just schema, is something that is also useful for search, but it's not meant just for search.
For instance, it can be useful for creating an internal knowledge graph that you can use for better governance of your own information about your catalog, your services, your knowledge base, and so on. So you can also use it with an internal value for the company.
Then you can eventually use it, and I think this is the most amazing thing to discover different ways to create an architecture of the content you have.
Martha van Berkel: Your taxonomy, right?
Gianluca Fiorelli: Yes, starting from the ontology and going to the taxonomy, and then doing true information gain, which is not just in a single piece of content.
I usually love to make two examples. One is Airbnb, when they invented this taxonomy based on weird ways of classifying an Airbnb, like houses in trees, houses in igloos, and all these kinds of things.
Another example that I usually use is Atlas Obscura, which is a wonderful website that has a very quirky and weird taxonomy based on the knowledge of all the UGC content that is hosting. So I think that is the most interesting use.
Martha van Berkel: Yes. One of the things I like to talk about, even if we're an SEO, is that this is the semantic data layer for marketing, right?
There are lots of people talking about the importance of semantic layers for grounding AI. So I would love SEOs to adopt this. This is your content knowledge graph or your marketing knowledge graph, right? And schema markup, or Schema.org is the vocabulary with which we're articulating that.
The other piece I love is that, especially when we get to taxonomy, what I'm hearing marketing leaders ask about is, “What am I an authority on?” And this is just E-E-A-T that we've been talking about for years, but now, with a knowledge graph, we can actually articulate and have measures for what your data is actually saying.
And what I find so interesting—we do this at Schema App because it's what we love to do—is that everyone is often looking at, “What does AI externally see me as an authority on?” instead of actually looking at your own data. So I love that it gives you that data source to ask questions of your own data.
Indeed. And I think connected to this concept is the very recent news—and today, the day of recording, is the 29th of July—that everybody can go and create a property for their own social media profile. So you can see the traffic from your other identity, your facet of your, let's say, rented entity in a rented space like X, Instagram, TikTok, or YouTube.
So it's substantially this. For instance, this is a way to say, “Yes, remember when we told you, for the company, to use this sameAs in Organization, or in Author or Person for a personal social media profile, and so on?” This is where all the things come together, I think. And I think that we are going in that direction.
Disambiguation, Identifiers, and Identity on the Web
Gianluca Fiorelli: However, Mark, don't you think that the biggest misconception, for instance, when it comes to structured data and AI—using AI as a synonym for LLM—is that, also because of the very bad training data LLMs have, there is this myth that structured data is used by LLMs as such in order to better understand content, when maybe it is not really so? It's something that happens before the code is fetched.
Mark van Berkel: There are maybe a couple of things I want to unpack and try to tie back to what was said earlier.
The knowledge graph, or even schema markup. Schema markup describes a concept or subject on a page. It has a property, and it has an object or a label on the end of it. It could be a product, it could be its sale price, or it could be lots of different things.
For many years, schema markup was kind of a reproduction of your web content as JSON-LD, Microdata, or RDFa. It was just a reproduction of that. But as we're going forward more recently, there is a real opportunity to use this in new ways.
The @id and sameAs properties are two really important things. The @id helps you create identity. It is literally an identifier for that subject. The sameAs property connects it to Instagram, YouTube, or other social media profiles. That was another influential one. But the @id didn't necessarily have the market demand for it.
But for companies that are looking to create a coherent set of facts for AI, the @ids are really important to start connecting across pages and say, “We're talking about the same thing over and over on these many pages.”
That same structure of connecting these subjects across many pages can also be done across websites. If you’re saying, “This is an entity that is described on this landing page on my site. It's referred to in a few blogs, but it's also the same as this other concept in Wikidata or Google's Knowledge Graph.” Creating those links is really important for creating a network and having one data set, rather than many siloed data islands.
The need to have one set allows you to start governing it, because how else would you update an entity that appears on 20 pages if you don't know that it's on those 20 pages? By having that explicit link to say, “It's connected to this ID,” you can update it in one spot. So there are savings in the total cost of ownership when maintaining and keeping a data set.
For AI, this is really important because, if you update one part of your site but forget that some other facts spread around your website are outdated, you can have problems. For example, one of our customers was asking, “Who's the CEO of the company?” They didn't post the fact that they had replaced the CEO, and it was in a blog post. So the answer came up from AI saying, “It's this CEO,” but that's incorrect because they had left that role.
Different things like this are really important to collect and to have in a governed data layer so that you're telling one brand story.
And just for the last little point I was thinking about: when you talk about taxonomy, to me, it's just another way of expressing information about the same entities. One entity can be in a concept scheme or a taxonomy, but it also has schema data. You're really talking about one thing.
So the more you can do to enrich that one entity, its information taxonomy, its identity, and its properties on the page, the more you create a really rich and consistent foundation for AI, but also for many other use cases in marketing.
Martha van Berkel: Gianluca, you're on mute.
Gianluca Fiorelli: So that's why I was saying that physically creating the knowledge graph, so also seeing the knowledge graph physically, even if sometimes it's very hard to see these maps of dots and axes. It's a way to discover unforeseen connections between things.
This is where I think the most amazing output of a knowledge graph is from a marketing point of view.
Discover what’s driving visibility.
Advanced Web Ranking helps you explore the impact of structured data, semantic SEO, and AI optimization across both traditional search and AI search experiences. Try AWR free.
Data Governance at Enterprise Scale
Gianluca Fiorelli: You were introducing the topic of governance, and I think this is a question for both of you. Let's be very unsophisticated: if you had to give governance a percentage compared with other factors and other things related to structured data and semantics, what percentage would governance represent in standard day-to-day work on semantic and structured data?
Mark van Berkel: I'm just going to spitball or guess, but I think today it's probably going to be… If you think of a content audit, how often are you doing a content audit or trying to get a sense of how well you're representing topics? It could be pretty often, but it's sort of a one-dimensional thing.
You'll do a report, but are you actually modifying, updating, and assigning ownership over some of these subjects and domains? So I think the investment in it is probably quite low right now, in terms of how well people are doing it. I don't think there are probably great tools to do that at the moment.
And I think it should be a bigger part of the SEO role in the future, helping to support that foundation of a consistent data layer for AI. It should be a bigger part of the work.
There are proxies for this. Governing data at enterprise scale is not a new thing. There are companies that have been working in semantic data governance and enterprise data governance. There's been a conference for this for as long as I can remember, 10 or 12 years. The one I was thinking of is in Washington.
It's a practice and a domain that has sat in IT for a long time, and they would say, “We need to do a good job of this.”
There are data.world and Informatica solving this. If you think about an IT budget, I don't know what percentage it is. Maybe it's 10% or 15% for things like master data management and enterprise data governance. It is absolutely a part of bigger companies today, but I think it is a practice that marketing is quite new at.
Martha van Berkel: I would even say, putting my SEO hat on again, that historically we would have talked about how you avoid schema drift. When you're implementing schema markup, the whole reason we built Schema Hub is because, to do this for enterprises, you have to build in the governance.
We do validations to make sure it's proper JSON-LD. We use open standards when we're storing the data in the ontology. There are things you have to do to make sure this is live and representative of the data across thousands and millions of pages. You have to build automations and rules to do that.
Historically, an SEO might be like, “Oh my goodness.” I talked to a customer this week who said, “We're using the Yoast plugin and then doing some manual schema on it.” That's not going to work in this new world because you need it to be automated and to live and breathe with the content.
Then I think the next step, where Mark is kind of going, is to move up the maturity scale. Are the business concepts or entities in my business well-defined? Are they duplicated?
I used to joke, using a very simple search example, that someone might be looking for the page that talks about a product, but you have four pages that talk about that product. There isn't one entity home that it can be confident is the one place, or you're not using things like @id to connect the dots between them. AI is just going to get confused.
So I think schema drift is the last decade's problem. That now gets exacerbated because we have AI and emerging use cases that we don't even know about hitting your site and asking those questions today.
I do think automation, standards, scalability, and being able to do maintenance at scale in an automated way are important. Again, these are things that we've thought about at Schema app because of the nature of our clientele.
Emerging Standards: WebMCP, MCP, NLWeb and Agent Discovery
Gianluca Fiorelli: Yes, and I'm also thinking about all the new things that are popping up. Many are not yet standards, as you were saying. They are proposed standards, or they are not even defined as standards because they are proprietary ideas.
For instance, I'm thinking of the open knowledge—okay, I never remember. I always confuse it with the Google one.
Martha van Berkel: We have an opinion on that.
Gianluca Fiorelli: I can imagine. Because of your work, especially with enterprise companies, clients may ask you, “We already started doing all these things, or we have already been doing them with you for many years. Now do we also have to start thinking about WebMCP, UCP, and ACP?" And all these new acronyms popping up? OKF, I think, is there. Or it's OKR. I never remember the correct one.
First of all, how do you make your clients breathe and think about how to prioritize all these things? Second, because they are not yet standards, how can you make them coexist with standard structured data?
Then, eventually, this is another question, perhaps for later, and perhaps Mark can also answer it. If you are working with knowledge graphs, you are using triples or RDF. How can all these different vocabularies be made to work in synchronicity without breaking the stuff?
Martha van Berkel: Okay. I'll talk about the alphabet soup and how we talk to leaders and SEOs about that. It was funny when we were starting the conversation today, Gianluca, that this is what I was thinking about. There's actually a bigger job to be done around education. Once we understand it, we have to make it make sense.
A lot of what we've been talking about is reframing things around how we bring clarity about the brand. So this is sort of like—are we describing what a page is and how it's connected to other things within the brand? This is schema markup 101, which I've been talking about for a decade.
Then we get to understanding, which is the next maturity level. We're not just connecting pages. Now we're being very sophisticated about defining entities, how we want them defined, how we want them disambiguated—perhaps that’s with the sameAs—and how they're connected within the schema markup in context.
This is really about how we add broader context. It involves not just automating the process but also managing and optimizing it, and really thinking about entity management. I loved you talking about taxonomy because this then becomes your data about your business that you're putting into the machine.
There's a sophistication level around entity management. I think about that as the health of your semantic data layer. Then there are these new use cases, and I like to think of them as business use cases rather than standards.
I got great feedback from an executive yesterday. They said, “Martha, it's great that agent discovery resources are coming out, along with WebMCP and MCP. But how does this impact my business?” Someone said to me, “Martha, I have four vendors doing onsite search.” I said, “That's where NLWeb can use your semantic data layer to help you have one conversational, natural-language search on your website.” Someone else might say, “The CEO is asking what we're doing for AI readiness.” NLWeb also plays a role there.
You also want to make sure you're thinking about entity management and actions. What are the actions on your website? Let's start planning how we're going to build the doorway for agents, grounded in data, and then think about actions.
Another use case might be, “We're building our internal chatbot, and I'm expected to bring all the marketing data.” Great. Let's get your internal team to use MCP because MCP goes into your knowledge graph.
All these things are very confusing as they're emerging, especially since we don't yet know how some of them are going to interact and work together. But we're working on those use cases and explaining them and making sure people know that they are grounded in that semantic data layer and knowledge graph and that they have this strategic marketing asset.
I think that's something SEOs can use to sell the importance of doing this. You're setting up the marketing team to have that strategic asset because the applications will change. Those use cases are going to keep evolving and changing, but the data is at the core of it.
Gianluca Fiorelli: Indeed. Now I'm curious: what is your take on the Open Knowledge Format?
Martha van Berkel: I see it as interesting that it came from the Google Cloud people.
Gianluca Fiorelli: Yes. This is something that I also highlighted in the post I wrote about it. This was not announced by the developers team or even the search team. It's Google Cloud.
Mark van Berkel: Yes.
Martha van Berkel: We often look at who's authoring it. Agent Discovery Resource is much more interesting because of who's behind it, and it's more likely to live on the web of the world that we live in versus the Open Graph pieces that look more like a tool in the toolkit for internal knowledge structuring. It's not that different from when people talk about semantic HTML. When I read it, I thought, “This is in the same vein as that.”
So I see it as more relevant for internal optimization than for agentic AI readiness. I see other standards as more relevant there. Mark, what do you think?
Mark van Berkel: Essentially the same. The team and the description of that table, or series of tables of data… I speak in graphs, and these are all tables of information, which is how relational databases are structured and where BI people live. They work in tabular formats, and this is the team that came out with it, so that makes total sense for them.
I think it also indicates another signal around making data readily available on the web. R.V. Guha, one of the founders of the Schema.org project who was at Google for a long time, proposed some ideas around 2022 about how we could get datasets onto the web and bring that data to the forefront.
There was some back and forth in the Schema.org issue about ways in which you could express data on the web. Open Knowledge Formats is another idea for that, but it flattens the structure.
You could put Schema.org data into it without a problem, but it would just be another expression or translation of what you would already be doing.
There are a few projects and people interested in re-expressing the data behind a website in a readily accessible format. I'm not sure publishing the raw data is necessarily the right means to the end. It's one way to do it. But wouldn't it be better to have a hosted service that helps propose, “Here are the kinds of questions and resources we'd like to make available to you”?
Think of MCP. MCP is a much better way of packaging the data into the relevant user stories that your customers might be interested in, rather than providing a data dump. A data dump is one thing, but until the big technology companies say, “We need this,” it's not a sure bet to put something like that on the web.
Discover what’s driving visibility.
Advanced Web Ranking helps you explore the impact of structured data, semantic SEO, and AI optimization across both traditional search and AI search experiences. Try AWR free.
How AI Systems Actually Use Structured Data
Martha van Berkel: It was interesting. Gianluca, I liked your question about how we get the data to the LLMs. When I was educating myself and making sure I understood the standards, I was having this rich conversation with ChatGPT. It said, “You can't submit the information to me. We crawl the site. We're still using search indexes.” I was laughing because I thought, “Why can't I just have an API and submit my graph to you? That would be so much easier.”
Then I had a conversation with Navah Hopkins, one of the product marketers at Microsoft. I'm going to quote their blog: “As of today, Microsoft grounding powers nearly every major AI assistant in the market.”
I find this fascinating because, if we step back macro, Google owns the market today in traditional search. There's no reason why ChatGPT, Perplexity, and others would partner with Google because Google is the competitor whose market share they're trying to take.
Microsoft has a smaller share, but it has the index because it has already done this at scale. When I interviewed the person from Microsoft who helped start Schema.org, around 2016 or 2017, it was clear that Microsoft had a sophisticated knowledge graph. They just don't talk about it externally as much as Google does.
So now we were exploring whether they are using structured data for training data. Well, Microsoft and Google have been very explicit that they’re using structured data.
Gianluca Fiorelli: Yes, indeed.
Martha van Berkel: I think about the profitability of ChatGPT and all these others that are really trying to scale. They’re not making enough money to scale, so they’re figuring that out. It makes total sense that they would borrow the index from Microsoft to do that. Google and Microsoft are using everyone’s structured data because they built it, have it, and it’s cheaper.
We were listening to Ryan Levering at Google Search Live, and he said, “It’s more efficient and effective to take the structured data you’re giving us than for us to figure out what’s on your website. Why wouldn’t we use that?” It’s cheaper for them to do it.
There are lots of arguments, and I know there are lots of debates online because people have been testing and saying, “When I train it on my local machine or with my ChatGPT instance, they’re not using structured data.”
But I would say those two pieces of evidence, along with what we see with our customers. When we take people through that enrichment piece—building clarity by describing what something is and then building rich entities and semantics into it—we see AI visibility increase across different platforms. We see hallucinations go away.
So I would say those two pieces of evidence mean that the semantic data layer we’ve historically called schema markup is still playing a role. Financially, it makes more sense, and we’re also seeing the results.
Gianluca Fiorelli: Indeed. I think that ChatGPT was also using Google through the SERP API. Now that SERP API has won the trial against Google, who is to say that ChatGPT is not going to use both Microsoft and Google through SERP API?
I think this is also based on a very old misunderstanding related to structured data and schema. They are not used in the true indexing part, but they are used in parsing, before indexing. They are substantially used to structure how Google, Bing, or any other search engine using schema, such as Yahoo, Yandex, and others, is going to organize the index.
They help determine the placement in the index, depending on the understanding, which has also been improved thanks to schema. I think people tend to forget where schema is working.
That’s why, when people say that ChatGPT, Claude, or another LLM is fetching a URL live, it is not reading the schema. It is stripping everything and transforming it into Markdown. It recognizes things, so it can tell you that there is a title tag or a meta description. It can rebuild the schema only if you explicitly ask it to rebuild the schema present on a page it is fetching.
But that is costly. When they are searching and fetching live, that is not where schema matters. That’s why schema, as Google also says, is not some sort of ranking factor in AI search.
Martha van Berkel: That’s right. Mark, you always have a story about how machines also know this. It’s a known vocabulary. Do you want to talk a little about that?
Mark van Berkel: Right. The benefit of using Schema.org is not just that you provide the data. It’s also in the model. The model is something that is well understood by LLMs because it’s all over the web. The specification is very public, and there is a lot of conversation about Schema.org.
The classes and properties themselves are part of the training set of LLMs. When you provide information, yes, there will be hallucinations. They’re obviously better now than the early versions of ChatGPT, but the vocabulary is pretty well known in the foundation models or in the training set. You’re almost getting the semantics for free with LLMs if you use the vocabulary.
As an aside, if you have your own proprietary data model and describe objects, classes, and properties using your own terminology, and then expect the LLM to understand it, you need to give it a lot of context and train it to understand those classes.
Using Schema.org, you get that data model almost out of the box. There is still some fine-tuning, but the semantics of those terms and the vocabulary are a benefit of using it.
Rather than trying to introduce LLMs to a new proprietary format, you can use Schema.org and extend it. You can use subclassing and create your own versions that largely use the Schema.org data model as a way of interacting with AI and bootstrapping the effort.
This is the main point: it has been part of the internet infrastructure on which all LLMs have been trained. So let’s use it.
Optimizing for Agents: ARIA, SKOS, and Actionability
Gianluca Fiorelli: Before, Martha, I noticed a little irony in your tone when you were talking about semantic HTML. I don't want to talk about semantic HTML, even if the ABC of classic standard on-page SEO is to use the headings.
Martha van Berkel: It's good best practice, right?
Gianluca Fiorelli: Yes. But I want to talk about another type of semantic labels, let's call them that, related to—I don't know if I'm going to pronounce this soup of letters correctly in English. In Italian, it would be pronounced “Aria,” so A-R-I-A. It's the vocabulary and all the rules related to usability, around how to label the code of a page correctly to make it usable and also help the machine understand it better.
One classic mistake, for instance, is having divs without labeling the nature of the div. This is one of the three things. Another one is the screenshot of a page. Google is also suggesting a third one, which I don't remember at this moment.
ARIA is one of the three methodologies that Google—not Google Search, but Google Developers and Google Chrome—is suggesting for agentic optimization. If we want to optimize a website for agents, we should make sure we do these things.
What kinds of semantic labels do you think are important but that SEOs, web designers, and others have not yet correctly considered?
Martha van Berkel: Maybe I'll frame it at a high level, and then Mark can jump in as well.
At a high level, I call it about labels or metadata because I also think about accessibility metadata and other pieces that are there to help understand by different users.
The agent and AI are new users, right? I often talk about it as a new marketing channel. Now we're not just thinking about humans or crawlers. We're thinking about AI as something that can take action.
Where my mind goes, and where I've been spending time thinking because I see it coming next, is how we look at the standards for actions. We know about the doorway for agents. NLWeb is a great open standard that we quite like and that we're building on for receiving agent requests.
How do we then think about actioning those requests so that there is an understanding of the content and knowledge of what page or entity home to go to in order to understand the area where the agent might make a decision? How do we then help agents take action?
I would say that's the next piece of the puzzle that needs to be solved. It's also where Mark and his team at Schema App are delving deeply and where we'll be early adopters on our own website, making sure that's happening.
For SEO teams, do you even know what actions you want to convert? It's great if AI understands, but we're still in the business of conversions. We're still trying to help organizations connect with their customers and help them take action to drive business.
That's the evolving standards space I'm most interested in. I'm eager to see what our team builds, what becomes standardized, what we action, and how we connect the dots between the alphabet soup of emerging standards.
But again, the intent is: how do we help convert? I think agent conversion is what Google, Shopify, and ChatGPT are talking about. They're starting in e-commerce, but what about everyone who isn't in e-commerce? That's where my mind goes when you ask that question.
Gianluca Fiorelli: Yes, I think so. Mark, do you have an opinion about this?
Mark van Berkel: I'm going in a slightly different direction on the technology. Two substitutes, almost, based on my knowledge of ARIA, are even in Schema.org, you have DefinedTerm as one of the classes. DefinedTerm helps you create specific representations with @ids and other kinds of labels.
What we primarily use to solve a broader set of problems and to represent terms and interrelated series of terms is a vocabulary called SKOS. SKOS has been around for quite a long time. It has a series of concepts, and the concepts have broader and narrower ways of expressing that one term is broader than another or narrower than another. So you can create taxonomies. Gianluca, you're on mute. I think you're jumping in.
Gianluca Fiorelli: Simple Knowledge Organization System. SKOS.
Mark van Berkel: It has a really nice, simple vocabulary. You can create synonyms and labels and define how things are organized in relation to one another. That's where we spend more of our time rather than on ARIA. I quite like SKOS. It was born out of the same semantic web community. It's very flexible and allows for more than just site synonyms and such.
Discover what’s driving visibility.
Advanced Web Ranking helps you explore the impact of structured data, semantic SEO, and AI optimization across both traditional search and AI search experiences. Try AWR free.
Balancing Creative Tone with Machine Clarity
Gianluca Fiorelli: I see. I have one last question. It's less technical and more conceptual. We've talked about clients, how to sell semantics to CMOs, and how to help SEOs and developers better understand the real purpose and nature of structured data and schema. But a website is also co-created by the people who create its content.
There has been a lot of discussion, even before AI, about the importance of natural language. It was first introduced because of Google Assistant and voice search, which never really exploded. But conversational search is essentially what LLMs are about.
On one side, natural language is also used in marketing. This includes rhetorical forms, playing with words, and creating ambiguity. These things may be needed for commercial and marketing purposes, or simply because they reflect the brand's style and tone of voice.
How do we reconcile this natural-language creativity with the clarity needed by machines?
Martha van Berkel: I think of this just like we started as SEOs thinking about sameAs as a way to clarify an organization with the social media piece. But entity disambiguation can use any property to bring clarity.
Sometimes that means using external definitions, whether that's Google's Knowledge Graph, Wikipedia, or Wikidata. That's external entity linking for disambiguation. If you're using a term that isn't clear, you can clarify it. You can also use things like SKOS to define the relationships between definitions, which is the more advanced approach.
When we talk about entity management, it's more than automating and identifying entities on a page. It's about being specific about what the brand means and making that meaning more accessible and clearly defined.
I also think of it as bringing clarity to relationships. Again, this goes back to the fundamentals. Ten years ago, we talked about the basis of the Schema.org vocabulary: you bring context and clarity when you define relationships between things.
This also comes back to your taxonomy, Gianluca. You can use relationships to define the taxonomy and the meaning of things. When you ask that question, my mind goes to these fundamentals. People who have practiced semantic SEO and used schema markup as it was designed already have a way to do this.
Then you can layer natural language, vector databases, and other things on top to understand those pieces. These things go together. Mark can talk a little more about how he sees them connecting and helping to answer that natural-language problem.
Mark van Berkel: You covered it well. I would just add that there's an expression: ambiguity is an LLM's kryptonite. If you have ambiguous information or represent the same thing differently in different places, what is the LLM going to do? It's going to make a decision about you, your brand, and the commitments you're making to the things you want to be known for.
Natural language is the top layer, but consistent data and entities are going to beat your clever prompts. You want to remove hallucinations so that whatever is underneath the natural language is an enduring asset. Those LLMs are going to come and go, but your data is forever.
Continue the conversation about semantic search with Jarno van Driel
Jarno takes a look at where structured data creates real value and where the industry may be expecting too much from it. The episode explores:
Schema.org beyond rich results,
the limits of structured data for LLM visibility,
internal knowledge graphs,
and why emerging agentic technologies may need more time before most businesses invest in them.
Working Together as Co-Founders and a Couple
Gianluca Fiorelli: Okay. We've reached almost one hour. We're already at 58 minutes, so let's stop here and promise to continue because I think we could go on for another hour.
But let me finish with a question about the two of you.
You're colleagues and co-founders who share a passion for the things you're working on. How are you able to switch between Martha and Mark, the CEO and CTO, working on everything related to schema, structured data, and clients, and Martha and Mark as a couple, when it's just the two of you? How much do you have to optimize that switch?
Martha van Berkel: It's funny because we just talked about clarity. I think it comes back to clarity. We're very clear about our roles at work and who makes which decisions, as well as our leadership team. That allows us to understand where we need to stay in our lanes.
We also have some rules. If you want to talk about work at home, you have to ask permission. If it's 10:00 at night and I'm ruminating on something, but Mark is ready to watch Game of Thrones, I have to ask permission. He can say, “Actually, no. We'll talk tomorrow.”
Our kids are also hilarious. Sometimes we'll be passionately debating something, and they'll say, “Stop talking about work.” We respect that request and stop talking about it. I think part of it is having that real clarity. It's very special that we get to work together and have so much fun. I don't know if it's for everybody, but it really works for us. Mark, what would you add?
Mark van Berkel: Yes, that's pretty much it. We like each other, so it's good. We do have conversations. We don't quite butt heads, but we have to have some hard conversations from time to time. I think that's something you want in a marriage and also in a co-founder relationship.
You need to have an intimate relationship with your co-founders or your life partner. Being open to having those hard conversations is the foundation of that trust, and then everything can flow from there.
Martha van Berkel: Yes. We also have clarity about where we're going. The whole leadership team and company discuss as a team and commit to where we're going. So much of it is also trust. You have to invest in the trust buckets, and some of that comes from getting to know people on a different, more personal level. We had that out of the gate. But I recently read something about how, as we age, we change too.
Gianluca Fiorelli: Indeed.
Martha van Berkel: So it's about continuing to be curious about who Mark is as he gets to be 50, or how he is different now that he has evolved through this time in his career. So we keep being curious too.
Gianluca Fiorelli: Indeed. Well, Mark and Martha van Berkel, it was a real pleasure to have you here on The Search Session. Here, we have talked many times with many people about semantics, such as Andrea Volpini, Jarno van Driel, and many others.
Martha van Berkel: Our friends, yes.
Gianluca Fiorelli: Yes. So maybe I'm starting to think about doing something bigger. Now that I'm refreshed on how to have three guests, maybe we can have five or six and do a special multi-panel episode. Thank you, Martha. Thank you, Mark.
Martha van Berkel: Thank you.
Mark van Berkel: It was a pleasure. Thank you for having us.
Gianluca Fiorelli: And thank you to all of you for being our guests too. Remember to subscribe to the channel and ring the bell to be notified whenever a new episode is published. Ciao.
Podcast Host
Gianluca Fiorelli
With almost 20 years of experience in web marketing, Gianluca Fiorelli is a Strategic and International SEO Consultant who helps businesses improve their visibility and performance on organic search. Gianluca collaborated with clients from various industries and regions, such as Glassdoor, Idealista, Rastreator.com, Outsystems, Chess.com, SIXT Ride, Vegetables by Bayer, Visit California, Gamepix, James Edition and many others.
A very active member of the SEO community, Gianluca daily shares his insights and best practices on SEO, content, Search marketing strategy and the evolution of Search on social media channels such as X, Bluesky and LinkedIn and through the blog on his website: IloveSEO.net.





