HomeReddit GrowthMonitorAI SnapshotBlog
HomeReddit GrowthMonitorAI SnapshotBlog
Start Growing

Schema Markup for AI Search: How Structured Data Gets You Cited by ChatGPT, Perplexity, and Google AI

Sep 16, 2026 · 12 min read

What structured data does for AI answers, what the engines actually document about it, and which schema types still earn their place in 2026. Written for the technical marketer who owns the implementation ticket.

Schema Markup for AI Search: How Structured Data Gets You Cited by ChatGPT, Perplexity, and Google AI
byHanna Forras

Table of contents

  1. Key Takeaways
  2. What Schema Markup Is — and Why AI Engines Depend on It More Than Google Ever Did
  3. Structured Data for SEO vs. Structured Data for LLMs: What Changes in 2026
  4. The Schema Types That Actually Drive AI Citations (Organization, FAQPage, HowTo, Article)
  5. Step-by-Step: How to Add Schema Markup to Your Site
  6. Schema Markup Guide: Common Mistakes That Get Your Structured Data Ignored
  7. How to Validate and Test Your Schema for AI Crawlers
  8. Schema Markup + Entity SEO: Building a Machine-Readable Knowledge Graph
  9. Where Schema Fits in a Klarivo Program
  10. Frequently Asked Questions

Schema markup tells machines what the things on your page are. Google recommends it for rich results, but its guide to generative AI features says AI Overviews and AI Mode do not require it.

The numbers are real. They measure something else. Google’s introduction to structured data reports 82% higher click-through for Nestlé pages shown as rich results and 25% higher click-through for marked-up Rotten Tomatoes pages. Neither figure measures an AI citation.

Schema belongs in answer engine optimization. Ship it correctly. Just don’t promise it will get you cited.

Key Takeaways

  • Google names schema as a myth, not a lever: its May 2026 generative AI guide puts "Overfocusing on structured data" in the list of things you can ignore for Google Search.
  • No engine documents schema as a citation factor: not Google, not OpenAI, not Anthropic, not Perplexity. Every "schema gets you cited" claim in this category is inference presented as documentation.
  • Two types every guide still recommends are retired at Google: the FAQ rich result stopped showing on 7 May 2026 and HowTo went in 2023. FAQPage vocabulary stays valid and still labels question and answer pairs, which is why it stays on the shortlist here.
  • The case-study numbers measure classic search: Nestlé's 82% and Rotten Tomatoes' 25% are rich-result click-through figures. Reusing them as AI-citation evidence is a false claim.
  • Organization markup is the one with durable value: it has no required properties, it carries sameAs, and it states your brand's facts in the one place a machine reads without inference.

What Schema Markup Is — and Why AI Engines Depend on It More Than Google Ever Did

Schema markup is a block of schema.org vocabulary embedded in a page, declaring the type of each thing on it. A page of prose is ambiguous by design: the string "$49" could be a price, a refund, or a competitor's price. Markup removes the guess: it labels that string as a price belonging to an Offer belonging to a Product.

Google recommends JSON-LD, a block of JavaScript Object Notation in a <script> tag. The documentation supports Microdata and RDFa too, and states all three are equally fine when implemented correctly. The practical argument for JSON-LD: it sits apart from the visible HTML, so a template change does not silently break it.

It helps to separate this from the two things it gets confused with. Access is a different layer: robots.txt decides which agents may fetch the page at all, and markup on a blocked page is invisible. Discovery is a third layer: the markdown index for AI crawlers proposes a file that says what to read. Schema says what the things you read actually are. Google's own guide names llms.txt among the files it does not use.

On the question in this heading, the honest answer is narrower than the phrasing. Google documents no dependency at all. Its AI features documentation sets one bar: the page must be indexed and eligible for a snippet. Then it adds, "There are no additional technical requirements." A second line says "specific optimization isn't required for AI Overviews and AI Mode", while keeping structured data on the list of worthwhile fundamentals. What decides inclusion is what decides it in classic search, which is what makes a passage quotable.

Structured Data for SEO vs. Structured Data for LLMs: What Changes in 2026

Structured data for SEO has a documented, testable payoff: rich result eligibility. You add the required properties for a feature. Google validates them, and the page becomes eligible for a different presentation. The mechanism is contractual: Google publishes the required properties per feature and honors them.

The LLM side has no contract. No engine publishes a schema requirement, a supported type list, or a validator. What markup does there is quieter. It states your facts in one unambiguous place, so a reading system has less room to infer wrongly.

Structured data for SEOStructured data read by AI systems
Documented requirementYes, per rich result typeNo
Published validatorYes, Rich Results TestNo
Failure modeIneligible for the rich resultNone documented
What it buys youA different SERP presentationFact clarity, no promised outcome
Worth shippingYesYes, for accurate page description and documented Search benefits. No AI citation is promised.

Three things changed in 2026. All three narrow the claim rather than widen it. Google published a generative AI guide in May that names AEO and GEO directly and lists the tactics you can ignore. Google retired the FAQ rich result on 7 May 2026 and removed its documentation the next month. Search Console added a generative AI performance report. The measurement question moved from inference to a report you can open.

One wrinkle is worth flagging, because two Google pages sit slightly apart. The AI features page says there are no additional technical requirements. The newer guide adds a second condition: the site must be included in Search generative AI features in Search Console. Read the newer page as the operative one. Check that setting before you conclude anything from an absence.

The Schema Types That Actually Drive AI Citations (Organization, FAQPage, HowTo, Article)

No engine documents a schema type as a citation factor. So treat this as a list of what each type declares, not a ranked list of levers. The useful question is not which type wins a citation. It is which types state the facts a machine would otherwise have to infer from prose.

Six useful types cover the pages most B2B SaaS teams will mark up first.

TypeWhat it declaresGoogle rich result status, 16 Sep 2026
OrganizationWho the company is, and which profiles are the same entityLive. Can help Google choose a logo for Search or a knowledge panel
ArticleAuthorship, publication date, and revision dateLive
BreadcrumbListWhere the page sits in the site hierarchyLive
SoftwareApplicationWhat the product is, its category and its priceLive, with the condition below
ReviewWho rated what, and on what scaleLive, with a hard restriction below
FAQPageWhich text is a question and which text is its answerRetired at Google, vocabulary still valid

Organization is the one to ship first. Google's Organization documentation states it has no required properties, so add whatever is relevant. Place it once, on your home page or on a single page describing the organization.

Article and BreadcrumbList are the low-effort pair. Article markup can use Article, NewsArticle, or BlogPosting. Google recommends author, headline, publication and modification dates, and image, but requires none of them for this type. BreadcrumbList needs only item, name, and position per entry. Google does not require an entry for your domain or for the page itself.

SoftwareApplication carries a catch most B2B teams walk into. Google requires name, offers.price, and either aggregateRating or review. WebApplication is the supported subtype for browser-based products.

That catch makes the review rules worth reading twice. Google does not show review stars for self-serving reviews of an Organization or LocalBusiness, even when they appear in a third-party widget. SoftwareApplication has a different rule: Google requires an aggregateRating or review, alongside the app name and price, for its rich result. A G2 or Trustpilot badge on your own page does not automatically qualify. Check the rules for the specific result you want.

FAQPage stays on this list despite the retirement. Google restricted the rich result to government and health sites in 2023, stopped showing it on 7 May 2026, and removed the documentation in June. The vocabulary is still valid schema.org. It still does the job that matters here: it labels which text is a question and which text answers it. HowTo is the genuine casualty. It was cut to desktop only on 8 August 2023. The documentation came down on 14 September 2023, once the result stopped showing at all.

Mapping the rest to page archetypes is the part generic guides skip. Pricing page: Product with Offer, only if the prices are on the page. Product page: SoftwareApplication or WebApplication. About page: Organization with sameAs. Docs article and blog post: Article or BlogPosting. Comparison page: Article, not Product. The page is analysis, not a listing. Schema describes a page. It is no substitute for coverage structured as clusters.

Step-by-Step: How to Add Schema Markup to Your Site

Pick the Type Before You Write Any JSON

Start from what the page is, not from what you want it to rank for. Google's specificity guideline is explicit: use the most specific applicable type schema.org defines. Mark up a comparison article as a Product and it fails the relevance guideline. Markup has to be a true representation of the page.

Write the JSON-LD

Here is a working Organization block for a B2B SaaS about page. Every value in it is also visible on the page.

html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Organization",
  "name": "Example Analytics",
  "url": "https://www.example.com/",
  "logo": "https://www.example.com/assets/logo.png",
  "description": "Product analytics for B2B software teams.",
  "foundingDate": "2019-03-01",
  "sameAs": [
    "https://www.crunchbase.com/organization/example-analytics",
    "https://github.com/example-analytics",
    "https://en.wikipedia.org/wiki/Example_Analytics"
  ]
}
</script>

And an Article block for a docs page or blog post:

html
<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Configuring Single Sign-On",
  "image": "https://www.example.com/assets/sso-guide.png",
  "datePublished": "2026-02-11",
  "dateModified": "2026-08-30",
  "author": {
    "@type": "Person",
    "name": "Dana Okoro"
  },
  "publisher": {
    "@type": "Organization",
    "name": "Example Analytics"
  }
}
</script>

Ship It Where the Page Can Be Crawled

Put the markup on the page it describes. Google's location guidance adds one step teams miss. Where duplicate pages exist, put the same markup on every duplicate, not on the canonical alone. Check access before you celebrate. The guidelines prohibit blocking marked-up pages to Googlebot via robots.txt or noindex. Markup on a blocked page is never read.

Handle the Templating Problem

On a site with many pages, schema is usually generated rather than typed. One common failure follows: a template ships the field, the CMS leaves it empty, and the whole template emits broken blocks. Generate from the same data source that renders the visible text. Then the two cannot drift apart.

Schema Markup Guide: Common Mistakes That Get Your Structured Data Ignored

The objection that opens most conversations is a fair one. "We already have schema from a plugin and nothing changed." Usually the plugin is right and incomplete. It emits valid markup that misses the required properties for any feature. The page never becomes eligible for one. That outcome is the most common one, and a syntax checker cannot tell it apart from success.

The rest of this schema markup guide covers the failures Google names in its general structured data guidelines.

Marking Up What the Reader Cannot See

This is the one with teeth. The guidelines say not to mark up content that readers cannot see. The example is direct: if the JSON-LD describes a performer, the HTML body must describe that same performer. Google's AI features page repeats it as a best practice, asking that structured data match the visible text.

Missing Required Properties

Items missing required properties are not eligible for rich results. Google's completeness guidance goes further. It prefers fewer complete properties over many badly formed ones. Adding every optional field you can find is not the win it looks like.

Marking Up the Wrong Thing

Google's relevance examples are blunt: a woodworking site labeling instructions as recipes, a streaming site labeling broadcasts as local events. A B2B version exists and is common: a pricing page marked up with Product and Offer when no price appears on the page.

Treating a Manual Action as a Ranking Penalty

It is not one. Google states that a manual action costs the page rich result eligibility. It does not affect how the page ranks in web search. Check the Manual Actions report before you blame a traffic drop on markup.

How to Validate and Test Your Schema for AI Crawlers

**The validator to use is the Schema Markup Validator.** It checks your markup against schema.org vocabulary itself, across every type. That is the right test when you ship schema for machine-readable meaning rather than for a Google feature. Paste a live URL or the snippet itself, and it reports what it parsed and what it could not.

One caveat, stated plainly so nobody sells you the opposite. No AI-specific validator exists. No tool tests how ChatGPT, Claude, Perplexity, or Gemini read your markup. None of those engines publishes a schema specification to test against. The schema.org validator is the closest honest thing, because valid vocabulary is valid for any system that parses it.

Google's Rich Results Test answers a narrower question. It reports eligibility for Google rich results on the types it detects, and stays silent on every type outside that set. Run it second, and only for the types Google still has a feature for.

ToolWhat it validatesUse it for
Schema Markup Validatorschema.org vocabulary, every typeThe default check on all markup
Rich Results TestGoogle rich result eligibilityTypes with a live Google feature

After deployment, Search Console's rich result status reports catch what a pre-launch test cannot. Templating and serving issues appear later. Outside Google, the honest method is behavioral rather than technical. That means running the audit by hand and recording what the engines say about you.

Schema Markup + Entity SEO: Building a Machine-Readable Knowledge Graph

This is where structured data for LLMs earns its place. The framing has to be precise to stay true. Markup does not build a knowledge graph. It contributes one unambiguous statement of your facts, in a format that needs no inference, at a URL you control.

sameAs is the property that does this work. It takes a list of URLs that refer to the same entity. Point your Organization at your Crunchbase page, your GitHub organization, and your Wikipedia entry where one exists. Google names the property explicitly, saying it can make general use of sameAs beyond the documented feature set.

Two practical rules keep it useful. Only list profiles you control or can verify. A wrong sameAs asserts an identity that is not yours. Keep the list short and current: a dead profile is a contradiction you published yourself. Reconciling what your profiles, listings, and site say about you is entity authority. That is a bigger job than markup. Schema declares the answer. It is not the process of arriving at it.

Where Schema Fits in a Klarivo Program

Schema sits in the technical baseline, alongside crawler access and page structure. It is worth doing correctly. This article has also spent most of its length on why nobody has documented it as a citation lever. Klarivo does not claim to change that. What Klarivo can do is tell you whether any of it moved the outcome you care about.

Klarivo Monitor runs targeted queries across five answer engines: ChatGPT, Claude, Perplexity, Grok, and Gemini. It analyzes how your brand appears in the answers they generate. The surface that matters here is "Top 10 Citation Sources", which reports the domains the engines drew on about your category. That is measurement at domain level, and it is the honest unit. It tells you which sources an engine trusts. It does not tell you which line of your markup it parsed.

Separate the two jobs, because they fail differently. Ship schema because it is correct. It earns rich result eligibility and keeps your facts consistent where machines read them. Measure citations because that is the outcome. Markup is one of a dozen inputs you would otherwise be guessing about. To see which sources the engines cite in your category before you decide what to build, book a Klarivo discovery call.

Frequently Asked Questions

Does ChatGPT read schema markup?

OpenAI publishes nothing saying that ChatGPT reads or weights schema.org markup. Neither do Anthropic, Perplexity, or Google. Systems that browse the live web retrieve the page HTML, and the JSON-LD comes with it. What they do with it is undocumented. Nobody outside those companies can verify a claim either way.

If schema is not required, why does Google keep recommending it?

Because rich result eligibility is real and separate from AI features. Google's position is consistent across both pages. Structured data earns presentation options in classic Search. It is not a requirement for generative AI features. Those two statements are compatible.

Will removing FAQPage markup hurt my traffic?

The rich result it powered stopped rendering on 7 May 2026. The presentation it bought you is already gone. Leaving valid FAQPage markup in place costs nothing, and it keeps your question and answer pairs explicit. Removing it costs nothing either. Rank it below Organization and Article work.

How long after shipping schema should I expect to see something?

Google's own method is a before-and-after test. Pick stable pages with several months of Search Console history. Add the markup, then confirm Google found it with the URL Inspection tool. Compare performance over a few months. Anything faster than that is noise.

Do I need different markup for each AI engine?

No. No engine publishes an engine-specific schema specification, so there is nothing to differentiate against. One correct, complete JSON-LD implementation is the whole job. Any vendor selling per-engine schema variants is selling a distinction no documentation supports.

Start accelerating your Reddit presence

See how Klarivo can shift your visibility across AI, search and buyer communities.

Book a Demo

Table of contents

  1. Key Takeaways
  2. What Schema Markup Is — and Why AI Engines Depend on It More Than Google Ever Did
  3. Structured Data for SEO vs. Structured Data for LLMs: What Changes in 2026
  4. The Schema Types That Actually Drive AI Citations (Organization, FAQPage, HowTo, Article)
  5. Step-by-Step: How to Add Schema Markup to Your Site
  6. Schema Markup Guide: Common Mistakes That Get Your Structured Data Ignored
  7. How to Validate and Test Your Schema for AI Crawlers
  8. Schema Markup + Entity SEO: Building a Machine-Readable Knowledge Graph
  9. Where Schema Fits in a Klarivo Program
  10. Frequently Asked Questions

Related articles

Topical Authority: What a Model Judges You an Authority On
Technical Foundations

Topical Authority: What a Model Judges You an Authority On

What is topical authority? Learn why answering distinct questions thoroughly, not publishing more pages, is what earns credibility with AI search engines.

by Isaac Tarrab· Sep 4, 2026· 11 min read
Entity SEO: How Models Learn What Your Brand Is
Technical Foundations

Entity SEO: How Models Learn What Your Brand Is

How search engines and language models build a picture of your brand as an entity, where your entity data already lives, and how to fix what contradicts it.

by Hanna Forras· Sep 2, 2026· 10 min read
llms.txt: What It Is and Whether It Does Anything
Technical Foundations

llms.txt: What It Is and Whether It Does Anything

What the llms.txt specification proposes, what each major engine has actually said about reading it, how it differs from robots.txt, and what to realistically expect.

by Isaac Tarrab· Sep 2, 2026· 9 min read
How it worksReddit GrowthKlarivo MonitorBlogFAQsLinkedIn
Privacy PolicyTerms of ServiceCookie policy

© 2026 Klarivo. All rights reserved.