Structured Data for AI Search: Schema for GEO That Matches Reality

By Dfeelings · Updated

Short answer

Structured data for AI search is JSON-LD markup that describes a business, its services, people and pages using schema.org vocabulary, connected through stable @id references. For GEO, structured data helps make entity facts explicit and consistent. Google states that no special markup is needed to appear in AI Overviews or AI Mode, and that markup must match visible content, so schema is a clarity layer, not a shortcut, and no one can guarantee an AI system will cite a business.

Key facts

  • Google says no new machine-readable files, AI text files or markup are needed to appear in AI Overviews or AI Mode.
  • Google recommends JSON-LD for structured data where a site’s setup allows it.
  • Google’s guidelines say not to mark up content that is not visible to readers of the page.
  • Google deprecated the FAQ rich result and states it is no longer shown in Google Search results (2026).
  • Schema.org recommends IETF BCP 47 codes, such as ar or en, for the inLanguage property.

What does structured data do for AI search?

Structured data for AI search states the facts on a page in a standard, machine-readable format, so systems do not have to infer them from layout and prose. Google describes structured data as a standardized format for providing information about a page and classifying its content. For GEO, structured data makes entity facts explicit; it does not replace clear visible content.

Google’s documentation on AI features is direct: there are no additional requirements to appear in AI Overviews or AI Mode, and site owners do not need to create new machine-readable files, AI text files or markup to appear in them. The same page lists “making sure your structured data matches the visible text on the page” among general good practices. So structured data is useful hygiene for search in general, not a special AI ranking lever.

Other AI platforms do not publish how, or whether, they use schema.org markup when composing answers. The honest position is that structured data reduces ambiguity for any system that reads it, and costs little when it is generated from the same source as the visible page. Treat any claim that a particular schema type “unlocks” AI citations as unverified.

Why is JSON-LD the recommended format for schema markup?

JSON-LD is the recommended format for schema markup because it sits in a single script block separate from the page’s HTML, which makes it easier to generate, maintain and validate than inline formats. Google states that it generally recommends JSON-LD where a site’s setup allows it, while also supporting Microdata and RDFa. The W3C defines JSON-LD as a JSON-based format to serialize Linked Data.

Because JSON-LD is separate from presentation, a server-rendered site can build it from the same data that renders the visible page. That is the most reliable way to keep markup and content in sync: one source, two outputs. Hand-written markup pasted into a template drifts as soon as someone edits the copy and forgets the script.

The “Linked Data” part matters for GEO. JSON-LD can give each thing an identifier with @id and refer to it from elsewhere, turning a collection of separate snippets into a connected description of one business.

What is a connected @id graph, and why does it matter?

A connected @id graph is structured data in which every entity, such as the Organization, WebSite, WebPage, Service, Person and Article, has a stable identifier, and the entities reference each other by that identifier instead of being redefined on every page. The W3C JSON-LD specification describes @id as the way to identify a node and reference it elsewhere. A connected graph states relationships explicitly.

A typical service business graph looks like this. The Organization has @id https://example.com/#organization. The WebSite has @id https://example.com/#website and names the Organization as publisher. Each WebPage has its own @id, is part of the WebSite, and has a BreadcrumbList. A Service names the Organization as provider and lists areaServed. A Person worksFor the Organization. An Article names the Person as author and the Organization as publisher.

The benefit is consistency. When the organization’s details live in one node referenced everywhere, a page cannot accidentally describe it with an old phone number or a different name. When a plugin injects a second, conflicting Organization block, the conflict becomes visible during validation.

  • Organization → the business entity (name, alternateName, url, logo, sameAs, address).
  • WebSite → the site, published by the Organization.
  • WebPage → each page, part of the WebSite, with a breadcrumb.
  • Service → provided by the Organization, with areaServed.
  • Person → works for the Organization; authors Articles.
  • Article → written by a Person, published by the Organization.
  • BreadcrumbList → the page’s position in the site hierarchy.

Which schema types matter most for a service business?

The schema types that matter most for a service business are Organization (or a specific LocalBusiness subtype where there is a physical location), WebSite, WebPage, Service, Person, Article and BreadcrumbList. Together these schema types describe who the business is, what it offers, where it serves, who does the work and how its pages are organised. Add other types only when the page genuinely contains that content.

Organization is the anchor. Google recommends placing it on the home page or a single page describing the organization, lists no required properties, and recommends name, alternateName, url, logo, address, telephone, sameAs and others. LocalBusiness is appropriate for a business location customers visit; Google lists name and address as its required properties.

Service is a schema.org type for a service provided by an organization, with properties such as provider, serviceType and areaServed. Google does not offer a Service rich result, so the value here is descriptive clarity rather than a search feature. Article markup supports author, datePublished and dateModified; Google advises listing each author separately and using Person or Organization as appropriate.

Why must structured data match the visible content of the page?

Structured data must match the visible content because Google’s general guidelines require markup to be a true representation of the page and say not to mark up content that is not visible to readers. Markup that describes hidden, irrelevant or misleading content can lead to a manual action that removes rich result eligibility. For GEO, mismatched markup also creates contradictory facts.

Google’s structured data guidelines also say to put the structured data on the page it describes unless documentation says otherwise, and not to mark up irrelevant or misleading content, such as fake reviews. Google adds that it does not guarantee structured data will show up in search results even when it is valid.

In practice, three mismatches are common: an Organization block listing services the site no longer offers, an Article dateModified that changes on every deploy although the text did not, and an Arabic page carrying English-only markup copied from the English template. Each one tells systems something the visible page does not.

Is FAQPage schema still useful for search and AI?

FAQPage schema no longer produces a Google rich result. Google’s documentation changelog records that the FAQ rich result was deprecated, would stop appearing from May 7, 2026, and is no longer shown in Google Search results; its documentation was removed in June 2026. FAQPage remains a valid schema.org type, but it is no longer a Google rich-result tactic.

This reverses years of common advice, and many older guides still recommend FAQPage for extra search real estate. Before that deprecation, Google had already limited FAQ rich results in 2023 to well-known, authoritative government and health websites.

What still matters is the visible FAQ content itself. Clear question-and-answer sections help readers and give AI systems self-contained passages to quote. If you keep FAQPage markup, it must match the visible questions and answers exactly, and you should not expect a Google search feature from it.

How should Arabic pages declare their language in structured data?

Arabic pages should declare their language in structured data with the inLanguage property, using a BCP 47 code such as ar, or a regional code such as ar-SA where relevant. Schema.org defines inLanguage as the language of the content and recommends IETF BCP 47 codes. The Arabic markup should also use the Arabic names and descriptions shown on the Arabic page.

inLanguage applies to CreativeWork types, which include WebPage, WebSite and Article. Set it on each page’s WebPage node and on Article nodes. The Organization can keep the same @id across languages; on the Arabic page, its name can be the official Arabic name with the Latin name in alternateName.

Structured data does not replace hreflang. Google uses hreflang annotations to understand localized versions of a page, and requires each language version to list itself and every other version. Keep both signals consistent: the page declared as ar in markup should be the page hreflang marks as ar.

How do you validate structured data?

Validate structured data with two tools for two purposes. Google’s Rich Results Test at search.google.com/test/rich-results checks eligibility for Google rich results. The Schema Markup Validator at validator.schema.org checks schema.org markup generally and summarises the extracted data graph. After deployment, Search Console’s URL Inspection tool confirms what Google found on the live page.

The Schema Markup Validator extracts JSON-LD, RDFa and Microdata and flags syntax errors, which makes it the right tool for types such as Service that have no Google rich result. The Rich Results Test is the right tool for Google features such as breadcrumbs or Article.

Validation catches syntax and missing properties. It does not check truth. A human still has to compare the extracted values with the visible page, in both languages, and confirm that the @id references resolve to the intended nodes.

  • Run the Rich Results Test for Google-supported features.
  • Run the Schema Markup Validator to inspect the full graph.
  • Compare every extracted value with the visible page text.
  • Re-check after template changes and after content edits.
Schema types for a service business: purpose, key properties and GEO relevance
Schema typePurposeKey propertiesGEO relevance
OrganizationDefines the business entityname, alternateName, url, logo, address, telephone, sameAsAnchors entity facts and links external profiles
LocalBusinessDescribes a physical location customers visitname, address, telephone, geo, openingHoursSpecificationStates where an office actually exists
WebSiteDescribes the site and its namename, alternateName, url, publisherConnects the site name to the organization
WebPageDescribes each pagename, inLanguage, isPartOf, breadcrumbDeclares page language and place in the graph
ServiceDescribes an offered servicename, serviceType, provider, areaServedMakes services and markets explicit
PersonDescribes authors and expertsname, jobTitle, worksFor, sameAsSupports attribution and authorship
ArticleDescribes editorial contentheadline, author, publisher, datePublished, dateModified, inLanguageSignals authorship and freshness
BreadcrumbListShows site hierarchyitemListElement with ListItem item and nameClarifies how pages relate

Frequently asked questions

Do I need special schema to appear in Google AI Overviews?

No. Google states that there are no special optimizations needed for AI Overviews or AI Mode, and that you do not need to create new markup to appear in them. Pages must be indexed and eligible to show with a snippet. Standard, accurate structured data remains good practice.

Is llms.txt a replacement for structured data?

No. Google says AI text files are not needed to appear in its AI features. An llms.txt file is a separate convention some sites publish; it does not replace schema.org markup, and neither replaces clear visible content. Decide on each on its own merits.

Should every page repeat the full Organization markup?

Google recommends Organization markup on the home page or a single page describing the organization, not every page. Other pages can reference the organization by its @id, which keeps one authoritative definition and avoids conflicting copies.

Can structured data hurt a site?

Inaccurate structured data can. Google’s guidelines say a structured data issue can lead to a manual action that removes rich result eligibility, though Google states it does not affect web ranking. Misleading markup also spreads wrong facts about the business.

Which language should the Organization name use on Arabic pages?

Use the official Arabic name on Arabic pages, matching the visible text, and include the official Latin-script name in alternateName. Keep the same @id in both languages so systems understand it is one organization.

How Dfeelings applies structured data

Structured data sits in stage 03 Technical of the Dfeelings GEO Framework and draws on the entity record from stage 02 Entity. Dfeelings builds websites in-house; its own site is server-rendered Next.js with JSON-LD structured data, bilingual hreflang and llms.txt, so markup is generated from the same source as visible content in Arabic and English. Dfeelings serves businesses in Saudi Arabia and Jordan from Amman, and validates markup against visible pages in both languages. No one can guarantee an AI system will cite a business.

Official sources