Schema Planning for AI Search: How to Build a Connected JSON-LD Graph That Machines Can Trust
Table of Contents (Click to show/hide)









Schema planning is deciding which entities, page types, facts, and relationships your website should describe through structured data before writing JSON-LD. Work from the business outward: identify the organisation, then the people who represent or create for it, then its products and services. Finally, look at each page's content and decide which of those real entities it describes. The JSON-LD should show those relationships rather than make every page look like an isolated fact sheet.
My rule is straightforward: every property should be true, useful, supportable, and maintainable.
Take an SEO audit page. The page describes a specific service, and the service is provided by a named business or practitioner. Give the page, service, and provider their own identifiers. Then let the page's mainEntity point to the service and the service's provider point to the business. Repeating separate Service and Organization blocks without those references leaves the relationship implicit. The markup should make the same connection a reader can see on the page.
This is the next practical step after retrieval context engineering: make the website's relationships explicit without making claims the content cannot support.
What Schema Can Actually Do For AI Search
Structured data provides a vocabulary for describing things. An article can identify its author and publisher. A service can identify its provider. An image can describe its subject, creator, credit, and copyright holder where those facts are known and accurate. These relationships make your intended meaning explicit for systems that consume the markup.
That is different from proving that a particular AI system will retrieve, rank, or cite the page because of your schema.
Google's AI search guidance says there are no special schema requirements for AI Overviews or AI Mode. Eligible supporting pages must be indexed and eligible for a search snippet. Ordinary SEO practices still apply, including crawl access, internal links, and structured data that agrees with visible content.
Google also says new AI text files are unnecessary for its Search AI features. That is a statement about Google, not a rule for every agent or retrieval product. Cloudflare's agent-readiness guidance looks beyond Google Search: it examines crawler access, sitemaps, and machine-readable content such as Markdown responses. Cloudflare describes llms.txt as a possible directory for agents, but does not include it in its default readiness check. Its guidance does not prove that every agent reads llms.txt or structured data.
My view is to prepare for more than one consumer without pretending to know each system's internal rules. Maintain crawlable pages, accurate sitemaps, and a clear, connected schema graph because they make your information easier to find and interpret. Consider llms.txt or alternate text formats where a specific audience or system can use them, and keep them in sync with the site. None of these is a universal citation requirement. Treat schema as an information quality project, then measure search outcomes separately.
Start With An Entity Inventory
Before opening a generator, write down what exists in the business and what exists on the website. Separate the two.
The business might include an organisation, a founder, three services, and a genuine office. The website might include a homepage, service pages, a contact page, articles, and category listings. Each webpage is a document about something; it is not automatically the same thing as that something.
For a consulting site, I would record five columns: entity, public name, authoritative URL, supporting evidence, and maintenance owner. This catches problems before they become code. Who owns the service description? Is the address public and current? Does the author have a biography? Is an old social account still an official identity?
Start small. You do not need a separate Brand node if it adds no useful distinction from the organisation. A founder is not necessarily the author of every post. A national service area does not justify inventing an office in every city.
Schema.org's Organisation vocabulary provides many possible properties. Treat that list as available vocabulary, not a form you must completely fill in.
The inventory's main job is to establish who can defend each fact. If nobody can maintain a price or confirm a claimed relationship, leave that property out until the underlying information is reliable.
Give Each Entity A Stable Identifier
In JSON-LD, @id identifies a node. Other nodes can use the same identifier to refer to it. The JSON-LD specification describes this linking model; @graph lets you group node descriptions within a document.
For example, a fictional agency might give its business and SEO audit service these IDs:
Follow the references: the page's mainEntity uses the service's @id, and the service's provider uses the organisation's @id. That is the connection. The values are fictional; replace them with supported facts rather than deploying this example unchanged.
Use absolute HTTPS identifiers and one canonical hostname. Keep the organisation identifier stable when its description changes. Give individual articles and services their own identifiers. Avoid timestamps or randomly generated IDs that change on each render.
The fragment after # is a naming convention. Choosing #organisation instead of #business does not create an SEO advantage. Consistent references matter more than the label.
An identifier is also not a promise that every consumer will fetch its URL and assemble your entire website graph. Include the relevant supporting node descriptions on the page where practical. Do not assume a reference alone supplies all the facts a search feature requires.
Separate Shared Identity From Page Meaning
I plan two layers, even if they ultimately render in one script.
The shared layer describes persistent entities: the organisation, website, relevant people, and genuinely distinct brands. The page layer describes a particular document and its main content: a service, article, product, or collection.
Shared does not have to mean one global custom-code block. It means one maintained source for those facts. A CMS template can repeat a consistent organisation definition alongside each article. That is different from maintaining five conflicting business descriptions by hand.
The service-page example above keeps each entity distinct while connecting them in one graph. It demonstrates relationships, not a complete rich-result implementation.
Service is useful vocabulary for describing an offer's provider and nature. Its existence in schema.org does not establish a dedicated Google service rich result.
Choose Types By What The Page Does
Ask what a visitor comes to the page to understand or do. Then select the most accurate type.
Use supporting BreadcrumbList and ImageObject descriptions when they represent real page elements. A case study can be an article; a portfolio index can be a collection. A contact form does not turn its whole page into a service.
One important 2026 correction: Google discontinued FAQ rich results from May 7, 2026. FAQPage remains vocabulary, but adding it should not be presented as a way to earn Google FAQ snippets. Keep useful questions for readers and decide whether their markup serves a real consumer.
Likewise, do not label every consulting engagement as a product to chase a search presentation. Start with the commercial reality, then check the current feature documentation for any intended rich result.
Apply The Four-Question Property Test
Most schema problems start with optional fields. Someone sees a property and assumes more information must be better. I use this decision matrix instead:
For identity, start with name, URL, and appropriate contact information. Use sameAs for genuine identity references, such as official profiles, rather than every article mentioning the business. Describe an actual founder relationship if it exists; do not assign one because the person currently manages the website.
For blog posts, prioritise headline, description, author, publisher, and accurate dates. Connect mainEntityOfPage to the article's canonical webpage. A modification date should represent a real update, not today's date inserted on every request.
A date detail worth testing: Search Engine Journal argues that a visible recent date may affect which result people click, while Google allows a publication date, an updated date, or both. That is not evidence that showing one date reliably beats two for CTR. Choose the display that helps readers understand the article's history, keep visible and structured dates consistent, and test CTR on your own pages after meaningful updates.
The full articleBody is optional vocabulary. Do not duplicate thousands of words solely to make a graph bigger. If maintaining a clean plain-text version creates an unnecessary second content system, omit it and keep the actual article accessible.
For services, explain the provider and service type before attempting complicated offers. Add pricing only when the commercial conditions are clear. A booking URL can describe an access channel if that accurately reflects how someone obtains the service.
For images, ImageObject can describe the actual asset, creator, dimensions, and credit. A filename is not proof of authorship. Do not claim copyright ownership of licensed or third-party work. Keep useful alt text in the image's HTML too; schema does not replace it. My image SEO guide covers the wider asset workflow.
Keep Ratings Out Of The Identity Shortcut
An aggregateRating is a rating claim. It should never be a decorative trust badge added because the business wants stars.
Google's review guidance makes the distinction explicit: organisation review markup is supported "only for sites that capture reviews about other organizations." The same principle applies to local businesses. A site that marks up ratings about itself is ineligible for Google's review-star feature, even if the reviews arrive through an embedded third-party widget.
A genuinely independent review site may mark up ratings about another eligible business or product when it meets Google's content and property rules. But one writer's product comparison is a Review, not automatically an AggregateRating. Use AggregateRating only when you have a real aggregate of user ratings for the specific item being rated; identify that item correctly and show the underlying rating information to readers. Google's rules for organisation/local-business ratings also say ratings must come directly from users, rather than an editor compiling them.
That does not mean testimonials are useless. They can be valuable visible evidence for visitors. It means testimonials and eligibility for a particular search feature are different questions.
During an audit, ask what the rating describes, where it comes from, whether the underlying reviews are visible, and whether the relevant feature permits that use case. Never invent a total or copy a rating across unrelated offers.
Do not strip away all useful business identity because one rating property is inappropriate. Remove the unsupported claim and keep the accurate relationships.
A Service-Site Audit I Would Run
For a professional-service website such as DEANLONG.io, I would begin with five representative URLs: homepage, service page, contact page, blog index, and one article. This is a proposed audit method, not a claim that the site's current graph has already been validated.
First, collect every JSON-LD block from those pages. Record entity identifiers and look for conflicting names, URLs, descriptions, and addresses. A repeated organisation node is not automatically a problem; contradictory definitions are.
Second, check page intent. The blog index should describe a collection, and an individual post should describe its actual article. The breadcrumb vocabulary can express the navigation hierarchy, but the hierarchy must make sense to someone using the site.
Third, trace relationships. Does the article point to its real author and publisher? Does the service identify its provider? Does the main image describe the asset displayed on that page? Are there unnecessary standalone identities that should share a maintained identifier?
Fourth, fix visible content where evidence is missing. If a service page never explains what an audit includes, adding a detailed hidden service description creates a mismatch. Improve the page first. The same principle applies to crawlable embedded content: markup cannot rescue information a crawler cannot access.
Finally, choose one improvement to roll out across the relevant template. Fixing a broken author binding across fifty articles is often more useful than adding ten optional properties to the homepage.
Validate In Layers Before Publishing
Validation has several jobs. Keep them separate so a green result does not become false confidence.
First, parse the JSON. This catches syntax problems such as missing commas and broken quotation marks. It says nothing about whether a claim is true.
Second, use the Schema Markup Validator to inspect vocabulary and relationships. Third, use Google's Rich Results Test for the search features Google supports. A valid Service description may not correspond to a feature that tool reports.
Fourth, compare the output with the rendered page. Google's structured-data policies require representative, accurate content and do not guarantee a rich result merely because validation passes.
My release checklist is:
After publishing, monitor indexation and applicable enhancement reports. Evaluate traffic and conversions alongside other changes. If you changed the introduction, internal links, and schema together, an uplift cannot honestly be attributed to JSON-LD alone.
My Take
I would start a schema project with an entity register, a page-intent map, and a maintenance owner. The code comes after those decisions.
Build a small graph that explains a real service page or article. Validate it against the content. Fix the template, then expand only where the extra information earns its place.
For AI search, that is a defensible investment in clearer website information. It is not a substitute for original expertise, useful pages, crawl access, or a coherent commercial offer.
If your current setup is a stack of unrelated snippets, the next move is a focused audit of one representative page. Get in touch to review the content, schema, and implementation together.
Sources
- Google Search: AI features and your website
- Google Search documentation updates: FAQ rich-result retirement and llms.txt clarification
- Cloudflare: Agent Readiness score
- Google Search: Byline dates
- Search Engine Journal: Article date strategies
- Google Search: General structured-data guidelines
- Google Search: Review snippet structured data
- Schema.org Markup Validator
- Google Rich Results Test
- W3C: JSON-LD 1.1
- Schema.org: Organization
- Schema.org: BlogPosting
- Schema.org: Service
- Schema.org: ImageObject
- Schema.org: BreadcrumbList







