Structured Data: Syntax, Vocabulary, and Page Content Alignment

Structured data gives machine-readable labels to the entities, properties, and relationships already present on a page. For a publishing team, that involves three separate jobs: establish the page facts, classify them with a vocabulary such as Schema.org, and serialize the classification in JSON-LD, Microdata, or RDFa. Search systems can use the result as an explicit clue, but valid code can still describe the wrong thing.

structured data: a large centered small globe tilted on its stand, interlocking gears, open laptop, closed folder, tilted tablet, pen, paper clips, and potted plant

Google’s working definition is a standardized format for providing information about a page and classifying its content. A recipe page, for example, can explicitly label its cooking time, ingredients, and author instead of asking a machine to infer every relationship from prose and layout.

Google uses in-page structured data as explicit information about page content. For Google Search, most supported markup uses Schema.org vocabulary expressed in JSON-LD, Microdata, or RDFa.

The phrase structured data also has a broader data-engineering meaning. AWS describes it as data organized under a predefined schema, often in rows and columns. A customer table in a relational database is structured data, but it is not automatically web-page markup. This article uses the search and publishing meaning: structured data embedded in a page to describe that page.

There is no structured-data formula. It is a semantic representation, not a calculated metric, and there is no accepted equation for a “good” amount of markup. There is also no universal benchmark for how many types or properties to add, or how much visibility they should produce. Implementation quality has to be examined as separate questions about truth, consumer requirements, rendering, and measured outcomes on the site using it.

Three terms that are often treated as synonyms belong to different layers:

  • Structured data is the general practice of expressing page information in a machine-readable form.
  • Schema.org is a vocabulary: a shared collection of types such as Article and Organization, and properties such as headline and datePublished.
  • JSON-LD is a syntax for serializing linked data in JSON. It can carry Schema.org vocabulary, but it is not the vocabulary itself.

Schema.org confirms that its vocabulary can also be expressed with Microdata or RDFa. The W3C’s JSON-LD 1.1 specification makes the deeper distinction explicit: JSON supplies the serialization syntax; linked-data contexts, identifiers, types, properties, and values supply the graph meaning.

A parser can approve the transport while the publisher is wrong about the cargo.

Use the three layers as a fault-isolation map

Stop treating a block of markup as one opaque “schema.” The first table is a diagnostic map: it assigns ownership, review questions, and characteristic failures to three layers.

LayerWhat it controlsA useful review questionTypical failure
Page factsWhat a reader can actually verify on the rendered pageIs this title, author, date, price, rating, or status visibly supported here?Markup describes hidden, stale, or different information
VocabularyWhat each entity and property meansIs this the most specific applicable type, with properties used in their documented sense?A page is labeled as the wrong kind of thing
SyntaxHow the statements are serialized in HTMLIs the JSON-LD, Microdata, or RDFa valid and available in the rendered page?Broken JSON, malformed nesting, or markup lost during rendering

Each layer can pass while another fails. Perfect JSON can contain the wrong Schema.org property. Valid vocabulary can describe a fact that the page never shows. Accurate markup in a source template can disappear from the rendered DOM. Those are different defects; “the validator is green” does not locate all of them.

Read one JSON-LD block token by token

Consider a small JSON-LD description of this article:

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "Structured Data Explained: Syntax, Vocabulary, and Visible-Content Alignment",
  "datePublished": "2026-07-02",
  "author": {
    "@type": "Organization",
    "name": "Orthotropy Team"
  }
}
</script>

This example is illustrative, not a complete implementation recipe for a particular search feature. Here the job is token anatomy—identifying what belongs to linked-data context, Schema.org vocabulary, JSON syntax, and HTML embedding:

  • @context tells a JSON-LD processor how the compact terms map to a vocabulary.
  • Article and Organization are Schema.org types.
  • headline, datePublished, author, and name are Schema.org properties.
  • The braces, quoted keys, nested object, and comma rules belong to JSON syntax.
  • The application/ld+json script element embeds the serialization in the page.

The token map establishes how the statement is built. A separate content condition still applies: the rendered page must present the same headline, date, and author identity. Changing a visible headline without updating the JSON-LD creates drift. Changing only the JSON-LD does not rewrite the article readers see.

The implementation remedy is to avoid two independently maintained copies of a fact. Store the canonical headline, author, date, availability, or price once; render the human-facing page and its structured-data projection from that same source. The markup becomes a semantic view of the content model, not a parallel database edited by hand.

Choose an encoding by its maintenance risk

Google accepts three formats for its supported structured-data features:

FormatWhere the markup livesOperational strengthMain maintenance risk
JSON-LDA JSON block in a script element, separate from visible proseKeeps nested entity descriptions together and is usually easier to generate from a content modelSeparate markup can drift from visible content if it uses a different source
MicrodataAttributes attached to the HTML elements that present the contentKeeps many values physically close to what readers seeDense attributes and nested items can become difficult to maintain in complex templates
RDFaHTML attributes that express linked-data statementsCan describe entities and relationships directly in the document structureTeams unfamiliar with RDFa can misread or break the relationships during template changes

Google generally recommends JSON-LD because it is usually easier to implement and maintain at scale, while also saying that all three formats are acceptable when valid and properly implemented for the feature. That recommendation concerns operation. JSON-LD does not make an inaccurate claim more accurate, and Microdata does not automatically prove that content is visible or current.

Choose from the table by the publishing system’s likely maintenance failure. The selected format must be generated consistently, tested in rendered output, and updated from the same content fields. A CMS that already produces reliable Microdata does not need a risky rewrite merely to imitate a preference. A new implementation with a clean content model will often find JSON-LD simpler.

Start feature eligibility with the consuming system

Schema.org maintains an open vocabulary for entities, their properties, their relationships, and actions. Google Search consumes parts of that vocabulary, but Google explicitly tells publishers to use its Search Central documentation as the authority for Google behavior. Schema.org contains types and properties that Google does not require or expose as rich results; another service may still use them.

The vocabulary-to-consumer boundary creates two different questions:

  1. Is this a coherent Schema.org description? That is a general vocabulary question.
  2. Is this page eligible for a named Google search feature? That is a product-specific question with supported types, required properties, quality policies, and access requirements.

Do not answer the second question by browsing Schema.org until you find a property that sounds useful. Start with the consuming feature’s documentation, then use Schema.org to understand the chosen terms. Google advises supplying fewer complete and accurate recommended properties instead of filling every possible property with weak, malformed, or uncertain values.

Specificity also needs judgment. Use the most specific type that truthfully describes the page’s main item, not the most elaborate type that appears to unlock a result. An article can discuss a product without becoming a Product page. A software documentation page can mention a question without becoming an FAQPage. Vocabulary selection is classification, not decoration.

Audit each marked claim against the visible state

JSON-LD normally sits in a script block and is not displayed as article prose. That is allowed. The requirement is about the facts described, not whether a reader sees curly braces.

Google says not to mark up content that users cannot see and requires structured data to represent the page truthfully. If markup describes a performer, the body must describe that performer. The same principle applies to a B2B SaaS page:

  • A price in markup needs a matching, current price visible under the same conditions.
  • A dateModified value needs a real content change, not an automated daily refresh.
  • An aggregate rating needs genuine ratings represented on the page and governed by the relevant feature policy.
  • An author or organization identity needs to match the visible byline and the entity the page actually attributes.
  • Availability, product name, FAQ answer, and event details cannot live only in JSON-LD as search-engine-facing claims.

Google distinguishes technical validity from quality. Google Search Central’s General Structured Data Guidelines note that markup can be syntactically correct yet remain ineligible when it is misleading, irrelevant, incomplete for a supported feature, or describes content hidden from users.

The review deliverable is a trace for one fact all the way through the system:

Source field → rendered statement → structured property → fetched DOM → consumer report

For an article headline, inspect the editorial source, the visible H1, the headline value, the HTML Google can render, and any production report that detects the item. For a product offer, repeat the trace for currency, amount, availability, and validity dates. The trace isolates a broken handoff more reliably than comparing two large blobs of HTML and JSON by eye.

When conditional content is involved, compare the same state. A logged-out price, regional offer, expiring promotion, or JavaScript-rendered value can be accurate in one session and contradictory in another. Document which version a crawler receives and which version a user can see before calling the markup aligned.

Match each validator to one question

Validation tools answer different questions. Use the table as a capability boundary for four checks, not as four interchangeable approvals:

CheckWhat it can establishWhat it cannot establish
JSON or HTML syntaxThe serialization can be parsedThe vocabulary is appropriate or the facts are true
Schema Markup ValidatorSchema.org items and properties can be recognized without Google feature-specific rulesEligibility for a Google rich result
Rich Results TestGoogle can detect supported markup and assess feature-specific technical eligibilityGuaranteed display, ranking, factual truth, or future template stability
Rendered-page and content reviewMarkup is present after rendering and agrees with what readers can verifyWhether a search system will choose to use or display it

Google recommends the Rich Results Test for Google features and the Schema Markup Validator for general Schema.org validation. Use URL-based testing when possible so the check includes fetching and rendering behavior, then inspect the relevant Search Console reports after deployment. A pasted code fragment can pass even when the production template serves something different.

Validation also belongs in the release path. A small implementation should cover representative page types, required and recommended properties that genuinely apply, both populated and missing-field states, and the rendered output. Re-run those checks when templates, content models, localization, pricing systems, or client-side rendering change. Structured data is production code tied to editorial data; either side can introduce drift.

Measure search outcomes instead of promising them

For search, the output of implementation is eligibility for a supported feature, not a promised appearance. Google does not guarantee that valid markup will produce an enhanced result, and its structured-data policies separate a manual action against rich-result eligibility from ordinary web-search ranking.

It is therefore unsafe to sell schema markup as a ranking switch. A page still needs to be crawlable, indexable, useful, relevant, and compliant with the policies for the intended feature. Even then, Google may choose a normal text result or no result at all.

There is no portable performance benchmark. Google recommends a before-and-after test on selected pages with enough historical data, followed through Search Console. Use that as a measurement design, not a promise of lift. Record which pages changed, which markup was added, when Google detected it, which search appearances occurred, and which business outcome matters. Seasonality, page edits, algorithm changes, and query mix can otherwise turn a simple comparison into a false causal story.

Google documents structured data as a way to enable eligibility for supported search appearances, not a guarantee that a rich result will be shown. It recommends measuring the effect on a site’s own pages rather than applying a universal lift benchmark.

Require separate evidence for AI-search claims

Structured data is often promoted as a direct route to AI citations. The documented evidence is narrower.

For Google’s AI Overviews and AI Mode, Google says normal Search eligibility applies. It recommends that important content be available as visible text and that structured data match that text. It also says there is no special Schema.org markup required for those features.

Google requires no special Schema.org data for AI Overviews or AI Mode and does not promise inclusion. Its explicit structured-data recommendation for these features is alignment with visible page text.

Accurate entity and relationship labels remain part of sound search publishing, and some consumers may document specific uses. The supported conclusion nevertheless stops before “add JSON-LD and earn citations.” For any other AI answer engine, look for current owner documentation that names the markup it consumes and the outcome it affects. In the absence of that evidence, treat a citation claim as a hypothesis to test, not a platform rule.

The practical GEO work stays with content integrity: put the substantive answer in visible text, support factual claims, keep entity names and dates consistent, and use markup to describe that same source—not to create an AI-only version of it.

Release in dependency order

The final checklist is an execution sequence. Use it when a real consumer recognizes a type that truthfully fits the page and the team can maintain the implementation:

  1. Name the consumer and intended feature or interoperability need.
  2. Inventory the facts already visible on the page and identify their source fields.
  3. Choose the most specific truthful type and only the properties that apply.
  4. Generate visible content and structured values from the same source wherever possible.
  5. Validate the general Schema.org description and the consumer-specific requirements separately.
  6. Inspect the rendered URL, not only a code fragment or source template.
  7. Deploy to a bounded page set, monitor detection and outcomes, and define a review trigger for content or template changes.

Stop the release when the implementation depends on hidden claims, guessed values, copied markup for a different page type, or a promised ranking or citation outcome that the consuming platform has not documented.

Approve structured data as a semantic projection, not as a second content source: the chosen consumer, truthful type, rendered claim trace, appropriate validator, and measurement plan must all be identifiable, and no marked claim may exceed what the visible page supports.

The useful output is not the largest graph or the most properties. It is an accurate description that reduces ambiguity, survives publishing changes, and remains legible to both people and machines.

Frequently asked questions

Can one page contain more than one structured data type?

A page can describe several truthful items, such as an article, its breadcrumb trail, and an embedded video. Google’s general structured data guidelines allow items to be nested under a main item or declared individually; when separate items need to refer to the same entity, connect them with a shared @id. Keep the page’s primary type explicit and mark only items visible to users, because adding unrelated types expands the graph without clarifying the page’s main purpose.

Does JSON-LD have to be placed in the HTML head?

JSON-LD may appear in a <script type="application/ld+json"> element in either the HTML <head> or <body>. Google’s structured data introduction recognizes both placements and can also read markup injected into the rendered page. Choose one template-owned location, generate it from the same content fields as the visible page, and inspect the rendered URL so component duplication does not emit conflicting copies.

Can JavaScript generate structured data after the initial HTML response?

Google can process JavaScript-generated structured data when it is present in the rendered DOM, and its JavaScript implementation guide documents both tag-manager and custom-script approaches. Test the deployed URL with the Rich Results Test rather than validating only the script source; for fast-changing product price or availability, prefer server-rendered values when possible because Google warns that dynamically generated markup can be crawled less frequently and less reliably for shopping uses.

One person. A whole marketing team.

Invite only