Google Reveals JSON-LD Change Causes Structured Data Errors

Google's stricter JSON-LD parser now decodes HTML entities only once, breaking double-escaped text in rich results.

By Central
Google's JSON-LD parser update decodes HTML entities once, causing validation failures for double-escaped schema markup.
Highlights
  • Google's JSON-LD parser changed from double to single HTML entity decoding, breaking double-escaped characters.
  • Double-escaped ampersands like & now render as raw & in rich snippets instead of the intended symbol.
  • The fix requires using native JSON encoders instead of HTML sanitization functions in your data pipeline.

Google has announced a significant change to how it processes JSON-LD structured data, a shift that is already causing validation errors for websites that rely on double-escaped HTML entities within their schema markup. The update, which alters the parser’s behavior regarding HTML entity decoding, means that content previously rendered correctly in rich results may now appear as raw, unparsed text.

Understanding the Parser Update: From Double to Single Decoding

The core of the change revolves around how Google’s parsing engine handles HTML entities embedded within JSON-LD blocks. Historically, the parser performed a multi-step decoding process, unescaping HTML entities more than once. This lenient approach allowed for a wider range of formatting errors to be corrected automatically. Under the new system, the parser decodes HTML entities only once. For the vast majority of correctly implemented JSON-LD, this change is entirely transparent and causes no issues. However, it creates a critical failure point for any structured data that contains double-escaped characters.

When a website’s content management system or a plugin passes data through an HTML escape function, and that data already contains HTML entities, the result is a double escape. For example, an ampersand (&) that is already encoded as & within a string will, after a second pass through an HTML entity encoder, become &. Google’s previous parser would decode this correctly back to &. The new parser, decoding only once, will leave the data in a partially escaped state, such as & or &, which is not the intended character.

Google confirmed this update via a LinkedIn post from the Google Search Central team. The announcement did not provide a detailed technical breakdown but clearly indicated a move to align the parser with modern JSON and schema.org standards. This is a textbook case of a parser becoming more strict and standards-compliant, which ultimately improves data integrity but creates a migration burden for sites with legacy or poorly implemented code.

Specific Examples of the Problem: What Breaks and What Stays

To understand the practical impact, it is helpful to examine how specific characters behave before and after the update. The primary class of breakage involves double-escaped entities.

The Ampersand and Double Escapes

Consider the ampersand. A single-escaped ampersand, written as & in the HTML source, continues to function perfectly. Google decodes it once, and it renders as &. The trouble begins with a double-escaped sequence. If the source contains &, the old parser would decode it twice, producing &. The new parser decodes it once, leaving & in the final output. This raw text, rather than the intended symbol, will appear in search result snippets.

Special Characters and Unicode Symbols

The same logic applies to any HTML entity, including those for checkmarks, arrows, or special punctuation. A double-escaped checkmark entity, ✔, was previously decoded correctly to ✔. Under the new rules, the single pass leaves it as ✔. Instead of a clean checkmark in a FAQ rich result, users and search engines will see the unparsed code.

The table below summarizes the behavior for key examples:

  • Single-escaped ampersand (&): Decodes correctly to &. No issue.
  • Double-escaped ampersand (&): Decodes to &. This is broken.
  • Double-escaped checkmark (✔): Decodes to ✔. This is broken.

What Are the Key Areas of Impact on Structured Data?

This change directly threatens the validity of several common structured data types. The primary risk is not just cosmetic; it can lead to Google being unable to parse critical fields, resulting in the loss of rich result eligibility. The specific fields most at risk are those that pull content from user-generated sources, WYSIWYG editors, or database fields that undergo multiple layers of sanitization and escaping.

Product Markup and Offers

E-commerce sites are particularly vulnerable. Product titles, descriptions, and offer URLs are common sources of double-escaped entities. For example, a product name like “Müller & Sohn” that is improperly encoded can appear as “Müller & Sohn” in the snippet. More critically, the offers.url or url field, if it contains a double-escaped ampersand in a query parameter (e.g., ?utm_source=x&utm_medium=y), will fail to validate. Google will not recognize the URL, and the product rich result can be suppressed entirely, not just displayed with an error.

FAQ and HowTo Markup

FAQ and HowTo schemas, which frequently pull question and answer text from rich text editors, are another major source of this problem. A checkmark used at the end of a list item, or an arrow used to indicate a step, will fail to render correctly. The visual impact on the rich result is significant, as users will see raw entity codes instead of the intended symbols, undermining trust and clarity.

Organization and Person Markup

Fields like sameAs (social media profile URLs) or address (which might contain ampersands or special characters) can also be affected. While less common, any text field that undergoes a double escape cycle is at risk.

Root Causes: Where Double Escapes Come From

Identifying the source of double escapes is the first step toward a fix. The problem almost always originates from a specific pattern in the content generation pipeline. The most common scenarios involve plugins, themes, or custom code that apply an HTML entity encoding function to a string that has already been processed by a WYSIWYG editor or a database retrieval function.

A WYSIWYG editor, such as TinyMCE or Gutenberg, will automatically encode special characters into HTML entities when saving content. This is standard and desirable. When a plugin then retrieves this content and passes it through another encoding function, such as PHP’s htmlspecialchars() or a similar JavaScript function, the previously encoded entities become double-escaped. The JSON-LD block then contains the double-escaped version, which the new Google parser cannot handle.

This is not a bug in the editor or the plugin per se, but a failure in the data handling pipeline. The correct approach is to ensure that strings placed into JSON-LD are encoded only once, using the proper method for the target format.

How to Fix JSON-LD Structured Data: A Technical Guide

The remedy is straightforward in concept but may require a careful review of code. The fundamental rule is: do not use HTML entity encoding on strings destined for JSON-LD. JSON-LD is a JavaScript Object Notation format, not HTML. It has its own set of escape rules, which are different from HTML.

Use Native JSON Encoders

The correct approach is to use the native JSON encoding functions provided by your programming language. In PHP, this is json_encode(). In JavaScript and Node.js, this is JSON.stringify(). These functions automatically handle the necessary escaping for JSON format, such as escaping double quotes within strings (using backslashes, e.g., \”). They do not, and should not, convert an ampersand to &.

For example, to include a product title like “Müller & Sohn”, you should place the literal string, with the real ampersand character, into your PHP array or JavaScript object, and then let json_encode or JSON.stringify produce the final output. The result will be a perfectly valid JSON-LD block where the ampersand is represented as its literal character (which is legal in JSON strings) or, if you prefer, as a Unicode escape (\u0026).

The One Exception: Preventing XSS in JSON-LD

There is one specific case where escaping within JSON-LD is necessary: preventing a closing script tag (</script>) from breaking the HTML page. If a text value contains the sequence </, the browser could interpret it as the end of the <script type=”application/ld+json”> tag, leading to a parsing error and potentially a security vulnerability. To avoid this, you must escape the less-than sign and slash sequence.

The correct method is to use a JSON Unicode escape, not an HTML entity. Replace < with \u003c. This prevents the browser from misinterpreting the character sequence while keeping the JSON valid. Using the HTML entity &lt; would trigger the very problem this update addresses.

How to Audit Your Site for This Issue

Performing a site-wide audit to identify and fix double-escaped entities is a critical next step for any SEO or development team. The process can be broken down into a manual check and a programmatic scan.

Manual Source Code Inspection

The simplest way to check a single page is to view the page source and search within any <script type=”application/ld+json”> block. Look specifically for the strings &amp; (the HTML entity for an ampersand) and &# (the beginning of a numeric HTML entity). If you find these strings inside the JSON-LD block, you have a problem. The Rich Results Test tool can then be used to see exactly how Google interprets the data, allowing you to compare the raw source with the parsed output.

Programmatic Site-Wide Scan

For larger sites, a manual check is impractical. Tools like Screaming Frog SEO Spider can be configured with a custom extraction feature to search for specific strings within HTML elements. A custom search for ld+json combined with a regex pattern to find &amp; or &# will return a list of all affected pages. This is the most efficient way to quantify the scope of the problem and prioritize fixes.

Once the affected pages are identified, focus remediation efforts on the data sources. This typically means auditing the plugins or custom functions responsible for building the JSON-LD blocks. The fix will involve removing the final HTML escape step from the pipeline, ensuring that data flows directly into a native JSON encoder.

What Does This Mean for SEO Practitioners and Developers?

This change from Google is a clear signal that they are moving toward stricter compliance with web standards. While it creates immediate work for sites with legacy code, it ultimately makes the ecosystem healthier. Structured data that is valid and clean is more reliable for both search engines and users. From an editorial and strategic perspective, this is an opportunity to clean up technical debt rather than a crisis.

For SEO practitioners, this underscores the importance of understanding the difference between HTML and JSON escaping. It is a technical detail that can have a direct impact on the visibility of rich results. The fix is not difficult, but it requires coordination between SEOs, developers, and content managers to ensure that the content pipeline is correctly configured.

The most important takeaway is a future-proofing one. Whenever you implement structured data, treat the JSON-LD block as a distinct environment with its own rules. Rely on native JSON encoders and avoid the temptation to use HTML-focused sanitization functions. This single change will prevent not only this specific error but also a host of other potential incompatibilities as parsers continue to modernize.

Ultimately, this update is a minor but necessary tightening of the rules. It rewards clean, standards-based implementation and penalizes shortcuts. For organizations that address it promptly, the impact will be minimal. For those that ignore it, the gradual loss of rich result features will be a tangible consequence. The path forward is clear: audit your JSON-LD, fix the double escapes, and align your encoding practices with the technical reality of the modern web.

Share This Article