wppaste
WordPress

wp_scrub_utf8( string $text ): string

Since
6.9.0
Source
wp-includes/utf8.php:109
Replaces ill-formed UTF-8 byte sequences with the Unicode Replacement Character.

Description

Knowing what to do in the presence of text encoding issues can be complicated.
This function replaces invalid spans of bytes to neutralize any corruption that may be there and prevent it from causing further problems downstream.

However, it’s not always ideal to replace those bytes. In some settings it may be best to leave the invalid bytes in the string so that downstream code can handle them in a specific way. Replacing the bytes too early, like escaping for HTML too early, can introduce other forms of corruption and data loss.

When in doubt, use this function to replace spans of invalid bytes.

Replacement follows the “maximal subpart” algorithm for secure and interoperable strings. This can lead to sequences of multiple replacement characters in a row.

Example:

// Valid strings come through unchanged.
'test' === wp_scrub_utf8( 'test' );

// Invalid sequences of bytes are replaced.
$invalid = "the byte xC0 is never allowed in a UTF-8 string.";
"the byte \u{FFFD} is never allowed in a UTF-8 string." === wp_scrub_utf8( $invalid, true );
'the byte � is never allowed in a UTF-8 string.' === wp_scrub_utf8( $invalid, true );

// Maximal subparts are replaced individually.
'.�.' === wp_scrub_utf8( ".\xC0." ); // C0 is never valid.
'.�.' === wp_scrub_utf8( ".\xE2\x8C." ); // Missing A3 at end.
'.��.' === wp_scrub_utf8( ".\xE2\x8C\xE2\x8C." ); // Maximal subparts replaced separately.
'.��.' === wp_scrub_utf8( ".\xC1\xBF." ); // Overlong sequence.
'.���.' === wp_scrub_utf8( ".\xED\xA0\x80." ); // Surrogate half.

Note! The Unicode Replacement Character is itself a Unicode character (U+FFFD).
Once a span of invalid bytes has been replaced by one, it will not be possible to know whether the replacement character was originally intended to be there or if it is the result of scrubbing bytes. It is ideal to leave replacement for display only, but some contexts (e.g. generating XML or passing data into a large language model) require valid input strings.

Compatibility

WordPress
since 6.9.0
PHP
7.4–8.6-dev
  • 6.7.7
  • 6.8.8
  • 6.9.7
  • 7.0.4
  • 7.1.0

Present in 3 of the 5 tracked releases, added in 6.9.0, and compiles on PHP 7.4 through 8.6-dev.

Parameters

$textstring
String which is assumed to be UTF-8 but may contain invalid sequences of bytes.

Return value

string
Input text with invalid sequences of bytes replaced with the Unicode replacement character.

Performance profile

How much work a call to wp_scrub_utf8() does, and what it touches: the algorithmic scaling, the Zend instruction count per call across PHP versions, the hooks it hands control to, and the core code that calls it. Measured from the compiled opcodes, not a stopwatch, so every number is identical on any machine running the same PHP version, and every function in core is ranked by cost.

Cost class
Trivial

Touches nothing outside its own arguments.

Scaling
Constant

No loop in the body: the same number of instructions runs whatever you pass in.

Instructions
16

Executed per call on PHP 8.5. The body compiles to 16.

Plugin surface
None

Nothing here hands control to plugin code.

Called by
2

2 places in core call this, so the cost is paid more often than your own code shows.

What one call costs · 1 distinct outcome

One number would be a lie: the work depends on which branch runs. These are every distinct cost wp_scrub_utf8() can have, taken from its control-flow graph on PHP 8.5.

WhenInstructionsCalls it makes
always16mb_substitute_character(), mb_substitute_character(), mb_scrub(), mb_substitute_character()

Across PHP versions

Compiles the same on PHP 7.4, 8.1, 8.2, 8.3, 8.4, 8.5 and 8.6-dev: 16 instructions, 16 executed per call, 0 branches. The work does not change between versions.

An instruction is not a fixed amount of time, so a matching count is not necessarily the same speed; what it rules out is a difference in the work itself.

Used by · 2

Source code

	function wp_scrub_utf8( $text ) {		/*		 * While it looks like setting the substitute character could fail,		 * the internal PHP code will never fail when provided a valid		 * code point as a number. In this case, there’s no need to check		 * its return value to see if it succeeded.		 */		$prev_replacement_character = mb_substitute_character();		mb_substitute_character( 0xFFFD );		$scrubbed = mb_scrub( $text, 'UTF-8' );		mb_substitute_character( $prev_replacement_character ); 		return $scrubbed;	}

Changelog

Introduced in 6.9.0. Unchanged from 6.9.7 through 7.1.0.

  1. 6.9.7
  2. 7.0.4
  3. 7.1.0

Signature, return type and hooks compared across 3 parsed releases.

About this page

Parsed data
Generated from the wordpress-develop 7.1.0 tag, from src/wp-includes/utf8.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it.
Corrections
Something wrong on this page? Report it and it gets fixed in the next regeneration.