wp_scrub_utf8( string $text ): string
- Since
- 6.9.0
- Source
wp-includes/utf8.php:109
Description
Knowing what to do in the presence of text encoding issues can be complicated.
This function replaces invalid spans of bytes to neutralize any corruption that may be there and prevent it from causing further problems downstream.
However, it’s not always ideal to replace those bytes. In some settings it may be best to leave the invalid bytes in the string so that downstream code can handle them in a specific way. Replacing the bytes too early, like escaping for HTML too early, can introduce other forms of corruption and data loss.
When in doubt, use this function to replace spans of invalid bytes.
Replacement follows the “maximal subpart” algorithm for secure and interoperable strings. This can lead to sequences of multiple replacement characters in a row.
Example:
// Valid strings come through unchanged.
'test' === wp_scrub_utf8( 'test' );
// Invalid sequences of bytes are replaced.
$invalid = "the byte xC0 is never allowed in a UTF-8 string.";
"the byte \u{FFFD} is never allowed in a UTF-8 string." === wp_scrub_utf8( $invalid, true );
'the byte � is never allowed in a UTF-8 string.' === wp_scrub_utf8( $invalid, true );
// Maximal subparts are replaced individually.
'.�.' === wp_scrub_utf8( ".\xC0." ); // C0 is never valid.
'.�.' === wp_scrub_utf8( ".\xE2\x8C." ); // Missing A3 at end.
'.��.' === wp_scrub_utf8( ".\xE2\x8C\xE2\x8C." ); // Maximal subparts replaced separately.
'.��.' === wp_scrub_utf8( ".\xC1\xBF." ); // Overlong sequence.
'.���.' === wp_scrub_utf8( ".\xED\xA0\x80." ); // Surrogate half. Note! The Unicode Replacement Character is itself a Unicode character (U+FFFD).
Once a span of invalid bytes has been replaced by one, it will not be possible to know whether the replacement character was originally intended to be there or if it is the result of scrubbing bytes. It is ideal to leave replacement for display only, but some contexts (e.g. generating XML or passing data into a large language model) require valid input strings.
Compatibility
- WordPress
- since 6.9.0
- PHP
- 7.4–8.6-dev
- 6.7.7
- 6.8.8
- 6.9.7
- 7.0.4
- 7.1.0
Present in 3 of the 5 tracked releases, added in 6.9.0, and compiles on PHP 7.4 through 8.6-dev.
Parameters
$textstring- String which is assumed to be UTF-8 but may contain invalid sequences of bytes.
Return value
string- Input text with invalid sequences of bytes replaced with the Unicode replacement character.
Performance profile
How much work a call to wp_scrub_utf8() does, and what it touches: the algorithmic scaling, the Zend instruction count per call across PHP versions, the hooks it hands control to, and the core code that calls it. Measured from the compiled opcodes, not a stopwatch, so every number is identical on any machine running the same PHP version, and every function in core is ranked by cost.
- Cost class
- Trivial
- Scaling
- Constant
- Instructions
- 16
- Plugin surface
- None
- Called by
- 2
Touches nothing outside its own arguments.
No loop in the body: the same number of instructions runs whatever you pass in.
Executed per call on PHP 8.5. The body compiles to 16.
Nothing here hands control to plugin code.
2 places in core call this, so the cost is paid more often than your own code shows.
What one call costs · 1 distinct outcome
One number would be a lie: the work depends on which branch runs. These are every distinct cost wp_scrub_utf8() can have, taken from its control-flow graph on PHP 8.5.
| When | Instructions | Calls it makes |
|---|---|---|
| always | 16 | mb_substitute_character(), mb_substitute_character(), mb_scrub(), mb_substitute_character() |
Across PHP versions
Compiles the same on PHP 7.4, 8.1, 8.2, 8.3, 8.4, 8.5 and 8.6-dev: 16 instructions, 16 executed per call, 0 branches. The work does not change between versions.
An instruction is not a fixed amount of time, so a matching count is not necessarily the same speed; what it rules out is a difference in the work itself.
Used by · 2
- WP_HTML_Processor::serialize_token()Serializes the currently-matched token.
- wp_check_invalid_utf8()Checks for invalid UTF8 in a string.
Source code
function wp_scrub_utf8( $text ) { /* * While it looks like setting the substitute character could fail, * the internal PHP code will never fail when provided a valid * code point as a number. In this case, there’s no need to check * its return value to see if it succeeded. */ $prev_replacement_character = mb_substitute_character(); mb_substitute_character( 0xFFFD ); $scrubbed = mb_scrub( $text, 'UTF-8' ); mb_substitute_character( $prev_replacement_character ); return $scrubbed; }Changelog
Introduced in 6.9.0. Unchanged from 6.9.7 through 7.1.0.
Signature, return type and hooks compared across 3 parsed releases.
About this page
- Parsed data
- Generated from the wordpress-develop 7.1.0 tag, from
src/wp-includes/utf8.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it. - Corrections
- Something wrong on this page? Report it and it gets fixed in the next regeneration.