_wp_scan_utf8( string $bytes, int $at, int $invalid_length, int|null $max_bytes = null, int|null $max_code_points = null, bool|null $has_noncharacters = null ): int
- Since
- 6.9.0
- Source
wp-includes/compat-utf8.php:47
Description
This is a low-level tool to power various UTF-8 functionality.
It scans through a string until it finds invalid byte spans.
When it does this, it does three things:
- Assigns
$atto the position after the last successful code point. - Assigns
$invalid_lengthto the length of the maximal subpart of the invalid bytes starting at$at. - Returns how many code points were successfully scanned.
This information is enough to build a number of useful UTF-8 functions.
Example:
// ñ is U+F1, which in ISO-8859-1/latin1/Windows-1252/cp1252 is 0xF1.
"Pi\xF1a" === $pineapple = mb_convert_encoding( "Piña", 'Windows-1252', 'UTF-8' );
$at = $invalid_length = 0;
// The first step finds the invalid 0xF1 byte.
2 === _wp_scan_utf8( $pineapple, $at, $invalid_length );
$at === 2; $invalid_length === 1;
// The second step continues to the end of the string.
1 === _wp_scan_utf8( $pineapple, $at, $invalid_length );
$at === 4; $invalid_length === 0; Note! While passing an options array here might be convenient from a calling-code standpoint, this function is intended to serve as a very low-level foundation upon which to build higher level functionality. For the sake of keeping costs explicit all arguments are passed directly.
Compatibility
- WordPress
- since 6.9.0
- PHP
- 7.4–8.6-dev
- 6.7.7
- 6.8.8
- 6.9.7
- 7.0.4
- 7.1.0
Present in 3 of the 5 tracked releases, added in 6.9.0, and compiles on PHP 7.4 through 8.6-dev.
Parameters
$bytesstring- UTF-8 encoded string which might include invalid spans of bytes.
$atint- Where to start scanning.
$invalid_lengthint- Will be set to how many bytes are to be ignored after
$at. $max_bytesint|nulloptional- Stop scanning after this many bytes have been seen.Default:
null $max_code_pointsint|nulloptional- Stop scanning after this many code points have been seen.Default:
null $has_noncharactersbool|nulloptional- Set to indicate if scanned string contained noncharacters.Default:
null
Return value
int- How many code points were successfully scanned.
Performance profile
How much work a call to _wp_scan_utf8() does, and what it touches: the algorithmic scaling, the Zend instruction count per call across PHP versions, the hooks it hands control to, and the core code that calls it. Measured from the compiled opcodes, not a stopwatch, so every number is identical on any machine running the same PHP version, and every function in core is ranked by cost.
- Cost class
- Trivial
- Scaling
- Constant
- Instructions
- 22–121
- Plugin surface
- None
- Called by
- 7
Touches nothing outside its own arguments.
No loop in the body: the same number of instructions runs whatever you pass in.
Executed per call on PHP 8.5, depending on the branch taken. The body compiles to 278.
Nothing here hands control to plugin code.
7 places in core call this, so the cost is paid more often than your own code shows.
What one call costs · 3 distinct outcomes
One number would be a lie: the work depends on which branch runs. These are every distinct cost _wp_scan_utf8() can have, taken from its control-flow graph on PHP 8.5.
| When | Instructions | Calls it makes |
|---|---|---|
!$i | 22 | none |
$i && $count && !$max_count && !$end && !$bytes && !$b1 && $b3 | 64–67 | strspn(), ord(), ord(), ord() |
$i && $count && !$max_count && !$end && !$bytes && !$b1 && $b3 | 85–121 | strspn(), ord(), ord(), ord(), ord(), ord() |
This body has more branch combinations than are worth enumerating, so the table covers the outcomes found first rather than every one that exists.
Across PHP versions
| PHP | Compiled | Executed | Branches | Notes |
|---|---|---|---|---|
| 8.6-dev | 278 | 22–121 | 82 | |
| 8.5 | 278 | 22–121 | 82 | |
| 8.4 | 278 | 22–121 | 82 | 13 fewer instructions than PHP 8.3 |
| 8.3 | 291 | 26–131 | 82 | |
| 8.2 | 291 | 26–131 | 82 | |
| 8.1 | 291 | 26–131 | 82 | 1 fewer instruction than PHP 7.4 |
| 7.4 | 292 | 26–131 | 82 |
An instruction is not a fixed amount of time, so a matching count is not necessarily the same speed; what it rules out is a difference in the work itself.
Used by · 7
- _mb_ord()Internal compat function to mimic mb_ord().
- _wp_is_valid_utf8_fallback()Fallback mechanism for safely validating UTF-8 bytes.
- _wp_scrub_utf8_fallback()Fallback mechanism for replacing invalid spans of UTF-8 bytes.
- _wp_utf8_codepoint_count()Returns how many code points are found in the given UTF-8 string.
- _wp_utf8_codepoint_span()Given a starting offset within a string and a maximum number of code points, return how many bytes are occupied by the span of characters.
- _wp_utf8_decode_fallback()Converts a string from UTF-8 to ISO-8859-1, maintaining backwards compatibility with the deprecated function from the PHP standard library.
- antispambot()Obscures email addresses in HTML to prevent spam bots from harvesting them.
Source code
function _wp_scan_utf8( string $bytes, int &$at, int &$invalid_length, ?int $max_bytes = null, ?int $max_code_points = null, ?bool &$has_noncharacters = null ): int { $byte_length = strlen( $bytes ); $end = min( $byte_length, $at + ( $max_bytes ?? PHP_INT_MAX ) ); $invalid_length = 0; $count = 0; $max_count = $max_code_points ?? PHP_INT_MAX; $has_noncharacters = false; for ( $i = $at; $i < $end && $count <= $max_count; $i++ ) { /* * Quickly skip past US-ASCII bytes, all of which are valid UTF-8. * * This optimization step improves the speed from 10x to 100x * depending on whether the JIT has optimized the function. */ $ascii_byte_count = strspn( $bytes, "\x00\x01\x02\x03\x04\x05\x06\x07\x08\x09\x0a\x0b\x0c\x0d\x0e\x0f" . "\x10\x11\x12\x13\x14\x15\x16\x17\x18\x19\x1a\x1b\x1c\x1d\x1e\x1f" . " !\"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\\]^_`abcdefghijklmnopqrstuvwxyz{|}~\x7f", $i, min( $end - $i, $max_count - $count ) ); if ( $count + $ascii_byte_count >= $max_count ) { $at = $i + ( $max_count - $count ); $count = $max_count; return $count; } $count += $ascii_byte_count; $i += $ascii_byte_count; if ( $i >= $end ) { $at = $end; return $count; } /** * The above fast-track handled all single-byte UTF-8 characters. What * follows MUST be a multibyte sequence otherwise there’s invalid UTF-8. * * Therefore everything past here is checking those multibyte sequences. * * It may look like there’s a need to check against the max bytes here, * but since each match of a single character returns, this functions will * bail already if crossing the max-bytes threshold. This function SHALL * NOT return in the middle of a multi-byte character, so if a character * falls on each side of the max bytes, the entire character will be scanned. * * Because it’s possible that there are truncated characters, the use of * the null-coalescing operator with "\xC0" is a convenience for skipping * length checks on every continuation bytes. This works because 0xC0 is * always invalid in a UTF-8 string, meaning that if the string has been * truncated, it will find 0xC0 and reject as invalid UTF-8. * * > [The following table] lists all of the byte sequences that are well-formed * > in UTF-8. A range of byte values such as A0..BF indicates that any byte * > from A0 to BF (inclusive) is well-formed in that position. Any byte value * > outside of the ranges listed is ill-formed. * * > Table 3-7. Well-Formed UTF-8 Byte Sequences * ╭─────────────────────┬────────────┬──────────────┬─────────────┬──────────────╮ * │ Code Points │ First Byte │ Second Byte │ Third Byte │ Fourth Byte │ * ├─────────────────────┼────────────┼──────────────┼─────────────┼──────────────┤ * │ U+0000..U+007F │ 00..7F │ │ │ │ * │ U+0080..U+07FF │ C2..DF │ 80..BF │ │ │ * │ U+0800..U+0FFF │ E0 │ A0..BF │ 80..BF │ │ * │ U+1000..U+CFFF │ E1..EC │ 80..BF │ 80..BF │ │ * │ U+D000..U+D7FF │ ED │ 80..9F │ 80..BF │ │ * │ U+E000..U+FFFF │ EE..EF │ 80..BF │ 80..BF │ │ * │ U+10000..U+3FFFF │ F0 │ 90..BF │ 80..BF │ 80..BF │ * │ U+40000..U+FFFFF │ F1..F3 │ 80..BF │ 80..BF │ 80..BF │ * │ U+100000..U+10FFFF │ F4 │ 80..8F │ 80..BF │ 80..BF │ * ╰─────────────────────┴────────────┴──────────────┴─────────────┴──────────────╯ * * @see https://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-3/#G27506 */ // Valid two-byte code points.Changelog
Introduced in 6.9.0. Unchanged from 6.9.7 through 7.1.0.
Signature, return type and hooks compared across 3 parsed releases.
About this page
- Parsed data
- Generated from the wordpress-develop 7.1.0 tag, from
src/wp-includes/compat-utf8.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it. - Corrections
- Something wrong on this page? Report it and it gets fixed in the next regeneration.