wppaste
WordPress

_wp_scan_utf8( string $bytes, int $at, int $invalid_length, int|null $max_bytes = null, int|null $max_code_points = null, bool|null $has_noncharacters = null ): int

Since
6.9.0
Source
wp-includes/compat-utf8.php:47
Finds spans of valid and invalid UTF-8 bytes in a given string.

Description

This is a low-level tool to power various UTF-8 functionality.
It scans through a string until it finds invalid byte spans.
When it does this, it does three things:

  • Assigns $at to the position after the last successful code point.
  • Assigns $invalid_length to the length of the maximal subpart of the invalid bytes starting at $at.
  • Returns how many code points were successfully scanned.

This information is enough to build a number of useful UTF-8 functions.

Example:

// ñ is U+F1, which in ISO-8859-1/latin1/Windows-1252/cp1252 is 0xF1.
"Pi\xF1a" === $pineapple = mb_convert_encoding( "Piña", 'Windows-1252', 'UTF-8' );
$at = $invalid_length = 0;

// The first step finds the invalid 0xF1 byte.
2 === _wp_scan_utf8( $pineapple, $at, $invalid_length );
$at === 2; $invalid_length === 1;

// The second step continues to the end of the string.
1 === _wp_scan_utf8( $pineapple, $at, $invalid_length );
$at === 4; $invalid_length === 0;

Note! While passing an options array here might be convenient from a calling-code standpoint, this function is intended to serve as a very low-level foundation upon which to build higher level functionality. For the sake of keeping costs explicit all arguments are passed directly.

Compatibility

WordPress
since 6.9.0
PHP
7.4–8.6-dev
  • 6.7.7
  • 6.8.8
  • 6.9.7
  • 7.0.4
  • 7.1.0

Present in 3 of the 5 tracked releases, added in 6.9.0, and compiles on PHP 7.4 through 8.6-dev.

Parameters

$bytesstring
UTF-8 encoded string which might include invalid spans of bytes.
$atint
Where to start scanning.
$invalid_lengthint
Will be set to how many bytes are to be ignored after $at.
$max_bytesint|nulloptional
Stop scanning after this many bytes have been seen.Default: null
$max_code_pointsint|nulloptional
Stop scanning after this many code points have been seen.Default: null
$has_noncharactersbool|nulloptional
Set to indicate if scanned string contained noncharacters.Default: null

Return value

int
How many code points were successfully scanned.

Performance profile

How much work a call to _wp_scan_utf8() does, and what it touches: the algorithmic scaling, the Zend instruction count per call across PHP versions, the hooks it hands control to, and the core code that calls it. Measured from the compiled opcodes, not a stopwatch, so every number is identical on any machine running the same PHP version, and every function in core is ranked by cost.

Cost class
Trivial

Touches nothing outside its own arguments.

Scaling
Constant

No loop in the body: the same number of instructions runs whatever you pass in.

Instructions
22–121

Executed per call on PHP 8.5, depending on the branch taken. The body compiles to 278.

Plugin surface
None

Nothing here hands control to plugin code.

Called by
7

7 places in core call this, so the cost is paid more often than your own code shows.

What one call costs · 3 distinct outcomes

One number would be a lie: the work depends on which branch runs. These are every distinct cost _wp_scan_utf8() can have, taken from its control-flow graph on PHP 8.5.

WhenInstructionsCalls it makes
!$i22none
$i && $count && !$max_count && !$end && !$bytes && !$b1 && $b364–67strspn(), ord(), ord(), ord()
$i && $count && !$max_count && !$end && !$bytes && !$b1 && $b385–121strspn(), ord(), ord(), ord(), ord(), ord()

This body has more branch combinations than are worth enumerating, so the table covers the outcomes found first rather than every one that exists.

Across PHP versions

PHPCompiledExecutedBranchesNotes
8.6-dev27822–12182
8.527822–12182
8.427822–1218213 fewer instructions than PHP 8.3
8.329126–13182
8.229126–13182
8.129126–131821 fewer instruction than PHP 7.4
7.429226–13182

An instruction is not a fixed amount of time, so a matching count is not necessarily the same speed; what it rules out is a difference in the work itself.

Used by · 7

Source code

function _wp_scan_utf8( string $bytes, int &$at, int &$invalid_length, ?int $max_bytes = null, ?int $max_code_points = null, ?bool &$has_noncharacters = null ): int {	$byte_length       = strlen( $bytes );	$end               = min( $byte_length, $at + ( $max_bytes ?? PHP_INT_MAX ) );	$invalid_length    = 0;	$count             = 0;	$max_count         = $max_code_points ?? PHP_INT_MAX;	$has_noncharacters = false; 	for ( $i = $at; $i < $end && $count <= $max_count; $i++ ) {		/*		 * Quickly skip past US-ASCII bytes, all of which are valid UTF-8.		 *		 * This optimization step improves the speed from 10x to 100x		 * depending on whether the JIT has optimized the function.		 */		$ascii_byte_count = strspn(			$bytes,			"\x00\x01\x02\x03\x04\x05\x06\x07\x08\x09\x0a\x0b\x0c\x0d\x0e\x0f" .			"\x10\x11\x12\x13\x14\x15\x16\x17\x18\x19\x1a\x1b\x1c\x1d\x1e\x1f" .			" !\"#$%&'()*+,-./0123456789:;<=>?@ABCDEFGHIJKLMNOPQRSTUVWXYZ[\\]^_`abcdefghijklmnopqrstuvwxyz{|}~\x7f",			$i,			min( $end - $i, $max_count - $count )		); 		if ( $count + $ascii_byte_count >= $max_count ) {			$at    = $i + ( $max_count - $count );			$count = $max_count;			return $count;		} 		$count += $ascii_byte_count;		$i     += $ascii_byte_count; 		if ( $i >= $end ) {			$at = $end;			return $count;		} 		/**		 * The above fast-track handled all single-byte UTF-8 characters. What		 * follows MUST be a multibyte sequence otherwise there’s invalid UTF-8.		 *		 * Therefore everything past here is checking those multibyte sequences.		 *		 * It may look like there’s a need to check against the max bytes here,		 * but since each match of a single character returns, this functions will		 * bail already if crossing the max-bytes threshold. This function SHALL		 * NOT return in the middle of a multi-byte character, so if a character		 * falls on each side of the max bytes, the entire character will be scanned.		 *		 * Because it’s possible that there are truncated characters, the use of		 * the null-coalescing operator with "\xC0" is a convenience for skipping		 * length checks on every continuation bytes. This works because 0xC0 is		 * always invalid in a UTF-8 string, meaning that if the string has been		 * truncated, it will find 0xC0 and reject as invalid UTF-8.		 *		 * > [The following table] lists all of the byte sequences that are well-formed		 * > in UTF-8. A range of byte values such as A0..BF indicates that any byte		 * > from A0 to BF (inclusive) is well-formed in that position. Any byte value		 * > outside of the ranges listed is ill-formed.		 *		 * > Table 3-7. Well-Formed UTF-8 Byte Sequences		 *  ╭─────────────────────┬────────────┬──────────────┬─────────────┬──────────────╮		 *  │ Code Points         │ First Byte │ Second Byte  │ Third Byte  │ Fourth Byte  │		 *  ├─────────────────────┼────────────┼──────────────┼─────────────┼──────────────┤		 *  │ U+0000..U+007F      │ 00..7F     │              │             │              │		 *  │ U+0080..U+07FF      │ C2..DF     │ 80..BF       │             │              │		 *  │ U+0800..U+0FFF      │ E0         │ A0..BF       │ 80..BF      │              │		 *  │ U+1000..U+CFFF      │ E1..EC     │ 80..BF       │ 80..BF      │              │		 *  │ U+D000..U+D7FF      │ ED         │ 80..9F       │ 80..BF      │              │		 *  │ U+E000..U+FFFF      │ EE..EF     │ 80..BF       │ 80..BF      │              │		 *  │ U+10000..U+3FFFF    │ F0         │ 90..BF       │ 80..BF      │ 80..BF       │		 *  │ U+40000..U+FFFFF    │ F1..F3     │ 80..BF       │ 80..BF      │ 80..BF       │		 *  │ U+100000..U+10FFFF  │ F4         │ 80..8F       │ 80..BF      │ 80..BF       │		 *  ╰─────────────────────┴────────────┴──────────────┴─────────────┴──────────────╯		 *		 * @see https://www.unicode.org/versions/Unicode16.0.0/core-spec/chapter-3/#G27506		 */ 		// Valid two-byte code points.

Changelog

Introduced in 6.9.0. Unchanged from 6.9.7 through 7.1.0.

  1. 6.9.7
  2. 7.0.4
  3. 7.1.0

Signature, return type and hooks compared across 3 parsed releases.

About this page

Parsed data
Generated from the wordpress-develop 7.1.0 tag, from src/wp-includes/compat-utf8.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it.
Corrections
Something wrong on this page? Report it and it gets fixed in the next regeneration.