WP_HTML_Decoder::read_character_reference( string $context, string $text, int $at = 0, $match_byte_length = null ): ?string
- Since
- 6.6.0
- Source
wp-includes/html-api/class-wp-html-decoder.php:258
Description
If a character reference is found, this function will return the translated value that the reference maps to. It will then set $match_byte_length the number of bytes of input it read while consuming the character reference. This gives calling code the opportunity to advance its cursor when traversing a string and decoding.
Example:
null === WP_HTML_Decoder::read_character_reference( 'attribute', 'Ships…', 0 );
'…' === WP_HTML_Decoder::read_character_reference( 'attribute', 'Ships…', 5, $token_length );
8 === $token_length; // …
null === WP_HTML_Decoder::read_character_reference( 'attribute', '¬in', 0 );
'∉' === WP_HTML_Decoder::read_character_reference( 'attribute', '∉', 0, $token_length );
7 === $token_length; // ∉
'¬' === WP_HTML_Decoder::read_character_reference( 'data', '¬in', 0, $token_length );
4 === $token_length; // ¬
'∉' === WP_HTML_Decoder::read_character_reference( 'data', '∉', 0, $token_length );
7 === $token_length; // ∉Compatibility
- WordPress
- since 6.6.0
- PHP
- 7.4–8.6-dev
- 6.7.7
- 6.8.8
- 6.9.7
- 7.0.4
- 7.1.0
Present in every tracked release (6.7.7 to 7.1.0), and compiles on PHP 7.4 through 8.6-dev.
Parameters
$contextstringattributefor decoding attribute values,dataotherwise.$textstring- Text document containing span of text to decode.
$atintoptional- Byte offset into text where span begins, defaults to the beginning (0).Default:
0 $match_byte_lengthoptional- Default:
null
Return value
?string- Decoded character reference in UTF-8 if found, otherwise null.
Performance profile
How much work a call to WP_HTML_Decoder::read_character_reference() does, and what it touches: the algorithmic scaling, the Zend instruction count per call across PHP versions, the hooks it hands control to, and the core code that calls it. Measured from the compiled opcodes, not a stopwatch, so every number is identical on any machine running the same PHP version, and every function in core is ranked by cost.
- Cost class
- Trivial
- Scaling
- Constant
- Instructions
- 10–80
- Plugin surface
- None
- Called by
- 3
Touches nothing outside its own arguments.
No loop in the body: the same number of instructions runs whatever you pass in.
Executed per call on PHP 8.5, depending on the branch taken. The body compiles to 143.
Nothing here hands control to plugin code.
3 places in core call this, so the cost is paid more often than your own code shows.
What one call costs · 5 distinct outcomes
One number would be a lie: the work depends on which branch runs. These are every distinct cost WP_HTML_Decoder::read_character_reference() can have, taken from its control-flow graph on PHP 8.5.
| When | Instructions | Calls it makes |
|---|---|---|
| always | 10–21 | none |
!$length | 30–41 | ->read_token() |
!$length && $replacement !== null && $context === "attribute" | 45–59 | ->read_token(), ord() |
!$length | 52–64 | strspn(), strspn() |
!$length && $digit_count !== 0 && !$max_digits | 68–80 | strspn(), strspn(), intval(), ::code_point_to_utf8_bytes() |
Across PHP versions
| PHP | Compiled | Executed | Branches | Notes |
|---|---|---|---|---|
| 8.6-dev | 143 | 10–80 | 26 | |
| 8.5 | 143 | 10–80 | 26 | |
| 8.4 | 143 | 10–80 | 26 | 4 fewer instructions than PHP 8.3 |
| 8.3 | 147 | 10–84 | 26 | |
| 8.2 | 147 | 10–84 | 26 | |
| 8.1 | 147 | 10–84 | 26 | 1 fewer instruction than PHP 7.4 |
| 7.4 | 148 | 10–85 | 26 |
An instruction is not a fixed amount of time, so a matching count is not necessarily the same speed; what it rules out is a difference in the work itself.
Uses · 1
- WP_HTML_Decoder::code_point_to_utf8_bytes()Encode a code point number into the UTF-8 encoding.
Used by · 3
- WP_HTML_Decoder::attribute_starts_with()Indicates if an attribute value starts with a given raw string value.
- WP_HTML_Decoder::decode()Decodes a span of HTML text, depending on the context in which it's found.
- WP_HTML_Tag_Processor::subdivide_text_appropriately()Subdivides a matched text node, splitting NULL byte sequences and decoded whitespace as distinct nodes prefixes.
Changelog
Introduced in 6.6.0. One change between 6.7.7 and 7.1.0.
Signature, return type and hooks compared across 5 parsed releases.
string|false to ?string.verified against sourceAbout this page
- Parsed data
- Generated from the wordpress-develop 7.1.0 tag, from
src/wp-includes/html-api/class-wp-html-decoder.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it. - Corrections
- Something wrong on this page? Report it and it gets fixed in the next regeneration.