wp_kses_normalize_entities( string $content, string $context = 'html' ): string
- Since
- 1.0.0, 5.5.0
- Source
wp-includes/kses.php:2089
Description
This function normalizes HTML entities. It will convert AT&T to the correct AT&T, : to :, &#XYZZY; to &#XYZZY; and so on.
When $context is set to 'xml', HTML entities are converted to their code points. For example, AT&T…&#XYZZY; is converted to AT&T…&#XYZZY;.
Compatibility
- WordPress
- since 5.5.0
- PHP
- 7.4–8.6-dev
- 6.7.7
- 6.8.8
- 6.9.7
- 7.0.4
- 7.1.0
Present in every tracked release (6.7.7 to 7.1.0), and compiles on PHP 7.4 through 8.6-dev.
Parameters
$contentstring- Content to normalize entities.
$contextstringoptional- Context for normalization. Can be either 'html' or 'xml'.
Default 'html'.Default:'html'
Return value
string- Content with normalized entities.
Performance profile
How much work a call to wp_kses_normalize_entities() does, and what it touches: the algorithmic scaling, the Zend instruction count per call across PHP versions, the hooks it hands control to, and the core code that calls it. Measured from the compiled opcodes, not a stopwatch, so every number is identical on any machine running the same PHP version, and every function in core is ranked by cost.
- Cost class
- Trivial
- Scaling
- Constant
- Instructions
- 26
- Plugin surface
- None
- Called by
- 3
Touches nothing outside its own arguments.
No loop in the body: the same number of instructions runs whatever you pass in.
Executed per call on PHP 8.5. The body compiles to 33.
Nothing here hands control to plugin code.
3 places in core call this, so the cost is paid more often than your own code shows.
What it touches
- regexregular expression over the whole input
preg_replace_callback()called directly
What one call costs · 1 distinct outcome
One number would be a lie: the work depends on which branch runs. These are every distinct cost wp_kses_normalize_entities() can have, taken from its control-flow graph on PHP 8.5.
| When | Instructions | Calls it makes |
|---|---|---|
| always | 26 | preg_replace_callback(), preg_replace_callback(), preg_replace_callback() |
Across PHP versions
| PHP | Compiled | Executed | Branches | Notes |
|---|---|---|---|---|
| 8.6-dev | 33 | 26 | 1 | |
| 8.5 | 33 | 26 | 1 | |
| 8.4 | 33 | 26 | 1 | 3 fewer instructions than PHP 8.3 |
| 8.3 | 36 | 29 | 1 | |
| 8.2 | 36 | 29 | 1 | |
| 8.1 | 36 | 29 | 1 | |
| 7.4 | 36 | 29 | 1 |
An instruction is not a fixed amount of time, so a matching count is not necessarily the same speed; what it rules out is a difference in the work itself.
Used by · 3
- _wp_specialchars()Converts a number of special characters into their HTML entities.
- esc_url()Checks and cleans a URL.
- wp_kses()Filters text content and strips out disallowed HTML.
Source code
function wp_kses_normalize_entities( $content, $context = 'html' ) { // Disarm all entities by converting & to & $content = str_replace( '&', '&', $content ); /* * Decode any character references that are now double-encoded. * * It's important that the following normalizations happen in the correct order. * * At this point, all `&` have been transformed to `&`. Double-encoded named character * references like `&` will be decoded back to their single-encoded form `&`. * * First, numeric (decimal and hexadecimal) character references must be handled so that * `	` becomes `	`. If the named character references were handled first, there * would be no way to know whether the double-encoded character reference had been produced * in this function or was the original input. * * Consider the two examples, first with named entity decoding followed by numeric * entity decoding. We'll use U+002E FULL STOP (.) in our example, this table follows the * string processing from left to right: * * | Input | &-encoded | Named ref double-decoded | Numeric ref double-decoded | * | ------------ | ---------------- | ------------------------- | -------------------------- | * | `.` | `.` | `.` | `.` | * | `.` | `.` | `.` | `.` | * * Notice in the example above that different inputs result in the same result. The second case * was not normalized and produced HTML that is semantically different from the input. * * | Input | &-encoded | Numeric ref double-decoded | Named ref double-decoded | * | ------------ | ---------------- | --------------------------- | ------------------------ | * | `.` | `.` | `.` | `.` | * | `.` | `.` | `.` | `.` | * * Here, each input is normalized to an appropriate output. */ $content = preg_replace_callback( '/&#(0*[1-9][0-9]{0,6});/', 'wp_kses_normalize_entities2', $content ); $content = preg_replace_callback( '/&#[Xx](0*[1-9A-Fa-f][0-9A-Fa-f]{0,5});/', 'wp_kses_normalize_entities3', $content ); if ( 'xml' === $context ) { $content = preg_replace_callback( '/&([A-Za-z]{2,8}[0-9]{0,2});/', 'wp_kses_xml_named_entities', $content ); } else { $content = preg_replace_callback( '/&([A-Za-z]{2,8}[0-9]{0,2});/', 'wp_kses_named_entities', $content ); } return $content;}Changelog
Introduced in 1.0.0. Unchanged from 6.7.7 through 7.1.0.
Signature, return type and hooks compared across 5 parsed releases.
$context parameter.from the docblockAbout this page
- Parsed data
- Generated from the wordpress-develop 7.0.4 tag, from
src/wp-includes/kses.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it. - Corrections
- Something wrong on this page? Report it and it gets fixed in the next regeneration.