Use this class in specific circumstances with a static set of lookup keys which map to a static set of transformed values. For example, this class is used to map HTML named character references to their equivalent UTF-8 values. This class works differently than code calling in_array() and other methods. It internalizes lookup logic and provides helper interfaces to optimize lookup and transformation. It provides a method for precomputing the lookup tables and storing them as PHP source code. All tokens and substitutions must be shorter than 256 bytes. Example: $smilies = WP_Token_Map::from_array( array(
'8O' => '😯',
':(' => '🙁',
':)' => '🙂',
':?' => '😕',
) );
true === $smilies->contains( ':)' );
false === $smilies->contains( 'simile' );
'😕' === $smilies->read_token( 'Not sure :?.', 9, $length_of_smily_syntax );
2 === $length_of_smily_syntax; <h2>Precomputing the Token Map.</h2> Creating the class involves some work sorting and organizing the tokens and their replacement values. In order to skip this, it's possible for the class to export its state and be used as actual PHP source code. Example: // Export with four spaces as the indent, only for the sake of this docblock.
// The default indent is a tab character.
$indent = ' ';
echo $smilies->precomputed_php_source_table( $indent );
// Output, to be pasted into a PHP source file:
WP_Token_Map::from_precomputed_table(
array(
"storage_version" => "6.6.0",
"key_length" => 2,
"groups" => "",
"long_words" => array(),
"small_words" => "8O\x00:)\x00:(\x00:?\x00",
"small_mappings" => array( "😯", "🙂", "🙁", "😕" )
)
); <h2>Large vs. small words.</h2> This class uses a short prefix called the "key" to optimize lookup of its tokens.This means that some tokens may be shorter than or equal in length to that key.Those words that are longer than the key are called "large" while those shorter than or equal to the key length are called "small." This separation of large and small words is incidental to the way this class optimizes lookup, and should be considered an internal implementation detail of the class. It may still be important to be aware of it, however. <h2>Determining Key Length.</h2> The choice of the size of the key length should be based on the data being stored in the token map. It should divide the data as evenly as possible, but should not create so many groups that a large fraction of the groups only contain a single token. For the HTML5 named character references, a key length of 2 was found to provide a sufficient spread and should be a good default for relatively large sets of tokens. However, for some data sets this might be too long. For example, a list of smilies may be too small for a key length of 2. Perhaps 1 would be more appropriate. It's best to experiment and determine empirically which values are appropriate. <h2>Generate Pre-Computed Source Code.</h2> Since the WP_Token_Map is designed for relatively static lookups, it can be advantageous to precompute the values and instantiate a table that has already sorted and grouped the tokens and built the lookup strings. This can be done with WP_Token_Map::precomputed_php_source_table(). Note that if there is a leading character that all tokens need, such as & for HTML named character references, it can be beneficial to exclude this from the token map. Instead, find occurrences of the leading character and then use the token map to see if the following characters complete the token. Example: $map = WP_Token_Map::from_array( array( 'simple_smile:' => '🙂', 'sob:' => '😭', 'soba:' => '🍜' ) );
echo $map->precomputed_php_source_table();
// Output
WP_Token_Map::from_precomputed_table(
array(
"storage_version" => "6.6.0",
"key_length" => 2,
"groups" => "si\x00so\x00",
"long_words" => array(
// simple_smile:[🙂].
"\x0bmple_smile:\x04🙂",
// soba:[🍜] sob:[😭].
"\x03ba:\x04🍜\x02b:\x04😭",
),
"short_words" => "",
"short_mappings" => array()
}
); This precomputed value can be stored directly in source code and will skip the startup cost of generating the lookup strings. See $html5_named_character_entities. Note that any updates to the precomputed format should update the storage version constant. It would also be best to provide an update function to take older known versions and upgrade them in place when loading into from_precomputed_table(). <h2>Future Direction.</h2> It may be viable to dynamically increase the length limits such that there's no need to impose them.The limit appears because of the packing structure, which indicates how many bytes each segment of text in the lookup tables spans. If, however, care were taken to track the longest word length, then the packing structure could change its representation to allow for that. Each additional byte storing length, however, increases the memory overhead and lookup runtime. An alternative approach could be to borrow the UTF-8 variable-length encoding and store lengths of less than 127 as a single byte with the high bit unset, storing longer lengths as the combination of continuation bytes. Since it has not been shown during the development of this class that longer strings are required, this update is deferred until such a need is clear.
Properties · 5
$key_lengthintprivate
How many bytes of each key are used to form a group key for lookup.
$large_wordsarrayprivate
Stores an optimized form of the word set, where words are grouped by a prefix of the $key_length and then collapsed into a string.
$groupsstringprivate
Stores the group keys for sequential string lookup.
$small_wordsstringprivate
Stores an optimized row of small words, where every entry is $this->key_size + 1 bytes long and zero-extended.
$small_mappingsstring[]private
Replacements for the small words, in the same order they appear.
Methods · 8
from_array()Create a token map using an associative array of key/value pairs as the input.
146classWP_Token_Map{147/**148 * Denotes the version of the code which produces pre-computed source tables.149 *150 * This version will be used not only to verify pre-computed data, but also151 * to upgrade pre-computed data from older versions. Choosing a name that152 * corresponds to the WordPress release will help people identify where an153 * old copy of data came from.154 */155constSTORAGE_VERSION='6.6.0-trunk';156157/**158 * Maximum length for each key and each transformed value in the table (in bytes).159 *160 * @since 6.6.0161 */162constMAX_LENGTH=256;163164/**165 * How many bytes of each key are used to form a group key for lookup.166 * This also determines whether a word is considered short or long.167 *168 * @since 6.6.0169 *170 * @var int171 */172private$key_length=2;173174/**175 * Stores an optimized form of the word set, where words are grouped176 * by a prefix of the `$key_length` and then collapsed into a string.177 *178 * In each group, the keys and lookups form a packed data structure.179 * The keys in the string are stripped of their "group key," which is180 * the prefix of length `$this->key_length` shared by all of the items181 * in the group. Each word in the string is prefixed by a single byte182 * whose raw unsigned integer value represents how many bytes follow.183 *184 * ┌────────────────┬───────────────┬─────────────────┬────────┐185 * │ Length of rest │ Rest of key │ Length of value │ Value │186 * │ of key (bytes) │ │ (bytes) │ │187 * ├────────────────┼───────────────┼─────────────────┼────────┤188 * │ 0x08 │ nterDot; │ 0x02 │ · │189 * └────────────────┴───────────────┴─────────────────┴────────┘190 *191 * In this example, the key `CenterDot;` has a group key `Ce`, leaving192 * eight bytes for the rest of the key, `nterDot;`, and two bytes for193 * the transformed value `·` (or U+B7 or "\xC2\xB7").194 *195 * Example:196 *197 * // Stores array( 'CenterDot;' => '·', 'Cedilla;' => '¸' ).198 * $groups = "Ce\x00";199 * $large_words = array( "\x08nterDot;\x02·\x06dilla;\x02¸" )200 *201 * The prefixes appear in the `$groups` string, each followed by a null202 * byte. This makes for quick lookup of where in the group string the key203 * is found, and then a simple division converts that offset into the index204 * in the `$large_words` array where the group string is to be found.205 *206 * This lookup data structure is designed to optimize cache locality and207 * minimize indirect memory reads when matching strings in the set.208 *209 * @since 6.6.0210 *211 * @var array212 */213private$large_words=array();214215/**216 * Stores the group keys for sequential string lookup.217 *218 * The offset into this string where the group key appears corresponds with the index219 * into the group array where the rest of the group string appears. This is an optimization220 * to improve cache locality while searching and minimize indirect memory accesses.221 *222 * @since 6.6.0223 *224 * @var string225 */
History
Introduced in 6.6.0. Unchanged from 6.7.7 through 7.1.0.
Signature, return type and hooks compared across 5 parsed releases.
About this page
Parsed data
Generated from the wordpress-develop 6.9.7 tag, from src/wp-includes/class-wp-token-map.php, and regenerated for each WordPress release so it tracks the code rather than a snapshot of it.
Corrections
Something wrong on this page? Report it and it gets fixed in the next regeneration.