Reference
Glossary
Plain-language definitions of the terms used across HotSymbol, from code points to Alt codes. Each entry links to its primary source.
A
- Alt code
A Windows keyboard method that enters a character by holding Alt and typing its number on the numeric keypad.
An Alt code is a Windows input method: hold the Alt key, type a decimal number on the numeric keypad, then release Alt. Numbers with a leading zero (Alt+0128 to Alt+0255) select characters from the Windows-1252 code page. Numbers without a leading zero (Alt+1 to Alt+255) select characters from the OEM code page, which is code page 437 on US English systems. Alt codes do not work with the number row, depend on the active code page, and cannot reach most Unicode characters.
Source: Unicode code page mappings (Microsoft vendor tables)
B
- Basic Multilingual Plane (BMP)
The first 65,536 Unicode code points, U+0000 to U+FFFF, where most common characters live.
Unicode divides its code space into 17 planes of 65,536 code points each. Plane 0, the Basic Multilingual Plane, covers U+0000 to U+FFFF and contains most characters for modern scripts, punctuation, and many symbols. Characters above U+FFFF, including most emoji, sit in supplementary planes and need a surrogate pair in UTF-16.
Source: Unicode Glossary
C
- Character
The smallest unit of written text with meaning, independent of how any font draws it.
In Unicode, a character is an abstract unit of text such as the letter A, the digit 7, or the rightwards arrow. Unicode assigns each character a code point and a name but not an appearance. What you see on screen is a glyph chosen by a font.
Source: Unicode Glossary
- Code page
A legacy table that maps up to 256 byte values to characters, such as Windows-1252 or code page 437.
A code page is a single-byte character encoding used before Unicode became universal. Each byte value from 0 to 255 maps to one character. Windows-1252 is the default for Western European Windows; code page 437 is the original IBM PC set with box-drawing characters and symbols such as ☺ and ♥. Windows Alt codes are numbers in these tables.
Source: Unicode code page mappings
- Code point
The unique number Unicode assigns to a character, written as U+ followed by hexadecimal digits.
A code point is a number in the Unicode code space, from U+0000 to U+10FFFF. It is written as “U+” followed by at least four hexadecimal digits; the rightwards arrow is U+2192, which is 8594 in decimal. HTML, CSS, and programming-language escapes all express the same code point in different notations.
Source: Unicode Glossary
- Combining character
A mark, such as an accent, that attaches to the character before it instead of standing alone.
A combining character is displayed together with the preceding base character. For example, U+0301 COMBINING ACUTE ACCENT after “e” renders as “é”. Some accented letters therefore have two valid encodings: one precomposed code point, or a base letter plus a combining mark. Normalization makes them comparable.
Source: Unicode Glossary
- CSS escape
A backslash followed by a hexadecimal code point, used to write a character inside CSS.
In CSS, a backslash followed by one to six hexadecimal digits represents a character by its code point, so \2192 is the rightwards arrow. It is most often used in the content property of ::before and ::after. If the next character is also a hexadecimal digit or a space, end the escape with a single space or use all six digits.
E
- Emoji presentation
Whether a character is shown as a colorful emoji or as a plain text symbol.
Some characters, such as ❤ and ☺, can be displayed either as a monochrome text symbol or as a colorful emoji. Each has a default presentation, which a variation selector can override: U+FE0E requests text style and U+FE0F requests emoji style. The final appearance also depends on the platform and font.
F
- Font fallback
When a font lacks a character, the system borrows the glyph from another installed font.
No single font covers all of Unicode. When the chosen font has no glyph for a character, browsers and operating systems substitute one from another font. If no installed font has it, a placeholder such as an empty box appears. This is why the same symbol can look different on different devices.
G
- General category
A Unicode property that classifies every character, for example as a letter, number, or symbol.
The General Category is a two-letter Unicode property assigned to every code point. The first letter gives the major class, such as L for letters, N for numbers, P for punctuation, S for symbols, and Z for separators. The second refines it: Sm is a math symbol, Sc a currency symbol, and So another symbol.
- Glyph
The visual shape a font uses to draw a character.
A glyph is the image a font draws for a character. One character can have many glyphs across fonts, and a font can combine several characters into one glyph. Unicode encodes characters, not glyphs, so copying a symbol copies the character; its look comes from whatever font displays it.
Source: Unicode Glossary
H
- Hexadecimal character reference
An HTML reference that writes the code point in hexadecimal, such as →.
A hexadecimal numeric character reference writes a character in HTML or XML as &#x, the code point in hexadecimal, and a semicolon. → is the rightwards arrow. It is equivalent to the decimal form → and matches the U+2192 notation, which makes it easy to read alongside Unicode charts.
- HTML entity
A named shortcut for a character in HTML, such as → for →.
An HTML entity, formally a named character reference, writes a character as an ampersand, a name, and a semicolon: © is © and → is →. The HTML Standard defines a fixed list of 2,125 names covering 1,511 characters, so most characters have no entity. Numeric references work for every character.
J
- JavaScript escape
A string escape such as \u2192 or \u{1F600} that writes a character in JavaScript source.
JavaScript strings accept Unicode escapes: \u followed by exactly four hexadecimal digits for characters in the Basic Multilingual Plane, or \u{…} with one to six digits for any code point. "\u2192" and "\u{2192}" both produce the rightwards arrow.
N
- Normalization
Converting text to a standard form so equivalent character sequences compare as equal.
Unicode normalization rewrites text into one of four standard forms. NFC composes characters where possible, so “e” plus a combining acute accent becomes “é”; NFD decomposes them. NFKC and NFKD also fold compatibility variants, such as superscript digits, into their plain forms, which changes meaning in some contexts.
- Numeric character reference (HTML code)
An HTML reference that writes the code point in decimal, such as →.
A numeric character reference writes a character in HTML or XML by its code point: &# followed by the decimal number and a semicolon, such as → for →. It works for every Unicode character, which makes it the most dependable way to include a symbol when the document encoding is uncertain. The hexadecimal form uses &#x.
S
- Script property
The writing system a character belongs to, such as Latin or Greek, or Common for shared symbols.
The Unicode Script property assigns each character to a writing system, such as Latin, Greek, or Devanagari. Punctuation, digits, and most symbols are used across many writing systems, so they have the value Common. Combining marks shared by several scripts have the value Inherited.
- Surrogate pair
Two UTF-16 code units that together encode one character above U+FFFF.
UTF-16 stores characters above U+FFFF as two 16-bit code units from a reserved range: a high surrogate followed by a low surrogate. This is why many emoji count as two characters in JavaScript’s string length even though they are one code point.
Source: Unicode Glossary
U
- Unicode
The universal character standard that gives every character in every writing system a unique number.
Unicode is the international standard for encoding text. It assigns each character a code point, an official name, and properties such as its category and script, so text can move between devices and programs without changing. HotSymbol’s character facts come from a pinned release of the Unicode Character Database.
Source: The Unicode Standard
- Unicode block
A named, contiguous range of code points, such as Arrows (U+2190–U+21FF).
A block is a named range of consecutive code points, always a multiple of 16 in size, such as Arrows or Mathematical Operators. Blocks organize the code charts, but they are not a classification: a block can contain unrelated characters, and related characters can be spread across several blocks.
Source: Unicode code charts
- Unicode character name
The permanent, uppercase name Unicode gives a character, such as RIGHTWARDS ARROW.
Every assigned character has an official name written in capital letters, such as RIGHTWARDS ARROW or EM DASH. Names never change once published, even when they contain mistakes; corrections are added as formal name aliases. HotSymbol may show a friendlier display name, but the official name stays on every symbol page.
- UTF-16
A Unicode encoding using 16-bit units, used internally by JavaScript, Java, and Windows.
UTF-16 encodes each code point as one or two 16-bit code units. Characters in the Basic Multilingual Plane take one unit; characters above U+FFFF take two, called a surrogate pair. JavaScript strings, Java, and the Windows API use UTF-16 internally.
Source: Unicode Glossary
- UTF-8
The dominant Unicode encoding on the web, using one to four bytes per character.
UTF-8 encodes each code point in one to four bytes and is identical to ASCII for the first 128 characters. It is the required encoding for new HTML documents. When a page is UTF-8, you can paste a symbol directly into the source instead of using an HTML code.
Source: RFC 3629: UTF-8
V
- Variation selector
An invisible code point that requests a specific visual variant of the character before it.
A variation selector follows a character to request a particular form. The most common pair is U+FE0E (text presentation) and U+FE0F (emoji presentation). Selectors are invisible, so two strings that look identical can differ by one; this matters when comparing or searching text.
Source: Unicode Glossary
Z
- Zero width joiner (ZWJ)
An invisible character, U+200D, that asks adjacent characters to join, for example into one emoji.
The zero width joiner, U+200D, has no appearance of its own. Between emoji it forms a ZWJ sequence that supporting platforms display as a single image, such as a family or a profession. Where a sequence is not supported, the individual emoji appear side by side.
Updated .