Type hello into a bold text generator, copy the bold result back into a search box, and
you get zero matches. The styled word consists of different code points than the one you
typed. That single fact explains most failures people hit with Unicode text styles, so
this article documents the exact mapping our own tool runs, including the thirteen
letters that break its arithmetic.
How the bold text generator computes each letter
Our bold text generator ships eight styles, and six of them are pure offset arithmetic.
The converter reads the ASCII code of each input character. Lowercase letters 97 through
122 map to a lowercase base plus the code minus 97. Uppercase letters 65 through 90 map
to an uppercase base plus the code minus 65. Every other character passes through
untouched.
Work the bold style as the example. Its lowercase base is U+1D41A, so a maps to U+1D41A
and z, 25 positions later, maps to U+1D433. The uppercase base is U+1D400, so A maps to
U+1D400. No lookup table exists for this style, only one addition per letter.
The eight style blocks and their bases
| Style | Uppercase base | Lowercase base | Exception letters |
| Mathematical Bold | U+1D400 | U+1D41A | none |
| Mathematical Italic | U+1D434 | U+1D44E | h |
| Bold Italic | U+1D468 | U+1D482 | none |
| Monospace | U+1D670 | U+1D68A | none |
| Double-Struck | U+1D538 | U+1D552 | C H N P Q R Z |
| Sans-Serif Bold | U+1D5D4 | U+1D5EE | none |
| Bold Script | U+1D4D0 | U+1D4EA | none |
| Fraktur | U+1D504 | U+1D51E | C H I R Z |
Eight styles times 52 letters each gives 416 mappings. Of those, 13 do not follow the
offset rule. Three styles carry exceptions, and all exceptions sit in the uppercase or
specific-letter slots described below.
The thirteen exception characters
Mathematical Italic has one exception. Lowercase h maps to U+210E, the Planck constant
symbol, because the italic block has no h at its expected slot. Any h you italicize
becomes physics notation. The uppercase H follows the normal offset.
Double-struck uppercase carries seven exceptions, and you already know these glyphs from
mathematics:
- C maps to U+2102, the complex numbers symbol
- H maps to U+210D, the quaternions symbol
- N maps to U+2115, the natural numbers symbol
- P maps to U+2119, the primes symbol
- Q maps to U+211A, the rationals symbol
- R maps to U+211D, the reals symbol
- Z maps to U+2124, the integers symbol
Fraktur uppercase carries five exceptions: C to U+212D, H to U+210C, I to U+2111, R to
U+211C, and Z to U+2128. These are the blackboard bold and letterlike symbols that
existed as standalone code points before the letter blocks were filled, so the blocks
skip their slots and the tool maps to the older characters.
What the tool never converts
The converters test only the two ASCII ranges. Digits, punctuation, spaces, and every
non-ASCII letter pass through unchanged. Type naïve 2026 and the ï and the digits come
out identical to the input. Turkish ş and ğ, German ß, accented French vowels, and
Cyrillic or Greek letters all survive conversion untouched. The output mixes styled and
plain glyphs inside one word, and the mix is invisible until you compare code points.
The input iteration itself is code point based. The tool spreads the string before
mapping, so an emoji in your input arrives at the output intact rather than split into
two broken halves.
Every styled letter is two UTF-16 code units wide
All eight block bases sit above U+FFFF, outside the Basic Multilingual Plane. UTF-16
encodes each of these characters as a surrogate pair, which means two code units. Style
the word hello in bold and a JavaScript length check reports 10, not 5. The practical
consequences follow directly:
1. Character counters that count code units double your text length.
2. Input fields that enforce length limits in code units reject styled text at half the
visible length.
3. String operations that slice by index can cut a surrogate pair in half and corrupt
the output.
The 13 exception characters sit on the Basic Multilingual Plane and take one code unit
each, so a Fraktur string with a C in it mixes one-unit and two-unit characters.
What you trade away when you paste
Styled text gives up several properties that plain text has. Search matching fails
because the code points differ from the query. Case manipulation fails because case is
baked into each code point, and a styled capital A has no lowercase partner the platform
can find. Screen readers announce mathematical alphanumerics inconsistently, and some
skip them entirely. Systems without a font that covers the mathematical alphanumeric
blocks render empty boxes instead of letters. Platforms that screen for confusable
characters in usernames sometimes reject this class of text outright.
Checklist before you post styled text
- You converted only A to Z and a to z content, and you accept every other character
stays plain.
- Your text contains no accented letters you expect to see styled.
- You checked the paste target accepts these code points at all.
- You do not need the text to match in search, on the page, or in a hashtag.
- You tested the result with one screen reader pass if accessibility matters.
- You kept the styled run short, a name or a label, because long runs tire readers.
The mapping is deterministic, so the same input always produces the same output, and you
can verify any character yourself with a code point inspector.
Count what happens with your own sample. Paste a 20 letter phrase into the
[bold text generator](https://webrecast.com/en/bold-text-generator), copy one styled
variant into your browser console, and check its length property against the visible
character count. If you find a letter that maps somewhere this table does not cover,
that finding is worth a message, because the table above comes from the shipping code.