Bold Text Generator Unicode Math: Eight Style Blocks and Thirteen Exceptions

Type hello into a bold text generator, copy the bold result back into a search box, and

you get zero matches. The styled word consists of different code points than the one you

typed. That single fact explains most failures people hit with Unicode text styles, so

this article documents the exact mapping our own tool runs, including the thirteen

letters that break its arithmetic.

How the bold text generator computes each letter

Our bold text generator ships eight styles, and six of them are pure offset arithmetic.

The converter reads the ASCII code of each input character. Lowercase letters 97 through

122 map to a lowercase base plus the code minus 97. Uppercase letters 65 through 90 map

to an uppercase base plus the code minus 65. Every other character passes through

untouched.

Work the bold style as the example. Its lowercase base is U+1D41A, so a maps to U+1D41A

and z, 25 positions later, maps to U+1D433. The uppercase base is U+1D400, so A maps to

U+1D400. No lookup table exists for this style, only one addition per letter.

The eight style blocks and their bases

StyleUppercase baseLowercase baseException letters
Mathematical BoldU+1D400U+1D41Anone
Mathematical ItalicU+1D434U+1D44Eh
Bold ItalicU+1D468U+1D482none
MonospaceU+1D670U+1D68Anone
Double-StruckU+1D538U+1D552C H N P Q R Z
Sans-Serif BoldU+1D5D4U+1D5EEnone
Bold ScriptU+1D4D0U+1D4EAnone
FrakturU+1D504U+1D51EC H I R Z

Eight styles times 52 letters each gives 416 mappings. Of those, 13 do not follow the

offset rule. Three styles carry exceptions, and all exceptions sit in the uppercase or

specific-letter slots described below.

The thirteen exception characters

Mathematical Italic has one exception. Lowercase h maps to U+210E, the Planck constant

symbol, because the italic block has no h at its expected slot. Any h you italicize

becomes physics notation. The uppercase H follows the normal offset.

Double-struck uppercase carries seven exceptions, and you already know these glyphs from

mathematics:

  • C maps to U+2102, the complex numbers symbol
  • H maps to U+210D, the quaternions symbol
  • N maps to U+2115, the natural numbers symbol
  • P maps to U+2119, the primes symbol
  • Q maps to U+211A, the rationals symbol
  • R maps to U+211D, the reals symbol
  • Z maps to U+2124, the integers symbol

Fraktur uppercase carries five exceptions: C to U+212D, H to U+210C, I to U+2111, R to

U+211C, and Z to U+2128. These are the blackboard bold and letterlike symbols that

existed as standalone code points before the letter blocks were filled, so the blocks

skip their slots and the tool maps to the older characters.

What the tool never converts

The converters test only the two ASCII ranges. Digits, punctuation, spaces, and every

non-ASCII letter pass through unchanged. Type naïve 2026 and the ï and the digits come

out identical to the input. Turkish ş and ğ, German ß, accented French vowels, and

Cyrillic or Greek letters all survive conversion untouched. The output mixes styled and

plain glyphs inside one word, and the mix is invisible until you compare code points.

The input iteration itself is code point based. The tool spreads the string before

mapping, so an emoji in your input arrives at the output intact rather than split into

two broken halves.

Every styled letter is two UTF-16 code units wide

All eight block bases sit above U+FFFF, outside the Basic Multilingual Plane. UTF-16

encodes each of these characters as a surrogate pair, which means two code units. Style

the word hello in bold and a JavaScript length check reports 10, not 5. The practical

consequences follow directly:

1. Character counters that count code units double your text length.

2. Input fields that enforce length limits in code units reject styled text at half the

visible length.

3. String operations that slice by index can cut a surrogate pair in half and corrupt

the output.

The 13 exception characters sit on the Basic Multilingual Plane and take one code unit

each, so a Fraktur string with a C in it mixes one-unit and two-unit characters.

What you trade away when you paste

Styled text gives up several properties that plain text has. Search matching fails

because the code points differ from the query. Case manipulation fails because case is

baked into each code point, and a styled capital A has no lowercase partner the platform

can find. Screen readers announce mathematical alphanumerics inconsistently, and some

skip them entirely. Systems without a font that covers the mathematical alphanumeric

blocks render empty boxes instead of letters. Platforms that screen for confusable

characters in usernames sometimes reject this class of text outright.

Checklist before you post styled text

  • You converted only A to Z and a to z content, and you accept every other character

stays plain.

  • Your text contains no accented letters you expect to see styled.
  • You checked the paste target accepts these code points at all.
  • You do not need the text to match in search, on the page, or in a hashtag.
  • You tested the result with one screen reader pass if accessibility matters.
  • You kept the styled run short, a name or a label, because long runs tire readers.

The mapping is deterministic, so the same input always produces the same output, and you

can verify any character yourself with a code point inspector.

Count what happens with your own sample. Paste a 20 letter phrase into the

[bold text generator](https://webrecast.com/en/bold-text-generator), copy one styled

variant into your browser console, and check its length property against the visible

character count. If you find a letter that maps somewhere this table does not cover,

that finding is worth a message, because the table above comes from the shipping code.