HTML Entity Escaping Rules: Five Characters, One Order, and a Table Bug We Found

Escaping HTML is a small ruleset done in a fixed order, and most breakage

comes from changing the order or escaping too much. We read the exact

replacements in our own encoder and tested the decode path. The rules fit in

one page, and one defect in our reference table proves why you should verify

tables instead of trusting them.

The five characters and why ampersand goes first

The encoder replaces exactly five characters, in this order:

& becomes &

< becomes &lt;

> becomes &gt;

" becomes &quot;

' becomes &apos;

The ampersand runs first because it is the only character that appears in

the replacement text itself. Escape it last and you re-encode your own

output: the &amp; you just wrote would turn into &amp;amp; and render as

visible text. Any correct escaper, ours included, touches the ampersand

before it touches anything else. If you write your own escaper, that is the

one ordering rule with no exceptions.

Five characters is also the ceiling. The encoder deliberately leaves every

other character alone, which is the next rule.

What the encoder leaves alone on purpose

Accented letters, dashes, currency signs, and emoji pass through as plain

UTF-8. Input containing é or © or an emoji produces output containing those

same characters, not &#233; or &#169;. That choice keeps output readable and

small, and it is safe when your page declares UTF-8. If you paste escaped

output into a system with no character encoding declared, non-ASCII

characters are the first thing to break. The fix is to declare the encoding,

not to entity-encode the whole document.

Decode uses the browser parser, so test what you paste

Decode mode does not carry its own entity table into the logic. It creates a

detached textarea element, assigns your input to its innerHTML, and reads

the value back. The browser parser resolves named forms like &amp; and

numeric forms like &#38; in one pass. Hexadecimal numeric forms also work,

because the browser accepts them.

The tool adds a warning pass over your input. A regular expression flags any

ampersand not followed by a letter or hash and a closing semicolon. Plain

prose like AT&T triggers the warning even though browsers display it

correctly, so read the warning as check this spot, not as a guarantee of

loss. The warning fires exactly where a human should look.

The reference table has 28 rows and one verified defect

Below the editor, the tool ships a reference table of 28 characters with

named and numeric forms: ampersand through the fraction three quarters. It

is the right thing to include, and it contains a defect we verified in the

source.

The em dash row lists a plain hyphen in the character column while listing

the mdash named form and the 8212 numeric form beside it. The source line

assigns a hyphen to that row. Anyone reading the table would conclude that

the mdash entity produces the hyphen key on their keyboard. It does not. The

row above it, the en dash, displays correctly, which makes the em dash row

easier to misread, not harder.

We are stating this in the open because the table itself is the argument:

verify character tables against the specification, including ours,

especially the rows that look obvious. A wrong character column in a

reference table propagates into every document written from it.

When to use entities at all

1. In text content, escape the ampersand, the less-than sign, and the

greater-than sign. These three change parsing when raw.

2. Inside double-quoted attributes, escape the double quote. Inside

single-quoted attributes, escape the apostrophe.

3. Use numeric forms when you cannot guarantee named-form support. The

apostrophe entity &apos; comes from XML and HTML5, and older HTML 4

parsers do not recognize it. &#39; works everywhere.

4. Escape at output boundaries, at the moment text enters markup. Do not

entity-encode stored data, or every later consumer must decode before it

can search, sort, or compare.

Steps to escape a block of text

1. Paste the text into encode mode.

2. Read the output and confirm five and only five character classes

changed.

3. Paste the output into decode mode and confirm it round trips to your

source text.

4. If the malformed warning fires, inspect each flagged ampersand.

5. Insert the output into your markup and view the rendered page once.

Checklist before you publish escaped markup

  • Ampersand escaped before the other four characters.
  • Quotes escaped in attributes, and the escape matches the quoting style.
  • Page declares UTF-8 before you rely on raw non-ASCII characters.
  • Apostrophes sent as &#39; where old parsers matter.
  • A decode round trip performed once as a smoke test.

If you depend on a character table for anything serious, check its rows

against the specification the way we checked ours, and tell us what you

find. Encode, decode, and audit the 28-row reference table in the HTML

entity encoder at https://webrecast.com/en/html-entity