Escaping HTML is a small ruleset done in a fixed order, and most breakage
comes from changing the order or escaping too much. We read the exact
replacements in our own encoder and tested the decode path. The rules fit in
one page, and one defect in our reference table proves why you should verify
tables instead of trusting them.
The five characters and why ampersand goes first
The encoder replaces exactly five characters, in this order:
& becomes &
< becomes <
> becomes >
" becomes "
' becomes '
The ampersand runs first because it is the only character that appears in
the replacement text itself. Escape it last and you re-encode your own
output: the & you just wrote would turn into &amp; and render as
visible text. Any correct escaper, ours included, touches the ampersand
before it touches anything else. If you write your own escaper, that is the
one ordering rule with no exceptions.
Five characters is also the ceiling. The encoder deliberately leaves every
other character alone, which is the next rule.
What the encoder leaves alone on purpose
Accented letters, dashes, currency signs, and emoji pass through as plain
UTF-8. Input containing é or © or an emoji produces output containing those
same characters, not é or ©. That choice keeps output readable and
small, and it is safe when your page declares UTF-8. If you paste escaped
output into a system with no character encoding declared, non-ASCII
characters are the first thing to break. The fix is to declare the encoding,
not to entity-encode the whole document.
Decode uses the browser parser, so test what you paste
Decode mode does not carry its own entity table into the logic. It creates a
detached textarea element, assigns your input to its innerHTML, and reads
the value back. The browser parser resolves named forms like & and
numeric forms like & in one pass. Hexadecimal numeric forms also work,
because the browser accepts them.
The tool adds a warning pass over your input. A regular expression flags any
ampersand not followed by a letter or hash and a closing semicolon. Plain
prose like AT&T triggers the warning even though browsers display it
correctly, so read the warning as check this spot, not as a guarantee of
loss. The warning fires exactly where a human should look.
The reference table has 28 rows and one verified defect
Below the editor, the tool ships a reference table of 28 characters with
named and numeric forms: ampersand through the fraction three quarters. It
is the right thing to include, and it contains a defect we verified in the
source.
The em dash row lists a plain hyphen in the character column while listing
the mdash named form and the 8212 numeric form beside it. The source line
assigns a hyphen to that row. Anyone reading the table would conclude that
the mdash entity produces the hyphen key on their keyboard. It does not. The
row above it, the en dash, displays correctly, which makes the em dash row
easier to misread, not harder.
We are stating this in the open because the table itself is the argument:
verify character tables against the specification, including ours,
especially the rows that look obvious. A wrong character column in a
reference table propagates into every document written from it.
When to use entities at all
1. In text content, escape the ampersand, the less-than sign, and the
greater-than sign. These three change parsing when raw.
2. Inside double-quoted attributes, escape the double quote. Inside
single-quoted attributes, escape the apostrophe.
3. Use numeric forms when you cannot guarantee named-form support. The
apostrophe entity ' comes from XML and HTML5, and older HTML 4
parsers do not recognize it. ' works everywhere.
4. Escape at output boundaries, at the moment text enters markup. Do not
entity-encode stored data, or every later consumer must decode before it
can search, sort, or compare.
Steps to escape a block of text
1. Paste the text into encode mode.
2. Read the output and confirm five and only five character classes
changed.
3. Paste the output into decode mode and confirm it round trips to your
source text.
4. If the malformed warning fires, inspect each flagged ampersand.
5. Insert the output into your markup and view the rendered page once.
Checklist before you publish escaped markup
- Ampersand escaped before the other four characters.
- Quotes escaped in attributes, and the escape matches the quoting style.
- Page declares UTF-8 before you rely on raw non-ASCII characters.
- Apostrophes sent as ' where old parsers matter.
- A decode round trip performed once as a smoke test.
If you depend on a character table for anything serious, check its rows
against the specification the way we checked ours, and tell us what you
find. Encode, decode, and audit the 28-row reference table in the HTML
entity encoder at https://webrecast.com/en/html-entity