Every web security talk eventually lands on the same slide: always escape user input before inserting it into HTML. The reason is HTML entities — the mechanism that turns <script> into <script> and prevents a user’s comment from becoming an XSS attack.
But escaping is only half the story. Unescaping — turning & back into & — is equally important and equally dangerous when done wrong. This post covers both directions, the characters that matter, and the mistakes that lead to real vulnerabilities.
The five characters that matter
HTML has five characters that are structurally significant — they delimit tags, attributes, and entities. If any of these appear in user content and aren’t escaped, the browser interprets them as HTML markup:
| Character | Entity | Why it matters |
|---|---|---|
< |
< |
Starts a tag. <script> is a tag. |
> |
> |
Ends a tag. |
& |
& |
Starts an entity. < is an entity. |
" |
" |
Delimits attribute values. |
' |
' |
Delimits single-quoted attributes. |
If you’re inserting text into an HTML element’s content, you need to escape <, >, and &. If you’re inserting text into an attribute value, you also need to escape " (or ' if the attribute uses single quotes).
How entities work
An HTML entity is a special sequence that the browser interprets as a single character:
- Named entities:
&,<,>,", — easy to read, limited set. - Numeric entities:
<(decimal),<(hex) — any Unicode character by code point. - Named entities for symbols:
©(©),€(€),—(—) — convenience names.
The & entity is special because it’s the escape for & itself. If you escape & first, then escape < and >, you don’t double-escape. The order matters:
Correct: & → & → &lt; (correct)
Wrong: < → < → &lt; (double-escaped)
Always escape & first. Then the other characters. Then you won’t double-escape.
Where escaping is required
User-generated content. Comments, forum posts, reviews, messages — any text that comes from a user and appears on a page must be escaped. A comment containing <img src=x onerror=alert(1)> is an XSS attack if you don’t escape the < and >.
Data in attributes. If you put user data in an href, src, alt, or title attribute, escape the entity characters. A value like foo" onclick="alert(1) breaks out of the attribute and injects an event handler.
JavaScript string interpolation. If you’re building HTML inside JavaScript:
// DANGEROUS: raw interpolation
element.innerHTML = `<p>${userComment}</p>`;
// SAFE: escaped interpolation
element.innerHTML = `<p>${escapeHtml(userComment)}</p>`;
innerHTML interprets the string as HTML. If userComment contains markup, it executes. Always escape before interpolating into HTML.
Server-side rendering. Template engines (Handlebars, Jinja, EJS, Razor) usually auto-escape by default. But if you use a “safe” or “raw” filter (like Handlebars’ triple-brace {{{ or Django’s |safe), you’re opting out of escaping. Only do that when you control the content.
When to unescape
Unescaping reverses the process: < → <, & → &. It’s needed when:
- Displaying encoded data. If an API returns escaped HTML (some do, to prevent rendering), you need to unescape before displaying.
- Processing form data. Form submissions may contain URL-encoded or entity-encoded values.
- Parsing user input. If a user pastes HTML-encoded text and you want to display it as-is, unescape first — then re-escape if you’re inserting it back into HTML.
The danger: unescape only when you’re about to display, never when you’re about to store. Store the original, escaped form. Unescape at render time.
XSS: one escaped character away
Cross-site scripting (XSS) happens when unescaped user input reaches the browser’s HTML parser. The most common variants:
Stored XSS. A user submits <script>steal(document.cookie)</script> as a comment. The server stores it unescaped. Every visitor who views the comment executes the script.
Reflected XSS. A URL contains ?q=<script>alert(1)</script>. The server reflects the q parameter into the page without escaping. Anyone who clicks the link executes the script.
DOM-based XSS. JavaScript reads location.hash and inserts it into the DOM via innerHTML. The hash contains markup. The browser renders it.
All three are prevented by the same thing: escape <, >, &, ", and ' before inserting content into HTML.
The escaping function
Every language has one. The pattern is the same:
function escapeHtml(str) {
return str
.replace(/&/g, '&')
.replace(/</g, '<')
.replace(/>/g, '>')
.replace(/"/g, '"')
.replace(/'/g, ''');
}
The order matters: & first, then the others. If you escape < before &, you get &lt; instead of <.
For high-throughput systems, use a library. DOMPurify sanitizes HTML (allowing safe tags, stripping dangerous ones). The Web platform’s built-in DOMParser and textContent also handle escaping implicitly.
Common mistakes
- Escaping
&after<— causes double-escaping:&lt;instead of<. - Forgetting attribute context — escaping
<and>but not"allows attribute injection. - Using
innerHTMLwithout escaping — the single most common XSS vector. - Trusting server responses — APIs that return “sanitized” HTML may not sanitize everything. Escape on the client too.
- Unescaping too early — unescaping at storage time means the raw payload sits in your database, ready to execute if rendering forgets to re-escape.
Try it
If you have HTML-encoded text you need to decode, or raw text you need to escape for safe insertion into HTML, use a browser-based tool so the conversion stays local.