Escape and unescape text — HTML entities, JavaScript strings and JSON strings
Three escape sets that people mix up with each other, in one page, both directions each. Escape for HTML and choose whether the result is going into element text or into an attribute value, because those are not the same job. Escape as a JavaScript string literal, or as a JSON string — which is stricter, and refuses several things JavaScript accepts. Everything runs on the characters you pasted, in this page: nothing is uploaded or saved, and nothing you type is ever inserted into this page as markup.
- Nothing is uploaded or saved
- Works offline
Only the characters that can change how a browser reads the surrounding markup are touched.
Tom & Jerry's <b class="loud">"best"</b> — 5 > 3 · café ☕
Every distinct character rewritten, with the code point it is and the sequence it became. Characters that cannot be shown on their own are named instead.
| Character | Code point | Became | Times |
|---|---|---|---|
| & | U+0026 | & | 1 |
| < | U+003C | < | 2 |
| > | U+003E | > | 3 |
- URL encoder — if the destination is an address rather than markup. It answers the one question that dominates percent-encoding — encodeURI or encodeURIComponent — by showing which characters your own input is treated differently by, which is why URL escaping is not repeated on this page.
- JSON formatter — if you have a whole JSON document rather than one string out of it. The mode here escapes and unescapes a single string literal; it does not parse objects, and it will refuse a raw line break because JSON does.
- Base64 converter — if the goal is to move bytes through something that only carries text — a header, a data URL, a config file. Escaping keeps text readable; Base64 does not, and that is the trade.
- Word to HTML — if you are going the other way and have messy pasted markup to clean rather than plain text to escape.
How it works
- 1
Pick the set, then the direction
HTML, JavaScript string, or JSON string sits above both panels because it changes the meaning of both. Escape and unescape sit beside it, with a Swap button that carries the result back into the input so you can round-trip something and see whether it survives.
- 2
For HTML, say where the text is going
Element text needs & < > handled. An attribute value needs the quote characters handled too, because a stray double quote closes the attribute and everything after it becomes new attributes. The context toggle lives in the output panel, where it changes the output and nothing else.
- 3
For JavaScript, pick your quote
A double-quoted literal must escape double quotes; a template literal must escape backticks and the “${” that starts an interpolation. The tool escapes exactly what the quote style you chose requires, and can optionally push every non-ASCII character out to \uXXXX for a file that has to stay pure ASCII.
- 4
Read the tables under the result
Escaping shows every distinct character it rewrote, with its code point and what it became. Unescaping shows the characters you got back that you cannot see on screen: control characters, non-breaking spaces, zero-width joiners, lone surrogates. That second table is usually the answer to “why does this string look identical and behave differently”.
- 5
Copy it, or download it
Both buttons take the output exactly as shown. There is no server step and no localStorage, so closing the tab is enough to leave nothing behind — which matters, because the strings people bring to an escaper are usually fragments of somebody’s source code.
Frequently asked questions
What is the difference between escaping for element text and for an attribute value?
Element text only has to survive the parser looking for tags and entities, so “&” and “<” are the characters that change meaning, and “>” is escaped as well because old parsers treated it inconsistently and it costs nothing. An attribute value has to survive all of that plus the quote that closes it. Paste the sample above and switch the context toggle: in text mode the double quotes come out untouched, and in attribute mode they become ". Drop the text-mode result into an attribute and you get <b class="loud"> ending early, with “best” read as two more attributes.
Do I need to escape the apostrophe, and why as ' rather than '?
You need it in an attribute, and only because you may not know which quote the surrounding template used — inside a double-quoted attribute a bare apostrophe is harmless. It is written ' because ' is not in HTML 4: it is defined in XML and in HTML5, so a modern browser handles it, but anything parsing your markup as HTML 4 passes it through as five literal characters, while the numeric reference has been understood by everything since the beginning. Note that no escaping saves an unquoted attribute, which HTML still permits: in <input value=hello world> the value is “hello” and “world” is a second attribute. Quote your attributes, then escape for the quote you used — or for both, which is what the attribute context here does.
What is actually different between a JavaScript string and a JSON string?
JSON is a strict subset, and the gaps are where things break. JavaScript has \' , \v, \0, \xHH, \u{1F600} and line continuations; JSON has none of them. JSON has \/ , which JavaScript does not need. JSON forbids single quotes as delimiters outright, and forbids raw control characters inside a string, so a real line break must be written \n. Try it: put \v into the JavaScript unescaper and you get a vertical tab; put the same text into the JSON unescaper and you get a sentence telling you JSON has no such escape and that \u000B is the JSON spelling.
What is a lone surrogate, and why does the tool warn about it?
JavaScript strings are UTF-16, so characters above U+FFFF are stored as a pair of code units — 😀 is \uD83D\uDE00. A lone surrogate is one half of a pair with no partner. It is a perfectly legal JavaScript string and a perfectly illegal Unicode one: TextEncoder replaces it with U+FFFD, a fetch body may reject it, and it cannot be written to a UTF-8 file as it stands. You can produce one here in one step — unescape \uD83D on its own in JavaScript mode — and the result carries the warning. It is the usual cause of a string that was truncated at a fixed byte count and now displays a replacement character.
Why does “\q” come back as just “q” instead of an error?
Because that is the JavaScript rule: an escape the language does not recognise loses its backslash and keeps the character. It is worth knowing precisely because it is silent — a Windows path written "C:\temp" quietly becomes C:<tab>emp, and "C:\qemu" stays readable while "C:\new" does not. The unescaper here says so in a note rather than hiding it. JSON, by contrast, refuses unknown escapes outright, which is the whole reason JSON survives being passed between languages.
What does the “escape / after <” option do?
It writes “<\/” instead of “</”. Inside an inline <script> block the HTML parser is still looking for the closing tag, and it finds “</script>” even in the middle of a string literal — so a JavaScript string containing that text ends your script early and dumps the rest onto the page. “<\/script>” is the same string to JavaScript, because \/ is an unknown escape that resolves to /, and invisible to the HTML parser. The same trick is why the JSON mode offers \/ as an option.
Is escaping the same as sanitising? Does this make user input safe?
No. Escaping is a per-context transformation, and it is only correct for the context you chose. Text escaped for element content is not safe in an attribute, and neither is safe inside a URL attribute — href="javascript:…" contains no character this page would touch. Nothing here is safe inside a <script> block, or inside a CSS value, or inside an unquoted attribute. Escaping at the right boundary is a real and necessary defence; escaping once and reusing the result everywhere is how the bug gets in. Escape at the point of output, in the context of that output.
Why is there no URL encoding mode here?
Because this site already has a page that does it properly. Percent-encoding has one question that dominates everything else — encodeURI or encodeURIComponent — and the URL encoder answers it by showing you which characters your own input is treated differently by, along with the “+ means space” form-data rule. Reproducing half of that here would give you two pages competing to answer the same question and neither doing it fully. Use the URL encoder when the destination is an address; use this page when the destination is markup or source code.
Which named entities can this decode?
The full HTML5 list runs past two thousand names, which is not worth carrying in a page budget. This one knows the Latin-1 block — every accented letter from through ÿ — plus the common punctuation, currency, arrow and mathematics names such as —, …, ’, €, ™, → and ≠. Numeric references are handled in full, in both decimal and hexadecimal, for every code point that exists. Anything it does not recognise is left exactly as it is and reported by name, rather than silently deleted — the exact count is printed under the result.