HTML encoder and decoder for text and entities
Convert special characters into HTML entities or decode character references back into readable text. The default escapes text for HTML while keeping accented letters, symbols, and emoji readable. Paste your text and copy the result, or open Options for other uses.
Loading the tool…
What is HTML entity encoding?
HTML uses characters such as < and & as part of its syntax. To display them as ordinary text, you can replace them with character references, often called HTML entities. For example, < displays a less-than sign and & displays an ampersand.
There are three common forms: a named reference such as ©, a decimal reference such as ©, and a hexadecimal reference such as ©. All three represent the copyright symbol, ©. Names are case-sensitive; include the final semicolon when writing references.
Encoding is useful when you want to show an HTML example on a page, insert literal text into a template, or represent a symbol with an explicit character reference. Decoding helps you read entity-encoded text from an export, source file, or copied snippet. It converts references to characters; it does not remove HTML tags or render a page.
For example, encoding this text:
<strong>Hello & welcome</strong>
produces HTML source that displays the tags literally:
<strong>Hello & welcome</strong>
Encoding already encoded text can double its entities: & becomes &amp;. Decode first only when you know the source is already entity-encoded.
For normal page output, use your template engine’s automatic escaping. HTML text escaping is not a general HTML sanitizer, JavaScript string escaper, or URL encoder; attribute values need appropriate quoting and context-specific validation.
Common HTML entities and reserved characters
Use these tables as a quick reference, not as a complete list. < and & have special roles in HTML text; quotes matter when they delimit attributes. The currency signs, mathematical symbols, and spacing characters below are useful references, not all reserved syntax. Decimal and hexadecimal references represent the same Unicode character; for example, © and © both mean ©. The x marks a hexadecimal value.
The HTML5 tab covers modern HTML. HTML4 shows commonly used legacy names and the different treatment of apostrophes. XML lists its five predefined names and explains when numeric references are needed. Changing a reference tab does not change the converter’s settings.
HTML5 character references
Quotes can remain literal in ordinary HTML text; escape the matching quote when inserting text into a quoted attribute.
The WHATWG named character reference list defines the names supported by modern HTML.
| Character | Entity name | Decimal reference | Hex reference | Description |
|---|---|---|---|---|
& | & | & | & | Ampersand; begins a character reference. |
< | < | < | < | Less-than sign; begins markup. |
> | > | > | > | Greater-than sign; usually literal in text, but may be escaped. |
" | " | " | " | Double quote; escape it inside a double-quoted attribute. |
' | ' | ' | ' | Apostrophe; escape it inside a single-quoted attribute. |
| Non-breaking space | |   |   | Keeps adjacent words on the same line. |
| © | © | © | © | Copyright symbol. |
| ® | ® | ® | ® | Registered trademark symbol. |
| ™ | ™ | ™ | ™ | Trademark symbol. |
| € | € | € | € | Euro currency sign. |
| £ | £ | £ | £ | Pound currency sign. |
| ¥ | ¥ | ¥ | ¥ | Yen or yuan currency sign. |
| × | × | × | × | Multiplication sign. |
| ÷ | ÷ | ÷ | ÷ | Division sign. |
| ≤ | ≤ | ≤ | ≤ | Less-than or equal-to sign. |
| ≥ | ≥ | ≥ | ≥ | Greater-than or equal-to sign. |
| → | → | → | → | Right arrow. |
| ✓ | ✓ | ✓ | ✓ | Check mark; this name is not in HTML4. |
HTML4 character references
The W3C HTML 4.01 entity reference documents the older named set. HTML4 does not define '; use ' for an apostrophe when that compatibility matters. These shared names also work in HTML5.
| Character | Entity name | Decimal reference | Hex reference | Description |
|---|---|---|---|---|
& | & | & | & | Ampersand; escape it when displaying a literal ampersand. |
< | < | < | < | Less-than sign; escape it when displaying literal text. |
> | > | > | > | Greater-than sign. |
" | " | " | " | Double quotation mark. |
' | No HTML4 name | ' | ' | Apostrophe; use the numeric reference. |
| Non-breaking space | |   |   | Space that prevents a line break. |
| © | © | © | © | Copyright symbol. |
| ® | ® | ® | ® | Registered trademark symbol. |
| ™ | ™ | ™ | ™ | Trademark symbol. |
| € | € | € | € | Euro currency sign. |
| £ | £ | £ | £ | Pound currency sign. |
| ¥ | ¥ | ¥ | ¥ | Yen or yuan currency sign. |
| × | × | × | × | Multiplication sign. |
| ÷ | ÷ | ÷ | ÷ | Division sign. |
| ≤ | ≤ | ≤ | ≤ | Less-than or equal-to sign. |
| ≥ | ≥ | ≥ | ≥ | Greater-than or equal-to sign. |
| → | → | → | → | Right arrow. |
XML character references
XML defines only five predefined entity names. They work without a DTD or custom entity declaration. References are case-sensitive and require a semicolon.
| Character | Entity name | Decimal reference | Hex reference | Description |
|---|---|---|---|---|
& | & | & | & | Escape a literal ampersand in text or attributes. |
< | < | < | < | Escape a literal less-than sign in text or attributes. |
> | > | > | > | Optional in ordinary text, except to avoid a literal ]]> sequence. |
" | " | " | " | Escape inside a double-quoted attribute. |
' | ' | ' | ' | Escape inside a single-quoted attribute. |
Names such as and © are not predefined in XML. Use   for a non-breaking space or © for ©, or declare the names in a DTD. Numeric references must still identify characters allowed by XML.
Choose Text with quotes to escape these five characters. This tool does not validate XML or remove forbidden characters, and Decode mode uses HTML5 rules rather than parsing XML.
HTML encoding and decoding in code
For repeatable conversions in an application, use a standard library or a maintained entity library. These examples encode the same short text and decode it again. They demonstrate string conversion, not rendering untrusted HTML. Libraries differ in their handling of quotes, Unicode, and incomplete references, so equivalent output need not use identical entity spellings.
JavaScript and TypeScript
The entities library works in Node.js and browser bundles. Install it with npm install entities. This example uses escapeText, the same function as this tool’s default HTML text (recommended) option. Use escapeUTF8 for Text with quotes, or encodeNonAsciiHTML for Extended entities.
import { escapeText, decodeHTML } from "entities";
const text = "<p>Tea & coffee</p>";
const encoded = escapeText(text);
const decoded = decodeHTML(encoded);
console.log(encoded); // <p>Tea & coffee</p>
console.log(decoded); // <p>Tea & coffee</p>
When putting plain text on a web page, assigning it to textContent avoids interpreting it as markup. encodeURIComponent performs URL encoding, which is a different operation.
Python
Python’s built-in html module includes both operations. html.escape escapes quotes by default while leaving ordinary Unicode text unchanged; html.unescape recognizes HTML5 named and numeric references.
import html
text = "<p>Tea & coffee</p>"
encoded = html.escape(text, quote=True)
decoded = html.unescape(encoded)
print(encoded) # <p>Tea & coffee</p>
print(decoded) # <p>Tea & coffee</p>
PHP
Use htmlspecialchars for special characters and html_entity_decode to decode references. Explicit flags select HTML5, escape both quote types, and replace invalid UTF-8 sequences during encoding. Ordinary Unicode letters remain unchanged.
<?php
$text = '<p>Tea & coffee</p>';
$flags = ENT_QUOTES | ENT_SUBSTITUTE | ENT_HTML5;
$encoded = htmlspecialchars($text, $flags, 'UTF-8');
$decoded = html_entity_decode($encoded, $flags, 'UTF-8');
// Run in the CLI to inspect both strings as text.
echo $encoded, PHP_EOL; // <p>Tea & coffee</p>
echo $decoded, PHP_EOL; // <p>Tea & coffee</p>
C# and .NET
System.Net.WebUtility provides HTML conversion without an additional package. HtmlEncode and HtmlDecode are convenient for strings; do not assume every HTML5 named reference is recognized by every platform’s decoder.
using System;
using System.Net;
string text = "<p>Tea & coffee</p>";
string encoded = WebUtility.HtmlEncode(text);
string decoded = WebUtility.HtmlDecode(encoded);
Console.WriteLine(encoded); // <p>Tea & coffee</p>
Console.WriteLine(decoded); // <p>Tea & coffee</p>
Go
Go’s standard html package escapes the five special characters and decodes a broader set of references. For an actual HTML template, prefer html/template, which applies escaping according to the surrounding context.
package main
import (
"fmt"
"html"
)
func main() {
text := "<p>Tea & coffee</p>"
encoded := html.EscapeString(text)
decoded := html.UnescapeString(encoded)
fmt.Println(encoded) // <p>Tea & coffee</p>
fmt.Println(decoded) // <p>Tea & coffee</p>
}
Rust
Add the html-escape crate with cargo add html-escape. Its encode_text function escapes &, <, and > for ordinary HTML text. It leaves quotes alone; use the crate’s appropriate attribute encoder for a quoted attribute value. decode_html_entities decodes terminated named and numeric references, with stricter behavior than a browser for malformed input.
use html_escape::{decode_html_entities, encode_text};
fn main() {
let text = "<p>Tea & coffee</p>";
let encoded = encode_text(text);
let decoded = decode_html_entities(encoded.as_ref());
println!("{}", encoded); // <p>Tea & coffee</p>
println!("{}", decoded); // <p>Tea & coffee</p>
}
Java
Add Apache Commons Text (org.apache.commons:commons-text) to your Maven or Gradle project. Its StringEscapeUtils provides escapeHtml4 and unescapeHtml4. These use HTML4 names, not the full HTML5 set; apostrophes are not escaped by escapeHtml4.
import org.apache.commons.text.StringEscapeUtils;
public class HtmlEntitiesExample {
public static void main(String[] args) {
String text = "<p>Tea & coffee</p>";
String encoded = StringEscapeUtils.escapeHtml4(text);
String decoded = StringEscapeUtils.unescapeHtml4(encoded);
System.out.println(encoded); // <p>Tea & coffee</p>
System.out.println(decoded); // <p>Tea & coffee</p>
}
}
Swift
Add SwiftSoup through Swift Package Manager and include its product in your target. Use Entities.escape for HTML text and Parser.unescapeEntities to decode references without stripping tags or rendering HTML. The false argument selects text rather than attribute context for decoding. Text escaping leaves quotes unchanged, so it is not a quoted-attribute encoder.
import SwiftSoup
let text = "<p>Tea & coffee</p>"
let encoded = Entities.escape(text)
do {
let decoded = try Parser.unescapeEntities(encoded, false)
print(encoded) // <p>Tea & coffee</p>
print(decoded) // <p>Tea & coffee</p>
} catch {
print("Could not decode HTML entities: \(error)")
}
Delphi / Object Pascal
Delphi’s TNetEncoding.HTML, from System.NetEncoding, provides Encode and Decode. It escapes &, <, >, and double quotes. Its documented decoder recognizes numeric references and the names amp, lt, gt, and quot; do not expect full HTML5 named-entity support or apostrophe escaping. This example uses Delphi’s runtime library, which is not shared by every Object Pascal compiler.
program HtmlEntitiesExample;
{$APPTYPE CONSOLE}
uses
System.NetEncoding;
var
Text, Encoded, Decoded: string;
begin
Text := '<p>Tea & coffee</p>';
Encoded := TNetEncoding.HTML.Encode(Text);
Decoded := TNetEncoding.HTML.Decode(Encoded);
Writeln(Encoded); // <p>Tea & coffee</p>
Writeln(Decoded); // <p>Tea & coffee</p>
end.
How to use this tool
- Choose a conversion. Select Encode HTML entities to turn characters such as
<and&into entity text. Select Decode HTML entities to turn references such as<and©back into readable characters. - Choose an encoding style if needed. In Encode mode, open Options to change the style. Keep HTML text (recommended) for ordinary text on a web page. Choose Text with quotes for quoted attribute values, or Extended entities when you need more characters converted. Options are hidden in Decode mode.
- Type or paste your text into Input. Output updates immediately as you type or change a setting. For example, encoding
Tea & coffeeproducesTea & coffee; decoding that result restores the original text. - Copy or reuse the result. Select Copy result beside Output. Inverse moves the result into Input without changing the conversion mode. To undo an encoding, switch to Decode after using Inverse. Clear beside Input empties both fields.
Choosing an encoding style
Modern HTML supports Unicode: accented letters, symbols, and emoji can stay readable in a UTF-8 page. HTML5 does not require converting them into entities. These options change how much is escaped, not which browser version you support. The HTML syntax reference describes the rules for text and attributes.
- HTML text (recommended) is the default for text between tags, such as a paragraph. It escapes
&,<, and>and writes non-breaking spaces as . Quotes and other Unicode characters stay unchanged. For example,café & tea ©becomescafé & tea ©. - Text with quotes escapes
&,<,>, and both quote types, using names that also work in XML. Choose it for text inside a quoted attribute such astitle="...", or for XML text. Accented letters, emoji, and non-breaking spaces stay unchanged. It does not validate XML or make URLs and JavaScript safe. - Extended entities also converts non-ASCII characters into named or numeric references. For example,
café ©becomescafé ©. Choose it when another system or your preferred source format calls for entity-based text; ordinary modern web pages do not need it.
The default also escapes > and non-breaking spaces for consistent output, although HTML does not always require this. For attribute values, use Text with quotes and keep the surrounding quotes in your markup. Decode recognizes HTML5 named and numeric references regardless of the encoding option you last used.
If the result is unexpected
Encoding text that already contains entities can encode it again: & becomes &amp;. Check whether your source needs decoding first. Decoding displays markup as text; it does not remove tags. If copying is unavailable, select the output and copy it with your keyboard or context menu.
Your tool data stays in your browser
This tool processes your data on your device. The text, files, or passwords you enter and the results it creates are not sent to our servers or any other server.
Page and asset requests still occur. Production pages also load Google Tag Manager and AdSense advertising; tool code does not send inputs or results to analytics or ads. See our privacy policy.
To inspect requests while using the tool, open your browser's developer tools and select Network. Try it with sample data.
Read How to check website network requests in Chrome, Firefox, Edge, and Safari.
Learn more and references
- Wikipedia: XML and HTML character entity references - an overview of named references, numeric forms, and differences between document standards.
- WHATWG: named character references - the authoritative list of HTML names and the Unicode characters they represent.
- W3C: XML predefined entities - the five built-in XML entity names and their relationship to numeric references.
- MDN: character references - a concise introduction to entity syntax and why literal markup characters need escaping.
- OWASP: cross-site scripting prevention - practical guidance on choosing output encoding for HTML text, attributes, JavaScript, and other contexts.