Every URL you’ve ever typed or clicked has been through URL encoding. The space in “my file.txt” becomes %20. The & in a query string becomes %26. The / in a path stays as / — but only because the encoder knows which characters are safe and which aren’t.
URL encoding (officially “percent-encoding”) is how URLs represent characters that aren’t in the unreserved character set. It’s a simple transformation, but the edge cases are where most API bugs live.
The unreserved characters
These characters are always safe in URLs and never need encoding:
A-Z a-z 0-9 - _ . ~
Everything else — spaces, slashes, ampersands, equals signs, non-ASCII characters — needs to be encoded as a percent sign followed by two hex digits:
- Space →
%20 &→%26=→%3D/→%2F?→%3F#→%23
The encoding is the character’s UTF-8 byte values, each represented as two hex digits. A multi-byte character like é (UTF-8: 0xC3 0xA9) becomes %C3%A9.
Where encoding matters
Query parameters. This is where URL encoding causes the most bugs. If a query value contains & or =, the URL parser splits on those characters and breaks the parameter structure:
/search?q=cats&dogs ← ambiguous: is "dogs" a separate parameter?
/search?q=cats%26dogs ← correct: "cats&dogs" is one value
Every HTTP library has a function for encoding query parameters. Use it. Never concatenate strings into a query string by hand.
Paths. Spaces in file names need encoding:
/download/my file.pdf ← broken
/download/my%20file.pdf ← works
But slashes inside a path segment need encoding too:
/file/path/segment ← two segments: "file/path" split at /
/file%2Fpath/segment ← one segment: "file/path"
Non-ASCII characters. URLs are ASCII-only. Any non-ASCII character (accented letters, CJK characters, emoji) must be percent-encoded:
café→caf%C3%A9日本語→%E6%97%A5%E6%9C%AC%E8%AA%9E🎉→%F0%9F%8E%89
Most browsers display the decoded version in the address bar, but the actual bytes sent over the wire are encoded.
The double-encoding trap
The most common URL encoding bug: encoding an already-encoded string.
Take the path /hello%20world. If you URL-encode it again, the % becomes %25:
/hello%20world ← original (correct)
/hello%2520world ← double-encoded (broken)
The server decodes %25 to %, then sees %20 and decodes that to a space. You end up with /hello world — but only if the server does a single decode. If it does two decodes (some do), you get the original back. The behavior is inconsistent and unpredictable.
The rule: encode once, decode once. If you receive an encoded string, decode it before re-encoding. Check whether your HTTP library’s URL builder expects raw or pre-encoded values — the difference between path and rawPath in most frameworks matters.
Query string encoding
Query strings have their own encoding rules. The key difference from paths: + represents a space in query strings (application/x-www-form-urlencoded), but %20 also represents a space. Both are valid, but they come from different standards:
application/x-www-form-urlencoded(form submissions) —+for spaces- Percent-encoding (URLs) —
%20for spaces
Most modern APIs accept both. But if you’re parsing a query string from a form submission, + → space. If you’re building a URL, %20 → space. Mixing them up causes subtle bugs where spaces turn into + signs or vice versa.
Use your language’s URL encoding function. It handles the distinction for you.
Encoding vs escaping
URL encoding is not HTML escaping. They solve different problems:
- URL encoding (
%20) — for characters in URLs. Prevents characters from being interpreted as URL syntax. - HTML escaping (
&,<) — for characters in HTML. Prevents characters from being interpreted as HTML tags or entities.
A URL inside an HTML href needs both: URL-encode the parameter values, then HTML-encode the entire URL:
<a href="/search?q=cats%26dogs&page=1">
The %26 keeps & from being interpreted as a query separator. The & keeps & from being interpreted as an HTML entity start.
Common gotchas
Spaces in different contexts. %20 in a URL, + in a form body, %2520 if you accidentally double-encoded. Know which context you’re in.
Hash fragments. Everything after # is not sent to the server. If you encode a URL with a fragment and the fragment contains ? or &, the server never sees them. The fragment is client-only.
Encoding the percent sign. % → %25. If a literal % appears in your data (like a password with % in it), it must be encoded. If you don’t encode it, the parser interprets the two characters after it as a hex sequence.
Unicode normalization. Some systems normalize Unicode before encoding. café (with combining accent) and café (with precomposed é) encode to different percent sequences. This causes issues with filenames and internationalized domain names.
Quick reference
| Character | Percent-encoded | Context |
|---|---|---|
| Space | %20 |
URL |
| Space | + |
Form body |
& |
%26 |
Query string |
= |
%3D |
Query string |
/ |
%2F |
Path segment |
? |
%3F |
Query string |
# |
%23 |
Path/query |
% |
%25 |
Everywhere |
+ |
%2B |
Query string |
Try it
If you have a URL or encoded string you need to decode (or raw data you need to encode), use a local tool so the conversion happens in your browser — no data sent anywhere.