Normalization vs Canonicalization in Pentest Input Handling
In security tooling, you often compare inputs against a signature database. The attacker controls the input, and you control the parser. If you both use the same rules, you win. If the attacker can express a payload in a way your parser does not normalize, they bypass your detection. This is the same mental model as bypassing a WAF.
Normalization
Normalization means transforming input into a common form before analysis. It removes encoding tricks that an attacker uses to evade exact or substring matching.
from urllib.parse import unquote
def normalize_path(path):
decoded = unquote(path) # %27 -> ', %20 -> space
lowered = decoded.lower()
collapsed = re.sub(r"[/\\]+", "/", lowered)
return collapsed
This is what you should do before signature matching.
Canonicalization
Canonicalization goes further. It answers: "what does this input actually resolve to?" Examples include resolving ../ sequences, following symlinks, or applying compression. In web paths, canonicalization might mean resolving /admin/../login to /login. Operating systems and web servers do this internally; your script should do it too if you want to compare paths safely.
from pathlib import Path
canon = Path("/admin/../login").resolve()
# /login
Why This Matters for Defense
If your detector sees /login%2f%2e%2e%2fadmin and you do not decode or collapse it, you will miss the traversal. The same technique bypasses WAFs and log-analysis tools. Normalization is not a silver bullet, but it removes the easiest bypass layer.
Normalization Pipeline for Log Analysis
A practical defensive pipeline looks like this:
- Decode URL encoding (
urllib.parse.unquoteorunquote_plusfor+as space). - Decode common encodings a second time if the input still contains
%sequences (double URL encoding). - Lowercase for case-insensitive matching.
- Collapse path separators and remove null bytes or other control characters.
- Resolve
.and..sequences (canonicalization). - Apply your signatures.
from urllib.parse import unquote
from pathlib import PurePosixPath
import re
def normalize_web_path(raw_path):
decoded = unquote(raw_path)
while "%" in decoded:
next_decoded = unquote(decoded)
if next_decoded == decoded:
break
decoded = next_decoded
cleaned = decoded.lower()
cleaned = re.sub(r"\x00+", "", cleaned)
cleaned = re.sub(r"[/\\]+", "/", cleaned)
cleaned = str(PurePosixPath(cleaned))
return cleaned
The Offensive Side
As a pentester, you already think this way. You try:
- URL encoding:
%27,%2f,%00 - Double URL encoding:
%2527 - Case changes:
Or,UnIoN,<ScRiPt - Path mangling:
../../admin,..%2f..%2fadmin - Unicode normalization:
%(full-width percent) instead of% - Comment injection or whitespace tricks
Your defensive tools must assume the input has been transformed and apply the inverse transformations first.
Trade-offs
- Normalization is lossy. If your detection needs the original value for a report, keep both raw and normalized copies.
- Canonicalization can change semantics.
?id=1+1as a query string vs a path fragment may mean different things to different servers. - Recursive decoding can be dangerous. Limit recursion depth to avoid denial of service from crafted input.
Related Concepts
- Defensive Input Handling for Security Scripts — whitelisting, validation, and safe failure.
- Regex Fundamentals — pattern syntax for signature detection.
- Path Traversal — why
../sequences matter. - Command Injection — input normalization failures in command contexts.