Normalization vs Canonicalization in Pentest Input Handling

In security tooling, you often compare inputs against a signature database. The attacker controls the input, and you control the parser. If you both use the same rules, you win. If the attacker can express a payload in a way your parser does not normalize, they bypass your detection. This is the same mental model as bypassing a WAF.

Normalization

Normalization means transforming input into a common form before analysis. It removes encoding tricks that an attacker uses to evade exact or substring matching.

from urllib.parse import unquote

def normalize_path(path):
    decoded = unquote(path)          # %27 -> ', %20 -> space
    lowered = decoded.lower()
    collapsed = re.sub(r"[/\\]+", "/", lowered)
    return collapsed

This is what you should do before signature matching.

Canonicalization

Canonicalization goes further. It answers: "what does this input actually resolve to?" Examples include resolving ../ sequences, following symlinks, or applying compression. In web paths, canonicalization might mean resolving /admin/../login to /login. Operating systems and web servers do this internally; your script should do it too if you want to compare paths safely.

from pathlib import Path

canon = Path("/admin/../login").resolve()
# /login

Why This Matters for Defense

If your detector sees /login%2f%2e%2e%2fadmin and you do not decode or collapse it, you will miss the traversal. The same technique bypasses WAFs and log-analysis tools. Normalization is not a silver bullet, but it removes the easiest bypass layer.

Normalization Pipeline for Log Analysis

A practical defensive pipeline looks like this:

  1. Decode URL encoding (urllib.parse.unquote or unquote_plus for + as space).
  2. Decode common encodings a second time if the input still contains % sequences (double URL encoding).
  3. Lowercase for case-insensitive matching.
  4. Collapse path separators and remove null bytes or other control characters.
  5. Resolve . and .. sequences (canonicalization).
  6. Apply your signatures.
from urllib.parse import unquote
from pathlib import PurePosixPath
import re

def normalize_web_path(raw_path):
    decoded = unquote(raw_path)
    while "%" in decoded:
        next_decoded = unquote(decoded)
        if next_decoded == decoded:
            break
        decoded = next_decoded
    cleaned = decoded.lower()
    cleaned = re.sub(r"\x00+", "", cleaned)
    cleaned = re.sub(r"[/\\]+", "/", cleaned)
    cleaned = str(PurePosixPath(cleaned))
    return cleaned

The Offensive Side

As a pentester, you already think this way. You try:

  • URL encoding: %27, %2f, %00
  • Double URL encoding: %2527
  • Case changes: Or, UnIoN, <ScRiPt
  • Path mangling: ../../admin, ..%2f..%2fadmin
  • Unicode normalization: % (full-width percent) instead of %
  • Comment injection or whitespace tricks

Your defensive tools must assume the input has been transformed and apply the inverse transformations first.

Trade-offs

  • Normalization is lossy. If your detection needs the original value for a report, keep both raw and normalized copies.
  • Canonicalization can change semantics. ?id=1+1 as a query string vs a path fragment may mean different things to different servers.
  • Recursive decoding can be dangerous. Limit recursion depth to avoid denial of service from crafted input.