Defensive Input Handling for Security Scripts

Security scripts often process data that comes from untrusted or messy sources: scan output, logs, network traffic, user uploads, or third-party tools. Writing code that assumes the input is clean is dangerous. A defensive mindset treats every input as potentially malformed, malicious, or surprising.

What Could Go Wrong

Functional Problems

  • Missing fields — a line has fewer columns than expected.
  • Extra whitespace — tabs, trailing spaces, leading newlines.
  • Empty lines — common at end of files or between sections.
  • Duplicate entries — cause double-counting or duplicate actions.
  • Encoding issues — files not in UTF-8, null bytes, control characters.
  • Unexpected types — numbers where strings expected, vice versa.

Security Problems

  • Path traversal — a filename like ../../../etc/passwd passed to file operations.
  • Command injection — building shell commands from untrusted strings.
  • Injection into templates, SQL, JSON, YAML — malicious payloads parsed as code.
  • Logic bypasses — malformed input skipping validation checks.
  • Resource exhaustion — enormous files, deeply nested data, or infinite loops.
  • Bare except: swallowing real errors — hides bugs and attack artifacts.

Defensive Patterns

Validate and parse early, fail explicitly

from pathlib import Path

def read_targets(path_str):
    path = Path(path_str).resolve()
    allowed_dir = Path("./data").resolve()
    if not path.is_relative_to(allowed_dir):
        raise ValueError(f"Path {path} is outside allowed directory {allowed_dir}")
    with open(path, "r", encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if not line or line.startswith("#"):
                continue
            yield line

Whitelist, do not blacklist

Instead of trying to reject every bad thing, define what is acceptable and reject everything else.

import re

def looks_like_ipv4(text):
    pattern = re.compile(r"^(\d{1,3}\.){3}\d{1,3}$")
    return bool(pattern.match(text))

Better yet, use ipaddress.IPv4Address() for real validation rather than regex alone.

Use specific exceptions

try:
    ip = ipaddress.IPv4Address(line)
except ValueError:
    skipped += 1

Avoid bare except: at all costs. It catches KeyboardInterrupt, SystemExit, and programming errors.

Sanitize output and log to different channels

Print results to stdout and diagnostics to stderr so scripts can be chained in pipelines.

print(f"Valid target: {ip}")
print(f"Skipped {skipped} lines", file=sys.stderr)

Limit resource consumption

MAX_LINES = 1_000_000
for i, line in enumerate(f):
    if i > MAX_LINES:
        raise RuntimeError("File too large")
    ...

The Pentest Mindset

Attackers think about what happens when the input is not what you expect. They look for:

  • Places where validation happens after use.
  • Exceptions that are silently swallowed.
  • Inputs that are concatenated into commands or paths.
  • Differences between how Python parses something and how another tool parses it.

When you write a security script, ask:

  1. Who controls this input?
  2. What is the worst thing they could put here?
  3. Am I failing safely when input is bad?
  4. Can I run this script on attacker-controlled data without trusting it?