Python Fundamentals for Security Tasks
Review core Python syntax, data structures, and control flow with a focus on writing small security utility scripts.
Updated in Obsidian: This lesson now includes extra practice on lists and dictionaries.
Practice Exercises
Complete these exercises using the patterns from this lesson: with open(...), pathlib, list/dict operations, and a clean main() + if __name__ == "__main__": structure.
Exercise 1: Clean and Deduplicate a Target List
You have a file targets_raw.txt that may contain:
- Duplicate hostnames or IPs
- Leading/trailing whitespace
- Empty lines
- Lines starting with
#(comments)
Write a script clean_targets.py that reads targets_raw.txt, removes duplicates, ignores comments and empty lines, sorts the remaining entries alphabetically, and writes them to targets_clean.txt.
Example input:
# web servers
192.168.1.1
10.0.0.5
192.168.1.1
# database
10.0.0.5
172.16.0.10
Example output:
10.0.0.5
172.16.0.10
192.168.1.1
Skills practiced: file I/O, string stripping, set deduplication, sorting, writing output files.
Exercise 2: Count Status Codes from a Simulated Scan Log
Write a script count_statuses.py that reads a file scan_results.txt where each line contains an IP address and an HTTP status code separated by a space, like this:
192.168.1.1 200
10.0.0.5 404
172.16.0.10 500
192.168.1.1 301
10.0.0.5 200
Your script should print a summary table showing how many times each status code occurred, sorted by status code. Example output:
200: 2
301: 1
404: 1
500: 1
Skills practiced: dictionaries, counting, parsing lines, sorting, formatted output.
Exercise 3: Build a Small Network Target Validator
Write a script validate_targets.py that reads targets_clean.txt and prints only the targets that look like valid IPv4 addresses. Use the ipaddress standard library module.
For each valid IPv4 address, print it in the format:
VALID: 192.168.1.1
For invalid lines, print:
SKIP: not-an-ip
Skills practiced: importing a standard library module, try/except for validation, combining file reading with conditional logic, and producing structured output.
Recommended Practice Method
For each exercise, follow the Read-Write-Close-Check loop from the lesson:
- Read the exercise carefully and plan your approach.
- Write the script from scratch without looking at example code.
- Close the exercise description and run your script.
- Check the result against the expected output. Fix errors, then write the script again from memory.
Assignment: Build clean_targets.py
Goal: Write a complete, working script that cleans and deduplicates a target list from a file.
Deliverables:
1. A file named clean_targets.py.
2. A sample input file named targets_raw.txt for testing.
3. A generated output file named targets_clean.txt.
Requirements:
- Use with open(...) or pathlib.Path(...).open(...) for file handling.
- Ignore lines that are empty after stripping whitespace.
- Ignore lines that start with # (comments).
- Remove duplicate entries.
- Sort the final entries alphabetically.
- Write the sorted, deduplicated entries to targets_clean.txt, one per line.
- Use a main() function and the if __name__ == "__main__": guard.
- Include a short docstring at the top of the file.
Sample input to test with:
# web servers
192.168.1.1
10.0.5
192.168.1.1
# database
10.0.5
172.16.0.10
# duplicate test
192.168.1.1
Expected output:
10.0.5
172.16.0.10
192.168.1.1
Submission: Paste your final clean_targets.py code here for review. Also mention any bugs or confusions you ran into while writing it.
Lesson Addition: Designing Scripts for Automation and Pipelines
A common mistake when learning Python for security is printing internal data structures or writing only to hardcoded files. In real pentest and forensics work, your Python script is usually one step in a larger pipeline. The output should be predictable, parseable, and composable.
Key Takeaways
- Output one item per line when the data is a list of targets, IPs, or domains. This format works with
sort,uniq,wc,xargs, andnmap -iL. - Default to stdout and provide an optional
-ooutput file. This lets other tools pipe your output. - Keep status messages and errors on
stderrso they never mix with the data stream. - Use
argparsefor real tools instead of hardcoded filenames.
Before Submitting a Script, Ask
- Can another tool pipe this output directly?
- Is the output format obvious and parseable?
- Are errors separate from the data stream?
- Can the user choose the output location?
This is the same philosophy behind the PowerShell + Python Acquisition Pipeline: one tool acquires or cleans, the next consumes. Master this design habit and your scripts become reusable across Bash, PowerShell, and Python workflows.
Why Following the Output Specification Matters
A common mistake when starting out is to skim an exercise, see the "spirit" of the task (e.g., "count status codes"), and then implement a different output format that feels more useful — such as writing JSON instead of a sorted text table.
This instinct is not bad. JSON is excellent for security automation because it is machine-readable and easy to chain into other tools. However, in exercises and in real client work, the requested deliverable format matters because:
- Other tools, scripts, or teammates may expect a specific structure.
- A pipeline may parse the output with
grep,cut, or regex assumptions. - Changing the format without asking breaks downstream consumers.
Practical habit
- Read the requested output format before writing code.
- Implement the spec exactly first.
- Then, if useful, add an optional flag or function to produce an alternate format (e.g.,
--json).
Example extension pattern:
import argparse
import json
from collections import Counter
def count_statuses(path):
counts = Counter()
with open(path, "r") as f:
for line in f:
line = line.strip()
if not line:
continue
parts = line.split()
if len(parts) == 2:
counts[parts[1]] += 1
return counts
def main():
parser = argparse.ArgumentParser()
parser.add_argument("input_file")
parser.add_argument("--json", action="store_true", help="Output JSON instead of text")
args = parser.parse_args()
counts = count_statuses(args.input_file)
if args.json:
print(json.dumps(dict(counts)))
else:
for status, count in sorted(counts.items()):
print(f"{status}: {count}")
if __name__ == "__main__":
main()
This keeps the default behavior matching the spec while giving you the flexibility you wanted.
Lesson Completion Summary
Completed the three core exercises:
- clean_targets.py — Deduplication, whitespace stripping, comment/empty-line filtering, sorted output.
- count_statuses.py — Counting HTTP status codes using dictionary/Counter patterns, sorted output,
main()guard. - validate_targets.py — IPv4 validation using
ipaddress.IPv4Address(), with specificValueErrorhandling andstderrfor diagnostics.
Key takeaways from the exercises:
- Use
with open(...)for automatic file cleanup. - Prefer
collections.Counterfor counting tasks. - Iterate over the file object directly (
for line in f) instead ofreadlines()for large files. - Use
line.strip().split()to parse lines cleanly. - Catch specific exceptions (
ValueError) instead of bareexcept:. - Send diagnostic messages to
sys.stderrto keepstdoutclean for data output. - Use
ipaddress.IPv4Address()when you want only IPv4, notip_address()which also accepts IPv6.
Common pitfalls observed:
- Adding extra output formats (JSON) when the exercise asks for plain text.
- Using bare
except:which hides real bugs. - Using
ipaddress.ip_address()when the task requires IPv4 only.