Structured Binary Parsing with struct and ctypes

Interpret C-like data structures, pack and unpack binary formats, and interface with system libraries.

Lesson Example: Reverse Engineering a PS2 Asset Archive (GRAPHRES.BIN)

A common reverse-engineering task is extracting assets from a proprietary game archive. The following script, used against a PlayStation 2 game's GRAPHRES.BIN, demonstrates how struct and binary offsets can turn a raw ROM blob into individual files.

The Script

import struct
import os

filename = "_GRAPHRES.BIN"
outdir = "extracted"

TABLE_OFFSET = 0x10

os.makedirs(outdir, exist_ok=True)

entries = []

with open(filename, "rb") as f:
    magic = f.read(4)                            # 1. read 4 raw bytes
    entry_count = struct.unpack("<I", f.read(4))[0] # 2. read a little-endian uint32

    print("Magic:", magic)
    print("Entries:", entry_count)

    f.seek(TABLE_OFFSET)                          # 3. jump to the file table

    for i in range(entry_count):
        entry_id, offset = struct.unpack("<II", f.read(8))
        entries.append((entry_id, offset))        # 4. collect (id, offset) pairs

    entries.sort(key=lambda x: x[1])              # 5. sort by offset

    f.seek(0, 2)                                  # 6. jump to end of file
    file_size = f.tell()

    for i, (entry_id, offset) in enumerate(entries):
        # 7. compute size from current offset to next offset (or EOF)
        if i < len(entries) - 1:
            next_offset = entries[i+1][1]
        else:
            next_offset = file_size

        size = next_offset - offset
        if size <= 0:
            continue

        f.seek(offset)
        data = f.read(size)

        # 8. naive magic-byte type detection
        header = data[:4]
        if header == b"TIM2":
            typ = "texture_TIM2"
        elif data[:2] in [b"\x06\x00", b"\x0C\x00", b"\x32\x00"]:
            typ = "struct_data"
        elif all(32 <= b < 127 for b in data[:16]):
            typ = "ascii_text"
        else:
            typ = "unknown"

        outname = f"{outdir}/res_{entry_id}_{offset:08X}_{typ}.bin"
        with open(outname, "wb") as out:
            out.write(data)

        print(f"ID {entry_id:5}  offset 0x{offset:08X}  size {size:7}  type {typ}")

What it is doing, line by line

  1. Magic bytes — f.read(4) pulls the first four bytes of the file. Many custom file formats start with a signature like b"TIM2" or a short identifier. This script reads the magic but does not validate it, which is fine for exploration but risky in a robust tool.

  2. Entry count — struct.unpack("<I", f.read(4))[0] reads four bytes and interprets them as a little-endian unsigned 32-bit integer. This is the number of files/assets stored inside the archive. The < means little-endian; I means 4-byte unsigned int.

  3. Seeking to the table — f.seek(TABLE_OFFSET) moves the read position to 0x10 (16 bytes in). The script assumes the file table starts there, after an 8-byte header (4 magic + 4 count). If the table actually starts elsewhere, the parsing will be wrong.

  4. Reading the table — Each entry is two 32-bit values: an ID and an offset. <II unpacks 8 bytes into two unsigned ints. The ID is used for the filename; the offset marks where the asset data begins inside the archive.

  5. Sorting by offset — The table might not be ordered by offset, but the script needs consecutive offsets to calculate each file's size. Sorting makes the next step possible.

  6. File size — f.seek(0, 2) seeks to the end of the file; f.tell() returns the total size. This is needed for the last asset, which runs until EOF.

  7. Size inference — Because the table only stores offsets, the script assumes each asset spans from its offset to the next asset's offset. The last asset uses EOF. This is a common pattern when the archive does not store explicit sizes. If the table is incomplete or the file has trailing padding, sizes can be wrong.

  8. Type detection — The script guesses the file type from its first few bytes:

  9. b"TIM2" is the PlayStation 2 texture format (TM2).
  10. A few 2-byte signatures are labeled as struct data (this is a guess).
  11. If every byte in the first 16 bytes is printable ASCII, it calls it text.
  12. Otherwise, unknown.

Important caveats

  • Hardcoded offset. The table starts at 0x10 only if the header is exactly 8 bytes. Always verify the header layout in a hex editor like 010 Editor, ImHex, or GHex.
  • No explicit sizes. Size-by-next-offset fails if the archive has gaps, compressed data, or a separate size table. If you see one huge file at the end, the table may be missing entries or the last entry may be padding.
  • Naive type detection. Magic bytes are useful first-pass heuristics, but they are not enough for accurate reconstruction. A TIM2 file, for example, has its own header fields for image width, height, and pixel format that this script ignores.
  • Little-endian assumption. PS2 data is little-endian, but other systems (e.g., GameCube, some older consoles) may be big-endian. The format string would change to >I for big-endian.

How to make this script safer and more useful

  • Replace hardcoded TABLE_OFFSET with a command-line argument or config.
  • Read the entry size table if one exists; do not always assume size-by-next-offset.
  • Validate the magic bytes and raise a clear error if the file does not look like the expected format.
  • Add a checksum or known-size sanity check for each extracted asset.
  • Use pathlib.Path for output paths instead of f-string concatenation.
  • struct for binary packing and unpacking
  • ctypes for C-style structure layouts
  • Capstone for disassembling extracted code sections
  • File handling pitfalls for binary files

Iterative Analysis: When Extraction Yields Mostly Unknown Files

A first-pass extractor rarely identifies every asset correctly. The magic bytes of the container (LINK) do not guarantee the format of assets inside. When most extracted files are labeled unknown or struct, the next phase is structure inference on the extracted blobs and the file table itself.

Why the original TIM2/struct/ascii heuristic failed

The script guessed asset types from the first bytes of each extracted chunk. In this archive, assets may not store their own magic bytes at the front. They may:
- Start with size or flag fields instead of a file signature.
- Be compressed, encrypted, or bit-packed.
- Be sub-structures whose type is encoded elsewhere (e.g. in the entry_id or a type field in the table).

The first 16 bytes were:

4C 49 4E 4B 06 04 00 00  00 08 00 00 00 00 00 00

Decoded:
- 4C 49 4E 4B → magic LINK (4 bytes)
- 06 04 00 00 → 0x00000406 = 1030 entries (4 bytes)
- 00 08 00 00 → 0x00000800 = 2048 (4 bytes, unknown purpose)
- 00 00 00 00 → 0 (4 bytes, possibly padding or a type/version field)

The field at 0x08 holding 2048 is suspicious. It could be:
- a base offset added to every table offset,
- an alignment value,
- a maximum asset size, or
- a second table size.

Before treating 2048 as padding, it should be inspected.

Next-step analysis workflow

  1. Inspect the table itself. Print the first 20 entries to look for patterns in entry_id and offset.
  2. Look at raw extracted files in a hex editor. Check whether any have recognizable headers slightly later in the file (not at offset 0).
  3. Correlate file size with entry_id. If certain entry_id ranges always produce files of the same size, the ID may encode a type.
  4. Check entropy. Compressed/encrypted data has high entropy. Textures and models often have lower entropy and repeating patterns.
  5. Search for known magic bytes inside the extracted files. Many formats start with R, RIFF, TGA, PSM, TMD, etc.
  6. Check the gap between table end and first asset.
  7. Test if the 2048 value is a base offset. Add it to every offset and see if the resulting bytes make more sense.

Python snippets for the next phase

from construct import Struct, Int32ul, Array, Bytes, this

Header = Struct(
    "magic" / Bytes(4),
    "entry_count" / Int32ul,
    "field_08" / Int32ul,
    "field_0C" / Int32ul,
    "entries" / Array(this.entry_count, Struct(
        "entry_id" / Int32ul,
        "offset" / Int32ul,
    )),
)

with open("_GRAPHRES.BIN", "rb") as f:
    parsed = Header.parse_stream(f)

for e in parsed.entries[:20]:
    print(f"entry_id={e.entry_id:5}  offset=0x{e.offset:08X}")

Search for known magic bytes inside extracted files

import os
import re

KNOWN_MAGICS = {
    b"TIM2": "TIM2",
    b"RIFF": "RIFF",
    b"TGA\x00": "TGA",
    b"PSM\x00": "PSM",
    b"\x01\x00\x00\x00": "maybe_struct",
}

for fname in os.listdir("extracted"):
    path = os.path.join("extracted", fname)
    with open(path, "rb") as f:
        data = f.read()
    for magic, label in KNOWN_MAGICS.items():
        if magic in data[:128]:
            print(f"{fname}: found {label} at start")

Compute entropy of a file

import math

def entropy(data):
    if not data:
        return 0.0
    counts = [0] * 256
    for b in data:
        counts[b] += 1
    ent = 0.0
    length = len(data)
    for c in counts:
        if c == 0:
            continue
        p = c / length
        ent -= p * math.log2(p)
    return ent

# usage
print(entropy(open("extracted/res_123_00001234_unknown.bin", "rb").read()))

Entropy interpretation:
- 8.0 = fully random (encrypted or compressed)
- 7.5–8.0 = compressed data
- 5.0–7.0 = executable code, textures, or structured data
- 4.0–5.0 = text or highly structured data
- near 0.0 = all same byte (null padding)

Check if 2048 is a base offset

for e in parsed.entries[:20]:
    adjusted = e.offset + 0x800
    print(f"entry_id={e.entry_id}  raw=0x{e.offset:08X}  adjusted=0x{adjusted:08X}")

If adjusted offsets point to recognizable data while raw offsets do not, the field at 0x08 is a base offset, not padding.

Practical goals for this phase

  • Determine whether every table entry is a real asset.
  • Identify whether entry_id encodes a type or category.
  • Find the real internal format of the most common unknown files.
  • Decide whether the archive needs a smarter parser or a post-processing step.

Interpreting the Asset Table

Once the archive header is parsed, the first useful step is to print a few table entries and look for patterns. For example, if the output looks like this:

entry_id=    5  offset=0x00506800
entry_id= 2579  offset=0x001BC820
entry_id= 3469  offset=0x00002860
entry_id= 3475  offset=0x0001E6E0
entry_id= 3536  offset=0x00000310

What this tells us

  1. IDs are not array indices. A sequential index would be 0, 1, 2, 3, 4. Values like 2579 or 3536 are usually engine-level identifiers, possibly hashes, or composite values where high bits encode a category.
  2. Offsets are scattered. The assets are not stored in the same order as the table. This confirms we must sort by offset before calculating sizes.
  3. Large gaps. Offsets range from 0x00000310 to 0x00506800, so the file is at least several megabytes and contains assets at very different file positions.

Hypotheses to test

Observation Possible meaning
Sequential IDs but shuffled offsets Table is sorted by ID, not by file position.
Large jumps in IDs IDs may be hashes, or split into category + index.
Round offsets (e.g., multiples of 0x800) Files may be aligned to a sector or block size.
Offsets smaller than the table end The table might be interleaved with assets, or a base offset is missing.

Next checks

  • Print the minimum and maximum offsets, and the table end position.
  • Check if offsets are aligned to a common boundary like 0x800, 0x1000, or 0x40.
  • Look at the first asset after sorting by offset, not the first table entry.
  • Test whether the unknown header field at 0x08 (e.g., 0x00000800) is a base offset added to each table entry.

This kind of pattern recognition is what turns a raw dump into a meaningful format understanding.

New Concepts from Analyzing the GRAPHRES.BIN Archive

When parsing an unknown binary archive, several new concepts come up repeatedly. These are worth recording explicitly because they appear in many reverse-engineering and binary-analysis tasks.

1. Index vs. Identifier

An index is a dense, sequential number: 0, 1, 2, 3.... It usually means "the Nth element in an array." An identifier is a semantic value chosen by the developer. It may have gaps, clusters, or hidden meaning. Game archives often use identifiers, not indices, because assets are referred to by stable engine IDs across files and across versions of the game.

# index: Nth slot in the table
0, 1, 2, 3, 4

# identifier: arbitrary engine-level IDs
5, 2579, 3469, 3475, 3536

A file table may contain identifiers, but the script still needs to sort them by file offset to extract the correct byte ranges.

2. Composite Identifiers (Bitfields)

A single 32-bit integer can encode multiple pieces of information. This is called a bitfield or composite key. For example, an ID like 0x00000A5D might mean:

upper 16 bits 0x0A = category (textures, models, sounds)
lower 16 bits 0x5D = index within that category

To test this hypothesis, print IDs in hex and look for shared upper or lower patterns:

for e in parsed.entries[:50]:
    print(f"entry_id={e.entry_id:5}  0x{e.entry_id:08X}  offset=0x{e.offset:08X}")

3. Offset-Based Slicing

When the archive does not store explicit sizes, the size of each asset is inferred from the distance between consecutive file offsets. This is why the table must be sorted by offset before slicing.

entries = sorted(parsed.entries, key=lambda e: e.offset)

for i, entry in enumerate(entries):
    if i < len(entries) - 1:
        next_offset = entries[i + 1].offset
    else:
        next_offset = len(file_data)

    size = next_offset - entry.offset
    asset = file_data[entry.offset : entry.offset + size]

This works only when assets are packed back-to-back with no gaps.

4. Heuristic Type Detection with Magic Bytes

Magic bytes are the first few bytes of a file that identify its format. After extracting an asset, checking the header can guess the format:

  • b"TIM2" = PlayStation 2 texture
  • b"RIFF" = WAV or other RIFF container
  • b"\x1f\x8b" = gzip
  • b"PK\x03\x04" = ZIP

Magic bytes are a first-pass heuristic. They are not reliable when the asset has no signature, is compressed, or starts with a size/flag field instead of a magic.

5. When Parsing Looks Wrong: Debugging the Layout

If every parsed entry looks identical, the declared layout is probably wrong. Common causes:

  • The script file is named after a module it imports (e.g., construct.py), causing a circular import.
  • The header size or entry size is wrong.
  • The file is not what the script expects.

A first debug step is always to print the raw header bytes and the total file size:

with open("_GRAPHRES.BIN", "rb") as f:
    data = f.read()

print(f"File size: {len(data)} bytes")
print(f"First 16 bytes: {data[:16].hex()}")
print(f"First 16 bytes as ASCII: {data[:16]}")

Then compare the hex dump to the construct struct. If the byte count or field order does not match, the struct is wrong.

6. Entropy as a Quick Analysis Tool

Entropy measures how random a blob of bytes looks. It helps decide whether an extracted file is compressed, encrypted, structured, or padding.

import math

def entropy(data):
    if not data:
        return 0.0
    counts = [0] * 256
    for b in data:
        counts[b] += 1
    ent = 0.0
    for c in counts:
        if c:
            p = c / len(data)
            ent -= p * math.log2(p)
    return ent
Entropy range Likely meaning
0.0 Single repeated byte (null padding)
4.0–5.0 Text, highly structured data
5.0–7.0 Code, textures, models
7.5–8.0 Compressed or encrypted data