The parts of a ZIP
Each file is stored as a local file header (signature 50 4B 03 04, “PK” and two
bytes, then name, compression method, sizes and CRC-32) followed by its data. After the last
file comes the central directory, one record per file (50 4B 01 02), and at the
very end the end of central directory record (50 4B 05 06).
| Offset and length | Part | Contents |
|---|---|---|
| 0 0x0000 512 bytes | Before the archive | Data before the archive |
| 512 0x0200 75 bytes | File | Local header and data of readme.txt |
| 587 0x024B 84 bytes | File | Local header and data of docs/notes.txt |
| 671 0x029F 116 bytes | Central directory | Central directory: 2 entries |
| 787 0x0313 22 bytes | End record | End of central directory record |
Why readers start at the end
The end of central directory record is 22 bytes long, plus an optional archive comment of up to 65,535 bytes (its length is a 16-bit field). A ZIP has exactly one. A reader therefore searches the last 22 + 65,535 bytes of the file for its signature, reads where the central directory starts and how large it is, and reads that directory. The file data is not needed to know what the archive contains.
The four signature bytes can also occur inside file data or inside the comment. UNQIRO first looks in the last 4 KB for a record whose comment ends exactly at the end of the file, then in the full range, and accepts a candidate only if it leads to a real central directory with the stated number of records. A signature that does not lead anywhere is not taken.
Measured on an example archive of 400 stored files (3.32 MB): UNQIRO read 30.7 KB to list all of them.
The central directory is the list
Every central directory record repeats the file’s name, method, sizes and CRC-32 and adds the offset of its local header. The local headers carry most of the same values again. When a writer did not know the sizes in advance, it sets bit 3 of the flags, writes zeros in the local header and puts the real values in a data descriptor after the data; the central directory still has them.
UNQIRO lists an archive from its central directory. Before it reads a file’s data, it checks that the local header at the stated offset has the right signature, method, encryption flag and name, and that the data does not run into the next entry. Entries whose data would overlap another entry are marked damaged: overlapping entries are how non-recursive ZIP bombs make a small file expand into a huge one.
Data before the archive: self-extracting ZIPs
A self-extracting ZIP is a program followed by a ZIP archive. APPNOTE counts offsets from the
start of the disk, which for a single file is the start of the file, so the stored offsets
should include the program. When a program is simply put in front of an existing ZIP, they do
not: they still count from where the archive starts. Info-ZIP’s zip -A adjusts
them afterwards.
A reader that starts at the end copes with both. The central directory ends where the end record begins, so its real position follows from its size, and the difference to the stored offset is the length of the data in front. In the example above, the first file’s local header is stored at offset 0 and found at 512; the directory is stored at 159 and found at 671.
Info-ZIP UnZip reports such a difference as “extra bytes at beginning or within zipfile” and continues. The warning means only that the end record sits later than the stored offsets say, so it also appears for archives that are not self-extracting.
Data before the archive is not damage. UNQIRO shows it in the archive details as “ Data before the archive” when the offsets were not adjusted; with adjusted offsets the archive is listed normally and no data in front is reported. The data in front is never run, and a listed archive says nothing about whether that program is safe.
Split archives are different
A split archive is one ZIP in several files (.z01, .z02, … and the
.zip last). Its central directory is in the last part, its first part begins with
the split marker 50 4B 07 08, and offsets count from the start of each part.
UNQIRO recognises the first and the last part and says so; it does not open split archives.
How to tell this from other failures is explained in
Why a ZIP file won’t open.
Large archives: ZIP64 in brief
When sizes, offsets or the number of entries do not fit the classic fields, the end record holds placeholder values and two more records sit right before it: the ZIP64 end record ( 56 bytes) and its locator (20 bytes). The details, and what “4 GB” really means, are in ZIP64: the real 4 GB and 65,535-file limits.
What UNQIRO checks
- End record candidates must lead to a central directory with the stated number of records.
- Each local header is compared with its central record before its data is read, and data may not run into the next entry or overlap another.
- Declared sizes that cannot be true are refused: stored files whose two sizes differ, and Deflate data that would expand more than about 1,032 times.
- A saved file must produce exactly its declared size and match its CRC-32; output beyond the declared size is stopped at once.
-
Paths with
.., absolute paths, drive letters and control characters are listed but never extracted. - Stored and Deflate data is extracted; every other method is named. Which methods programs support is explained in ZIP compression methods and encryption.
Check your own file
The ZIP tool lists your archive in the browser from its central directory. Its details show the format, the size of the file list, the compression methods, data before the archive and how much of the file was read. Nothing is extracted until you save a file, and nothing is uploaded.
Inspect a ZIP without extracting it (View ZIP Contents)