Reading a ZIP, step by step
The steps below are the order in which UNQIRO’s reader checks an archive. The first step that fails decides what it reports.
| Step | The reader needs | UNQIRO says |
|---|---|---|
| 1 | A ZIP: its records, or other data followed by one |
|
| 2 | The end of central directory record at the end of the file |
|
| 3 | One archive in one file |
|
| 4 | ZIP64 records, when the end record points to them |
|
| 5 | A central directory at its offset with the stated number of records |
|
| 6 | For each file: a method the program reads, and no encryption it lacks |
|
| 7 | For each file: a safe path and data where the list says |
|
| 8 | When a file is saved: exactly its size and its CRC-32 |
|
Is it a ZIP at all?
ZIP records begin with the bytes 50 4B, the letters “PK”. An ordinary ZIP starts
with a local file header; a self-extracting one starts with its program, with the ZIP after
it. Two cases are common: a web page was saved instead of the download (an error or sign-in
page), or another format was renamed to .zip, such as RAR, 7z or gzip.
UNQIRO recognises many formats from their first bytes, among them gzip, bzip2, xz, 7z, RAR,
Zstandard, PDF, programs and images, and names them. TAR and gzip archives (also
.tar.gz) open in the same archive workspace as a ZIP; the others are named, not
opened. It does not inspect web pages or text: a file named .html, .htm or .txt is named from its name, and a web page saved as .zip is reported as not recognised. Opened in a text editor, such a file starts
with <!DOCTYPE html> or <html.
Other programs tend to report all of this with one message. Python’s zipfile
raises “File is not a zip file” whenever it finds no end record; Info-ZIP UnZip says the
end-of-central-directory signature was not found and suggests the file is not a ZIP or is one
disk of a multi-part archive.
Is it complete?
The end record is the last thing in a ZIP, so a file that was cut short loses it first, for example after an interrupted download or a copy that stopped early. The data at the start may be intact, but the list that says what the archive contains is gone.
When a file starts like a ZIP but no end record ends at its end, UNQIRO says the end is missing. It does not treat the record signature found somewhere inside the data as the end. Downloading or copying the file again is the only real fix. A tool that scans local headers from the front can sometimes recover files that lie before the cut; nothing can recreate the missing part.
Is it one part of several?
A split archive is one ZIP in several files, usually .z01, .z02 …
and
.zip last. Only the last part has the central directory, and its end record names
a disk other than the first. The first part begins with the split marker
50 4B 07 08. Neither part can be opened alone. UNQIRO names both cases; it does
not open split archives. Python’s zipfile refuses archives that span multiple
disks as well.
Is the file list damaged?
Here the end record is in its place, but it does not lead to a usable central directory: the directory is not at the stated offset, a record is broken, or the number of records does not match. UNQIRO reports this apart from a missing end: the file is complete, but its list cannot be trusted. Data before the archive, as in a self-extracting ZIP, is not damage; how readers handle it is explained in how a ZIP file is read.
The list opens, but files fail
A valid archive can contain files a program cannot read: a compression method it does not
have, or encryption it does not support. Some programs then call the whole archive invalid.
Java’s java.util.zip.ZipFile already rejects encrypted entries and methods other
than Stored and Deflate when it opens the archive (“invalid CEN header”). Microsoft states
that Windows File Explorer does not support encrypted archives.
UNQIRO lists such files, names the method and marks encryption, and extracts the others. Which programs read which methods is explained in ZIP compression methods and encryption.
Large archives and older programs
An archive with files of 4 GiB or more, with data beyond 4 GiB or with more than 65,535 entries uses ZIP64 records. Programs without ZIP64 support fail on such archives even though they are valid, and archives written without the ZIP64 records they need are broken for every reader. ZIP64: the real 4 GB and 65,535-file limits tells the two apart.
Real damage, and what is refused on purpose
Damage inside a file shows when it is read: the data does not match its CRC-32, or it produces more or less than the declared size. UNQIRO stops at once and does not offer such a file; the other files stay available. Entries whose data would overlap another entry, as in a non-recursive ZIP bomb, are marked damaged before anything is read.
Entries whose names contain .., an absolute path, a drive letter or control
characters could be written outside the target folder. UNQIRO lists them as unsafe and never
extracts them. Nothing is repaired: the archive is read as it is.
Check your own file
The ZIP tool reads your archive in the browser and shows the step at which it fails, or lists it with every file’s method, encryption and problems marked. Nothing is uploaded and nothing is repaired or changed.
Inspect the archive structure (View ZIP Contents)