Which fields are limited
| Field | Where | Largest classic value | In ZIP64 |
|---|---|---|---|
| Uncompressed size of a file | Local and central header | 4,294,967,295 0xFFFFFFFF | ZIP64 extra field (0x0001) |
| Compressed size of a file | Local and central header | 4,294,967,295 0xFFFFFFFF | ZIP64 extra field (0x0001) |
| Offset of a file’s local header | Central header | 4,294,967,295 0xFFFFFFFF | ZIP64 extra field (0x0001) |
| Number of entries | End of central directory record | 65,535 0xFFFF | ZIP64 end record |
| Size of the central directory | End of central directory record | 4,294,967,295 0xFFFFFFFF | ZIP64 end record |
| Offset of the central directory | End of central directory record | 4,294,967,295 0xFFFFFFFF | ZIP64 end record |
Because the largest value doubles as the marker, a file of exactly 4,294,967,295 bytes should already use ZIP64: readers take that value as “look in ZIP64”. In the same way, 65,535 is the largest count a classic end record can state, and a reader that sees 0xFFFF may look for ZIP64 records. UNQIRO opens an archive of exactly 65,535 entries that has none.
What “4 GB” really means
“A ZIP cannot be larger than 4 GB” is too coarse. The limits apply to single fields, and each can be reached on its own:
- One 5 GB file needs ZIP64 for its size, even if the archive holds nothing else.
- A thousand files of 10 MB each need ZIP64 although none is large: the later local headers and the central directory start beyond 4 GiB.
- 70,000 small files need ZIP64 for the entry count, even if the archive is only a few megabytes.
If a writer puts more than 65,535 entries into a classic archive without ZIP64, the 16-bit count wraps around. UNQIRO lists such archives completely by reading the central directory to its end and checking the count modulo 65,536.
APPNOTE also allows ZIP64 for small files. That an archive uses ZIP64 does not mean it needed to.
The ZIP64 records
For each file whose size or header offset does not fit, the central directory record holds 0xFFFFFFFF and a ZIP64 extra field with the ID 0x0001 carries the real 64-bit values, in a fixed order: uncompressed size, compressed size, header offset, disk number. Only the values whose classic field holds the marker appear. At the end, a 56-byte ZIP64 end record carries the entry counts and the directory’s size and offset, and a 20-byte locator just before the end record says where that record is. The layout of the classic records around them is explained in How a ZIP file is read.
| Offset and length | Part | Contents |
|---|---|---|
| 87 83 bytes | Central directory | Sizes and offset hold the marker; the real values follow in the 0x0001 extra field |
| 170 56 bytes | ZIP64 end record | 64-bit entry counts, directory size and directory offset |
| 226 20 bytes | ZIP64 locator | Where the ZIP64 end record is, and the total number of disks |
| 246 22 bytes | End record | Counts 0xFFFF, directory size and offset 0xFFFFFFFF |
UNQIRO reads this archive as ZIP64: its file large.bin has an uncompressed size of 28 bytes and its local header at offset 0, values it took from the extra field because the classic fields hold the marker. It reads a ZIP64 field only where a marker asks for it, so an empty or unneeded 0x0001 field does not stop it. The locator is expected 20 bytes before the end record; when data before the archive shifted the stated offset, UNQIRO also looks right before the locator.
Old programs and broken archives
Two different problems look the same to a user. A valid ZIP64 archive fails in a program without ZIP64 support: Info-ZIP UnZip checks the “version needed to extract”, which is 4.5 for ZIP64, and a build without ZIP64 skips such files with a message like “need PK compat. v4.5 (can do v2.1)”. And a writer can produce a broken archive: one analysis of OneDrive exports larger than 4 GiB found a locator that says 0 disks instead of 1, which several readers refuse. UNQIRO accepts 0 there and treats more than one disk as a split archive.
An archive that needed ZIP64 but was written without it cannot be read reliably by anyone: its sizes or offsets no longer fit the fields they were written into. Why a ZIP won’t open in general is explained in Why a ZIP file won’t open.
Java: “Invalid CEN header (invalid zip64 extra data field size)”
Oracle’s JDK updates 11.0.20, 17.0.8, 20.0.2 and 21 (July 2023) made
java.util.zip.ZipFile validate ZIP64 extra fields more strictly. Archives from
some build tools with empty ZIP64 extra fields were then rejected with this message; the JDK
21 release notes name fixes in Apache Commons Compress, Ant and BND. The check can be switched
off for ZipFile with the system property
jdk.util.zip.disableZip64ExtraFieldValidation=true, not for the ZIP file system
provider. Oracle lists a fix for this error, JDK-8313765, in later updates such as 17.0.8.0.2.
The message comes from one Java component and from those releases; it says nothing about other programs. UNQIRO ignores empty ZIP64 extra fields, as described above, and does not predict what a given Java version will do.
Check your own file
The ZIP tool reads the end of your archive in the browser. Its details show “ ZIP64 (large archive format)” when the archive uses the ZIP64 end records, and how large its file list is. Nothing is uploaded.
Check whether this archive uses ZIP64 (View ZIP Contents)
Sources
- PKWARE APPNOTE.TXT 6.3.10: .ZIP File Format Specification
- Oracle JDK 17.0.8 release notes: improved ZIP64 extra field validation
- Oracle JDK 21 release notes: known issue JDK-8313765 (ZIP64 extra fields)
- OpenJDK source: java.util.zip.ZipFile
- Python documentation: zipfile (3.14)
- Info-ZIP UnZip 6.0 source: extract.c (method and version messages)
- bitsgalore: OneDrive ZIP exports over 4 GiB and the ZIP64 locator