Moving a 20 GB Model Through a 2 GiB Limit
A quantized model is one file until you try to move it. Then the first question is not about the model. It is about the weakest limit on the path.
For the Intern-S2 ternary builds, the path ran through a GitHub release. GitHub's documentation says each file in a release must be under 2 GiB, which is 2,147,483,648 bytes. There is no limit on the total size of a release. The limit is per file, so the answer is to split.
The first split
The corrected TQ2_0 parent model totals 10,827,311,232 bytes. Divide that by six and each part is 1,804,551,872 bytes. That is about 1.68 GiB, under the limit with room to spare. The README of intern-s2-tq2 still describes this layout: parent and feral models each split into six parts so every asset stays inside GitHub's 2 GiB limit.
The second split
The current release lists 24 parts of 451,137,968 bytes each, about 430 MiB. The 4-bit build is also 24 parts, of 836,169,021 bytes each, about 797 MiB. Parts this small are not required by the limit. They do have a practical benefit: each downloads independently, and a damaged part costs 430 MiB to fetch again rather than 1.7 GiB.
Make reassembly boring
A good rule is that anyone should rebuild the file with one shell command and no special tool. The parts are named in sorted order, parent-part-00 through parent-part-23, so cat parent-part-*.bin in a shell glob concatenates them in the right sequence. Zero-padded numbers matter: without them, part 10 sorts before part 2.
- Name parts with fixed-width numbers so sorting is the order.
- Keep part sizes equal so a missing byte is obvious from the listing.
- State the reassembly command in the README, next to the files.
- Ship a checksum for the final file, since cat will happily join a damaged part.
What is missing
The release lists no checksum file. A concatenation that finishes without error says nothing about whether the bytes are right, so a checksum for the final GGUF is the missing piece.
The release also carries a tarball of a patched llama.cpp build, 78,967,256 bytes. A model file is only usable with a runtime that understands its format, so shipping the runtime beside the weights removes one more step between a reader and a running model.