Ternary Weights, 10.8 GB on Disk

A language model is mostly a very large pile of numbers. Making the pile smaller is the cheapest way to make the model easier to move, load, and run. The usual step down from 16-bit weights is a 4-bit format. I wanted to see what happens one step further, where every weight is one of three values: minus one, zero, or plus one.

The repository is intern-s2-tq2. It holds corrected TQ2_0 ternary builds of Intern-S2, post-trained with absmean scaling. The source model is internlm/Intern-S2-Preview-FP8, released under Apache-2.0. I did not train a ternary model from scratch. I took a finished model and ternarized it afterward.

What TQ2_0 actually is

TQ2_0 is a llama.cpp quantization type for ternary models. It was added alongside TQ1_0 in the llama.cpp pull request that introduced ternary packing for TriLMs and BitNet b1.58 models. TQ1_0 stores 1.6875 bits per weight. TQ2_0 stores 2.0625 bits per weight, which is two bits per weight plus a small per-block scale. The extra bits buy simpler, faster unpacking.

Absmean is the scaling rule from the BitNet b1.58 paper. Take the mean of the absolute values of a weight matrix, divide every weight by it, then round to the nearest of minus one, zero, and plus one. The scale is kept so the output can be multiplied back.

What the byte counts say

The corrected release has 24 files for the parent model, each exactly 451,137,968 bytes. Together that is 10,827,311,232 bytes, about 10.8 GB. My earlier Q4_0 build of the same parent is 20,068,056,512 bytes across 24 parts, about 20 GB, with 753 tensors.

So the ternary file is 54 percent of the size of the 4-bit file. That is a real saving. The bit widths alone, 2.0625 against the 4.5 of Q4_0, would predict about 9.2 GB. The file is 1.6 GB larger than that. I have not broken the gap down tensor by tensor.

What this does not show

File size is a fact. Model quality is a separate question, and I have not published an evaluation for these files. A smaller model that answers worse is not a win. The point of releasing the parts is that anyone can reassemble them, run their own benchmarks, and tell me where the ternary version breaks.

Post-training ternarization is also the harder route. Models trained as ternary from the start, as BitNet b1.58 was, get to adapt to three-level weights while they learn. A model trained at higher precision has to absorb the rounding error all at once. The release is tagged v1-corrected, and its notes call the files corrected TQ2_0 absmean models.

Reassembly

The files are plain binary parts. The README gives the rule: concatenate the parts in order with cat to rebuild the GGUF. No custom tool is needed.

The comparison that matters next is perplexity against the Q4_0 build on the same text.

← back to the journal