Researchers Evangelos Georganas, Alexander Heinecke, and Pradeep Dubey published a paper on September 14 detailing a method to surpass the 1.58-bit storage limit for ternary large language models (LLMs). Their work focuses on optimizing how ternary weights, which take values from the set {-1, 0, +1}, are stored more efficiently than the conventional 1.625 bits per weight format used in current deployments, according to arxiv.org.
The team analyzed the actual distribution of ternary symbols in LLMs, which are not equally probable, and developed a new encoding scheme that reduces the effective storage bit-width below the traditional information-theoretic limit of log2(3) ≈ 1.585 bits per weight. This approach improves upon the standard five-trit packing method, which packs five ternary weights into one byte but rounds up to 1.625 bits per weight, as explained in their paper on arxiv.org.
This advancement matters because ternary LLMs offer a promising trade-off between model size and computational efficiency, making them attractive for deployment in resource-constrained environments. By reducing storage requirements further, the new method could enable more compact and faster models compared to existing quantization techniques. The research contributes to ongoing efforts to optimize large language models, which are central to many AI applications today.
The paper titled "Breaking the 1.58-bit Barrier for Ternary LLMs" was submitted on September 14, 2026, and is available on arxiv.org, providing detailed technical insights and benchmarks demonstrating the improved storage efficiency achieved by the authors.