NanoQuant: Efficient Sub-1-Bit Quantization of Large Language Models
International Conference on Machine Learning (ICML),
NanoQuant uses low-rank binary factorization for sub-1-bit LLM post-training quantization. It compresses Llama 2 70B by 25.8× in 13 hours on one H100, enabling inference on a consumer 8 GB GPU.