Current - Issue

Year 2026 · Volume 5 · Issue 2

Original Article

The Effects of Low-Bit Quantization on Confidence Calibration in Small Language Models

Iddhant Gupta1
1 Department, Credence High School, UAE.

Published Online: May-August 2026

Pages: 938-950

References

1. Dettmers, T., & Zettlemoyer, L. (2022). The case for 4-bit precision: k,-bit inference scaling laws (arXiv:2212.09720). arXiv.
https://arxiv.org/abs/2212.09720
2. Frantar, E., Ashkboos, S., Stock, P., & Alistarh, D. (2022). GPTQ: Accurate post-training quantization for generative pre-trained transformers
(GPTs). In Proceedings of the 11th International Conference on Learning Representations (ICLR). OpenReview.
https://arxiv.org/abs/2210.17323
3. Zhong, M., Chen, L., Wang, Y., Zhang, R., & Zhang, M. (2025). Quantized can still be calibrated: A unified framework to calibration in
quantized large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025).
Association for Computational Linguistics.
4. https://aclanthology.org/2025.acl,long.1473
5. Hubara, I., Nahshan, Y., Hanani, Y., Banner, R., Soudry, D., & Yang, H. (2021). Accurate post training quantization with small calibration
sets. In Proceedings of the 38th International Conference on Machine Learning (ICML 2021). PMLR.
6. https://proceedings.mlr.press/v139/hubara21a.html
7. Nagel, M., Amjad, R. A., Van Baalen, M., Louizos, C., & Blankevoort, T. (2021). An underexplored dilemma between confidence and
calibration in quantized neural networks. In Proceedings of the 9th International Conference on Learning Representations (ICLR 2021).
OpenReview. https://arxiv.org/abs/2111.08163
8. Tomani, C., Buettner, M., Hu, W., & Cremers, D. (2022). Parameterized temperature scaling for boosting the expressive power in post hoc
uncertainty calibration. In Computer Vision – ECCV 2022 (Lecture Notes in Computer Science). Springer.
https://www.ecva.net/papers/eccv_2022/papers_ECCV/papers/136730554.pdf
9. Zhong, M., Chen, L., Wang, Y., Zhang, R., & Zhang, M. (2025). Quantized can still be calibrated: A unified framework to calibration in
quantized large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL 2025).Association for Computational Linguistics.
10. https://aclanthology.org/2025.acl,long.1473
11. The Case for 4-bit Precision: k-bit Inference Scaling Laws – Dettmers & Zettlemoyer, arXiv 2022
12. GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers – Frantar et al., ICLR 2023
13. Secure or Suspect? Investigating Package Hallucinations in Quantized Code Models – 2025 arXiv https://arxiv.org/html/2512.08213v1
14. Ultimate Guide to LLM Quantization for Faster, Leaner AI Models – Lamatic AI 2024 https://labs.lamatic.ai/p/llm,quantization/
15. General LLM calibration papers (ECE, over/under-confidence)
16. https://aclanthology.org/2025.acl,long.1473/
17. OpenAI. (2025). Overconfidence and hallucinations in large language models. OpenAI Research Blog.
18. https://arxiv.org/abs/2509.04664
19. Sliced-Wasserstein Distribution Alignment Loss Improves the Ultra-Low-Bit Quantization of Large Language Models
20. https://arxiv.org/abs/2601.07878
21. CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs https://arxiv.org/abs/2405.17233
22. SDQ-LLM: Sigma-Delta Quantization for 1-bit LLMs of any size
23. https://arxiv.org/abs/2510.03275
24. End-to-End On-Device Quantization-Aware Training for LLMs at Inference Cost https://arxiv.org/abs/2509.00031
25. xviii. ICQuant: Index Coding enables Low,bit LLM Quantization
26. https://arxiv.org/abs/2505.00850
27. xix. Enhancing Ultra,Low,Bit Quantization of Large Language Models Through Saliency,Aware Partial Retraining
https://arxiv.org/abs/2504.13932
28. xx. Liu, J., Gong, R., Wei, X., Dong, Z., Cai, J., & Zhuang, B. (2023). QLLM: Accurate and Efficient Low,Bitwidth Quantization for Large
Language Models.
29. xxi. Huang, W., Ma, X., Qin, H., Zheng, X., Lv, C., Chen, H., Luo, J., Qi, X., Liu, X., & Magno, M. (2024). How Good Are Low,bit Quantized
LLaMA3 Models? An Empirical Study. xxii. Fasoli, A., Chen, C. Y., Serrano, M., Sun, X., Wang, N., Venkataramani, S., Saon, G., Cui, X.,
Kingsbury, B., Zhang, W., Tüske, Z., & Gopalakrishnan, K. (2021). 4,bit Quantization of LSTM,based Speech Recognition Models.
30. xxiii. Gupta, K. (2022). Towards Efficient and Reliable Deep Neural Networks. xxiv. Liu, J., Gong, R., Wei, X., Dong, Z., Cai, J., & Zhuang,
B. (2023). QLLM: Accurate and Efficient Low,Bitwidth Quantization for Large Language Models

Related Articles

2026

Artificial Intelligence in Learning and Teaching

2026

Admin Assist: An AI – Driven Configuration and Orchestration for Enterprise Application

2026

Enhancing Blood Group Identification using pigeon inspired optimization: An Innovative Approach

2026

Eco-Genius: Power Up Smart, Power Down Waste

2026

Crowd-Sourced Disaster Response and Rescue Assistant

2026

Unveiling Deepfake Detection Using Vision Transformers: A Survey and Experimental Study

Share Article

X
LinkedIn
Facebook
WhatsApp

Or copy link

https://www.indjcst.com/archives/the-effects-of-low-bit-quantization-on-confidence-calibration-in-small-language-models

*Instagram doesn't support direct link sharing from web. Copy the link and share it in your Instagram story or post.