跳到正文
原文
Hugging Face Blog·· 2026-08-25精选AI 评分62

Quantization-Aware Healing 论文:4-bit 压缩模型在 7/9 基准上超越其全精度原版

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

AI 导读

Multiverse Computing 发布论文《Quantization-Aware Healing》,提出从压缩前原模型直接蒸馏的修复方法,应用于 GPT-OSS 120B 压缩至 60B 并量化到 MXFP4 后,在 7/9 基准上超越其 bfloat16 版本,其中长上下文推理 +7.4、数学 +5.6。

推荐理由

论文提出从压缩前原模型蒸馏的修复方法,让 4-bit 模型在多数基准上反超自身 bfloat16 版本,并对比了与 QAT 的稳定性差异。

来源:Hugging Face Blog · huggingface.co