"Beats Q8_0 perplexity at half the size -- and even beats F16"
I'm sure you know that perplexity pales against KLD for this kind of comparison, so claiming better than Q8 seems pretty disingenuous.
I don't think I need to explain how claiming the quantization is better than the full precision model is kind of a mathematically insane claim.
If you quantize a model, and then the quantized model is better at one benchmark than the base model, that's still quantization error and it has most likely regressed somewhere else.
Apex Quality has ~2.5x worse KLD than Q8_0, you should compare it to other bit sizes for a better size/accuracy comparison.
Apex Quality has ~4.5x worse KLD than Unsloth UD-Q8_K_XL
"Beats Q8_0 perplexity at half the size -- and even beats F16"
I'm sure you know that perplexity pales against KLD for this kind of comparison, so claiming better than Q8 seems pretty disingenuous.
I don't think I need to explain how claiming the quantization is better than the full precision model is kind of a mathematically insane claim.
If you quantize a model, and then the quantized model is better at one benchmark than the base model, that's still quantization error and it has most likely regressed somewhere else.
Apex Quality has ~2.5x worse KLD than Q8_0, you should compare it to other bit sizes for a better size/accuracy comparison.
Apex Quality has ~4.5x worse KLD than Unsloth UD-Q8_K_XL