Error during quantization

#1
by thirteenbit - opened

Hi!

Trying to quantize this model using llama.cpp's llama-quantize and getting NaNs error:

$ hf download SicariusSicariiStuff/Wingless_Imp_8B_Abliterated
$ rm -rf ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF
$ mkdir -p ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF
$ python ./convert_hf_to_gguf.py \
 --outfile ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-{ftype}.gguf \
 --outtype auto \
 --model-name Wingless_Imp_8B_Abliterated \
 --verbose \
 ~/.cache/huggingface/hub/models--SicariusSicariiStuff--Wingless_Imp_8B_Abliterated/snapshots/9f0cbc2cc35967302e833baddce46d1ab023a5b0
$ ~/src/llama.cpp/build/bin/llama-quantize \
 ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-bf16.gguf \
 ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-Q8_0.gguf \
 Q8_0
...
[  45/ 292] blk.4.attn_v.weight                  - [  4096,   1024,      1,      1], type =   bf16, converting to q8_0 .. size =     8.00 MiB ->     4.25 MiB
[  46/ 292] blk.4.ffn_down.weight                - [ 14336,   4096,      1,      1], type =   bf16, converting to q8_0 .. size =   112.00 MiB ->    59.50 MiB
[  47/ 292] blk.4.ffn_gate.weight                - [  4096,  14336,      1,      1], type =   bf16, converting to q8_0 .. size =   112.00 MiB ->    59.50 MiB
[  48/ 292] blk.4.ffn_norm.weight                - [  4096,      1,      1,      1], type =    f32, size =    0.016 MiB
ggml_validate_row_data: found 3 NaNs in row of 58720256 BF16 values
llama_model_quantize: failed to quantize: tensor 'blk.4.ffn_up.weight' has invalid data
main: failed to quantize model from './models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-bf16.gguf'

Do you know if this is the bug in llama.cpp (either convert_hf_to_gguf.py or llama-quantize) or in the model weights?

Depending on that I'll either report a bug in llama.cpp repository or ask for your help and post full logs here (if this helps troubleshooting).

Looked at some old closed similar llama.cpp issues and tried:

  • using --outtype f32 for convert_hf_to_gguf.py: same error
  • updating numpy in convert_hf_to_gguf.py venv: same error

hmmm seems like a corruption, thanks for the catch, will test and reupload if needed.
there was a bug with uploads some time ago, could be related✍️

reuploaded, please let me know if the issue is resolved

Thank you, this is fixed.

thirteenbit changed discussion status to closed

Sign up or log in to comment