Error during quantization
#1
by thirteenbit - opened
Hi!
Trying to quantize this model using llama.cpp's llama-quantize and getting NaNs error:
$ hf download SicariusSicariiStuff/Wingless_Imp_8B_Abliterated
$ rm -rf ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF
$ mkdir -p ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF
$ python ./convert_hf_to_gguf.py \
--outfile ./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-{ftype}.gguf \
--outtype auto \
--model-name Wingless_Imp_8B_Abliterated \
--verbose \
~/.cache/huggingface/hub/models--SicariusSicariiStuff--Wingless_Imp_8B_Abliterated/snapshots/9f0cbc2cc35967302e833baddce46d1ab023a5b0
$ ~/src/llama.cpp/build/bin/llama-quantize \
./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-bf16.gguf \
./models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-Q8_0.gguf \
Q8_0
...
[ 45/ 292] blk.4.attn_v.weight - [ 4096, 1024, 1, 1], type = bf16, converting to q8_0 .. size = 8.00 MiB -> 4.25 MiB
[ 46/ 292] blk.4.ffn_down.weight - [ 14336, 4096, 1, 1], type = bf16, converting to q8_0 .. size = 112.00 MiB -> 59.50 MiB
[ 47/ 292] blk.4.ffn_gate.weight - [ 4096, 14336, 1, 1], type = bf16, converting to q8_0 .. size = 112.00 MiB -> 59.50 MiB
[ 48/ 292] blk.4.ffn_norm.weight - [ 4096, 1, 1, 1], type = f32, size = 0.016 MiB
ggml_validate_row_data: found 3 NaNs in row of 58720256 BF16 values
llama_model_quantize: failed to quantize: tensor 'blk.4.ffn_up.weight' has invalid data
main: failed to quantize model from './models/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF/SicariusSicariiStuff_Wingless_Imp_8B_Abliterated-GGUF-bf16.gguf'
Do you know if this is the bug in llama.cpp (either convert_hf_to_gguf.py or llama-quantize) or in the model weights?
Depending on that I'll either report a bug in llama.cpp repository or ask for your help and post full logs here (if this helps troubleshooting).
Looked at some old closed similar llama.cpp issues and tried:
- using
--outtype f32forconvert_hf_to_gguf.py: same error - updating
numpyinconvert_hf_to_gguf.pyvenv: same error
hmmm seems like a corruption, thanks for the catch, will test and reupload if needed.
there was a bug with uploads some time ago, could be related✍️
reuploaded, please let me know if the issue is resolved
Thank you, this is fixed.
thirteenbit changed discussion status to closed