zackliqcom commited on
Commit
aafdf65
·
verified ·
1 Parent(s): bc202db

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +44 -45
README.md CHANGED
@@ -51,61 +51,60 @@ build/bin/llama-quantize --tensor-type-file <override-file> \
51
  model-f16.gguf model-q4_0-override.gguf q4_0
52
  ```
53
 
 
54
 
55
- ## Performance Metrics
56
- Performance on IQ9 (QCS9075M)
57
 
58
- | Compute | CTX | TTFT (ms) | Prefill (tok/s) | Decode (tok/s) |
59
- |---------|----:|----------:|----------------:|---------------:|
60
- | CPU | 512 | 3221.0 ± 248.1 | 109.2 ± 9.4 | 19.7 ± 3.0 |
61
- | CPU | 1024 | 7820.4 ± 245.5 | 117.1 ± 3.6 | 22.3 ± 0.4 |
62
- | CPU | 4096 | 38454.0 ± 28.2 | 101.3 ± 0.1 | 19.4 ± 0.1 |
63
- | GPU | 512 | 3059.0 ± 0.7 | 114.4 ± 0.0 | 11.3 ± 0.0 |
64
- | GPU | 1024 | 8386.1 ± 0.8 | 109.1 ± 0.0 | 10.0 ± 0.0 |
65
- | GPU | 4096 | 40601.2 ± 1.6 | 96.0 ± 0.0 | 7.6 ± 0.0 |
66
- | HTP | 512 | 703.4 ± 1.3 | 497.6 ± 0.9 | 16.4 ± 0.2 |
67
- | HTP | 1024 | 1827.0 ± 2.6 | 500.8 ± 0.7 | 16.1 ± 0.2 |
68
- | HTP | 4096 | 8110.8 ± 15.1 | 480.5 ± 0.9 | 15.7 ± 0.2 |
69
 
70
- ## Accuracy Metrics
 
 
 
 
 
 
 
 
 
71
 
72
- It is measured by 100-prompt suite (`text_prompts.yaml`) from AIMET, one `llama-cli` single-turn per prompt, native gemma chat template, reasoning on, natural EOS, `-n 2048`, `--seed 42`. Only the **first assistant turn** is scored (the `[Start thinking]…[End thinking]` block is excluded; the final answer after it is judged). Per-prompt invocation, Scored by Claude (LLM-as-judge) on below:
73
 
74
- | Score | Meaning |
75
- |---|---|
76
- | **0** | Complete gibberish |
77
- | **2** | Catastrophic failure (repeated words, gibberish words) |
78
- | **7** | Sound grammar, but strange/incoherent content |
79
- | **10** | No issues |
80
 
 
 
 
 
 
 
 
 
81
 
82
- <!-- QUALITY_BEGIN -->
83
- | Model (backend) | n | mean | median | 0 | 2 | 7 | 10 |
84
- | ----- | --: | ---: | -----: | -: | -: | -: | --- |
85
- | [Unsloth Q4_0](https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF) | 100 | 9.86 | 10.0 | 0 | 1 | 2 | 97 |
86
- | [Google's Q4_0 qat](https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-gguf) | 100 | 9.83 | 10.0 | 0 | 1 | 3 | 96 |
87
- | Our Q4_0 | 100 | 9.75 | 10.0 | 0 | 2 | 3 | 95 |
88
- <!-- QUALITY_END -->
89
 
90
- Also, the MMLU-Pro is measured:
91
 
92
- | Subject | Google QAT GGUF | Unsloth PTQ GGUF | Ours |
93
  | --- | ---: | ---: | ---: |
94
- | **mmlu_pro** | **0.4178** | **0.4084** | **0.4131** |
95
- | biology | 0.6262 | 0.6248 ± 0.0181 | 0.6248 ± 0.0181 |
96
- | business | 0.5019 | 0.4956 ± 0.0178 | 0.4740 ± 0.0178 |
97
- | chemistry | 0.3127 | 0.3021 ± 0.0137 | 0.3004 ± 0.0136 |
98
- | computer_science | 0.4854 | 0.4683 ± 0.0247 | 0.4683 ± 0.0247 |
99
- | economics | 0.4976 | 0.4538 ± 0.0171 | 0.4431 ± 0.0171 |
100
- | engineering | 0.3375 | 0.2859 ± 0.0145 | 0.2931 ± 0.0146 |
101
- | health | 0.3423 | 0.3961 ± 0.0171 | 0.3985 ± 0.0171 |
102
- | history | 0.2310 | 0.2730 ± 0.0229 | 0.3097 ± 0.0237 |
103
- | law | 0.2289 | 0.2416 ± 0.0129 | 0.2480 ± 0.0130 |
104
- | math | 0.6366 | 0.6366 ± 0.0131 | 0.6373 ± 0.0131 |
105
- | other | 0.3528 | 0.3561 ± 0.0158 | 0.3431 ± 0.0156 |
106
- | philosophy | 0.3848 | 0.3387 ± 0.0212 | 0.3487 ± 0.0214 |
107
- | physics | 0.4003 | 0.3580 ± 0.0133 | 0.3926 ± 0.0136 |
108
- | psychology | 0.5113 | 0.4875 ± 0.0177 | 0.5013 ± 0.0177 |
109
 
110
  ## License
111
 
 
51
  model-f16.gguf model-q4_0-override.gguf q4_0
52
  ```
53
 
54
+ ## Performance Measurement Commands
55
 
56
+ CPU uses `--device none -ngl 0`; HTP uses `--device HTP0 -ngl 99`. For each (model, backend, CTX ∈ {512, 1024, 4096}) two `llama-bench` runs were issued — one for prefill, one for decode:
 
57
 
58
+ ```sh
59
+ # environment on device
60
+ export LD_LIBRARY_PATH=./lib
61
+ export ADSP_LIBRARY_PATH=./lib
 
 
 
 
 
 
 
62
 
63
+ # Prefill (Prefill tok/s; TTFT = CTX / Prefill × 1000)
64
+ ./bin/llama-bench --device <none|HTP0> -m <model.gguf> \
65
+ --poll 1000 -t 6 --cpu-mask 0xfc --cpu-strict 1 --ubatch-size 1024 -fa on \
66
+ -ngl <0|99> -p <CTX> -n 0
67
+
68
+ # Decode at depth = CTX (Decode tok/s)
69
+ ./bin/llama-bench --device <none|HTP0> -m <model.gguf> \
70
+ --poll 1000 -t 6 --cpu-mask 0xfc --cpu-strict 1 --ubatch-size 1024 -fa on \
71
+ -ngl <0|99> -p 0 -n 128 -d <CTX>
72
+ ```
73
 
 
74
 
75
+ ## Performance Metrics
76
+ Performance on IQ9 (QCS9075M)
 
 
 
 
77
 
78
+ | Compute | CTX | Unsloth PTQ GGUF | Google QAT GGUF | Ours |
79
+ |---------|----:|:-------:|:------:|:----:|
80
+ | CPU | 512 | 82.2 / 22.25 | 107.2 / 22.51 | 98.9 / **23.78** |
81
+ | CPU | 1024 | 76.6 / 21.65 | 97.8 / 21.94 | 90.8 / **23.16** |
82
+ | CPU | 4096 | 61.3 / 18.09 | 76.5 / 18.26 | 72.3 / **19.10** |
83
+ | HTP | 512 | 273.2 / 19.22 | 582.5 / 16.40 | 582.4 / **18.33** |
84
+ | HTP | 1024 | 266.9 / 18.93 | 552.3 / 16.45 | 552.2 / **18.17** |
85
+ | HTP | 4096 | 254.9 / 18.30 | 502.5 / 15.87 | 501.8 / **17.53** |
86
 
87
+ ## Accuracy Metrics
 
 
 
 
 
 
88
 
89
+ The MMLU-Pro is measured:
90
 
91
+ | Subject | Unsloth PTQ GGUF | Google QAT GGUF | Ours |
92
  | --- | ---: | ---: | ---: |
93
+ | **mmlu_pro** | **0.4084** | **0.4178** | **0.4131** |
94
+ | biology | 0.6248 ± 0.0181 | 0.6262 | 0.6248 ± 0.0181 |
95
+ | business | 0.4956 ± 0.0178 | 0.5019 | 0.4740 ± 0.0178 |
96
+ | chemistry | 0.3021 ± 0.0137 | 0.3127 | 0.3004 ± 0.0136 |
97
+ | computer_science | 0.4683 ± 0.0247 | 0.4854 | 0.4683 ± 0.0247 |
98
+ | economics | 0.4538 ± 0.0171 | 0.4976 | 0.4431 ± 0.0171 |
99
+ | engineering | 0.2859 ± 0.0145 | 0.3375 | 0.2931 ± 0.0146 |
100
+ | health | 0.3961 ± 0.0171 | 0.3423 | 0.3985 ± 0.0171 |
101
+ | history | 0.2730 ± 0.0229 | 0.2310 | 0.3097 ± 0.0237 |
102
+ | law | 0.2416 ± 0.0129 | 0.2289 | 0.2480 ± 0.0130 |
103
+ | math | 0.6366 ± 0.0131 | 0.6366 | 0.6373 ± 0.0131 |
104
+ | other | 0.3561 ± 0.0158 | 0.3528 | 0.3431 ± 0.0156 |
105
+ | philosophy | 0.3387 ± 0.0212 | 0.3848 | 0.3487 ± 0.0214 |
106
+ | physics | 0.3580 ± 0.0133 | 0.4003 | 0.3926 ± 0.0136 |
107
+ | psychology | 0.4875 ± 0.0177 | 0.5113 | 0.5013 ± 0.0177 |
108
 
109
  ## License
110