Instructions to use zackliqcom/gemma4-E2B-Q40-custom with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use zackliqcom/gemma4-E2B-Q40-custom with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("zackliqcom/gemma4-E2B-Q40-custom", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use zackliqcom/gemma4-E2B-Q40-custom with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0 # Run inference directly in the terminal: llama cli -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0 # Run inference directly in the terminal: llama cli -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0 # Run inference directly in the terminal: ./llama-cli -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Use Docker
docker model run hf.co/zackliqcom/gemma4-E2B-Q40-custom:Q4_0
- LM Studio
- Jan
- Ollama
How to use zackliqcom/gemma4-E2B-Q40-custom with Ollama:
ollama run hf.co/zackliqcom/gemma4-E2B-Q40-custom:Q4_0
- Unsloth Desktop
- Pi
How to use zackliqcom/gemma4-E2B-Q40-custom with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "zackliqcom/gemma4-E2B-Q40-custom:Q4_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use zackliqcom/gemma4-E2B-Q40-custom with Docker Model Runner:
docker model run hf.co/zackliqcom/gemma4-E2B-Q40-custom:Q4_0
- Lemonade
How to use zackliqcom/gemma4-E2B-Q40-custom with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Run and chat with the model
lemonade run user.gemma4-E2B-Q40-custom-Q4_0
List all available models
lemonade list
- Hermes Agent
How to use zackliqcom/gemma4-E2B-Q40-custom with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use zackliqcom/gemma4-E2B-Q40-custom with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf zackliqcom/gemma4-E2B-Q40-custom:Q4_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "zackliqcom/gemma4-E2B-Q40-custom:Q4_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Update README.md
Browse files
README.md
CHANGED
|
@@ -51,61 +51,60 @@ build/bin/llama-quantize --tensor-type-file <override-file> \
|
|
| 51 |
model-f16.gguf model-q4_0-override.gguf q4_0
|
| 52 |
```
|
| 53 |
|
|
|
|
| 54 |
|
| 55 |
-
|
| 56 |
-
Performance on IQ9 (QCS9075M)
|
| 57 |
|
| 58 |
-
|
| 59 |
-
|
| 60 |
-
|
| 61 |
-
|
| 62 |
-
| CPU | 4096 | 38454.0 ± 28.2 | 101.3 ± 0.1 | 19.4 ± 0.1 |
|
| 63 |
-
| GPU | 512 | 3059.0 ± 0.7 | 114.4 ± 0.0 | 11.3 ± 0.0 |
|
| 64 |
-
| GPU | 1024 | 8386.1 ± 0.8 | 109.1 ± 0.0 | 10.0 ± 0.0 |
|
| 65 |
-
| GPU | 4096 | 40601.2 ± 1.6 | 96.0 ± 0.0 | 7.6 ± 0.0 |
|
| 66 |
-
| HTP | 512 | 703.4 ± 1.3 | 497.6 ± 0.9 | 16.4 ± 0.2 |
|
| 67 |
-
| HTP | 1024 | 1827.0 ± 2.6 | 500.8 ± 0.7 | 16.1 ± 0.2 |
|
| 68 |
-
| HTP | 4096 | 8110.8 ± 15.1 | 480.5 ± 0.9 | 15.7 ± 0.2 |
|
| 69 |
|
| 70 |
-
#
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 71 |
|
| 72 |
-
It is measured by 100-prompt suite (`text_prompts.yaml`) from AIMET, one `llama-cli` single-turn per prompt, native gemma chat template, reasoning on, natural EOS, `-n 2048`, `--seed 42`. Only the **first assistant turn** is scored (the `[Start thinking]…[End thinking]` block is excluded; the final answer after it is judged). Per-prompt invocation, Scored by Claude (LLM-as-judge) on below:
|
| 73 |
|
| 74 |
-
|
| 75 |
-
|
| 76 |
-
| **0** | Complete gibberish |
|
| 77 |
-
| **2** | Catastrophic failure (repeated words, gibberish words) |
|
| 78 |
-
| **7** | Sound grammar, but strange/incoherent content |
|
| 79 |
-
| **10** | No issues |
|
| 80 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 81 |
|
| 82 |
-
|
| 83 |
-
| Model (backend) | n | mean | median | 0 | 2 | 7 | 10 |
|
| 84 |
-
| ----- | --: | ---: | -----: | -: | -: | -: | --- |
|
| 85 |
-
| [Unsloth Q4_0](https://huggingface.co/unsloth/gemma-4-E2B-it-GGUF) | 100 | 9.86 | 10.0 | 0 | 1 | 2 | 97 |
|
| 86 |
-
| [Google's Q4_0 qat](https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-gguf) | 100 | 9.83 | 10.0 | 0 | 1 | 3 | 96 |
|
| 87 |
-
| Our Q4_0 | 100 | 9.75 | 10.0 | 0 | 2 | 3 | 95 |
|
| 88 |
-
<!-- QUALITY_END -->
|
| 89 |
|
| 90 |
-
|
| 91 |
|
| 92 |
-
| Subject |
|
| 93 |
| --- | ---: | ---: | ---: |
|
| 94 |
-
| **mmlu_pro** | **0.
|
| 95 |
-
| biology | 0.
|
| 96 |
-
| business | 0.
|
| 97 |
-
| chemistry | 0.
|
| 98 |
-
| computer_science | 0.
|
| 99 |
-
| economics | 0.
|
| 100 |
-
| engineering | 0.
|
| 101 |
-
| health | 0.
|
| 102 |
-
| history | 0.
|
| 103 |
-
| law | 0.
|
| 104 |
-
| math | 0.6366
|
| 105 |
-
| other | 0.
|
| 106 |
-
| philosophy | 0.
|
| 107 |
-
| physics | 0.
|
| 108 |
-
| psychology | 0.
|
| 109 |
|
| 110 |
## License
|
| 111 |
|
|
|
|
| 51 |
model-f16.gguf model-q4_0-override.gguf q4_0
|
| 52 |
```
|
| 53 |
|
| 54 |
+
## Performance Measurement Commands
|
| 55 |
|
| 56 |
+
CPU uses `--device none -ngl 0`; HTP uses `--device HTP0 -ngl 99`. For each (model, backend, CTX ∈ {512, 1024, 4096}) two `llama-bench` runs were issued — one for prefill, one for decode:
|
|
|
|
| 57 |
|
| 58 |
+
```sh
|
| 59 |
+
# environment on device
|
| 60 |
+
export LD_LIBRARY_PATH=./lib
|
| 61 |
+
export ADSP_LIBRARY_PATH=./lib
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 62 |
|
| 63 |
+
# Prefill (Prefill tok/s; TTFT = CTX / Prefill × 1000)
|
| 64 |
+
./bin/llama-bench --device <none|HTP0> -m <model.gguf> \
|
| 65 |
+
--poll 1000 -t 6 --cpu-mask 0xfc --cpu-strict 1 --ubatch-size 1024 -fa on \
|
| 66 |
+
-ngl <0|99> -p <CTX> -n 0
|
| 67 |
+
|
| 68 |
+
# Decode at depth = CTX (Decode tok/s)
|
| 69 |
+
./bin/llama-bench --device <none|HTP0> -m <model.gguf> \
|
| 70 |
+
--poll 1000 -t 6 --cpu-mask 0xfc --cpu-strict 1 --ubatch-size 1024 -fa on \
|
| 71 |
+
-ngl <0|99> -p 0 -n 128 -d <CTX>
|
| 72 |
+
```
|
| 73 |
|
|
|
|
| 74 |
|
| 75 |
+
## Performance Metrics
|
| 76 |
+
Performance on IQ9 (QCS9075M)
|
|
|
|
|
|
|
|
|
|
|
|
|
| 77 |
|
| 78 |
+
| Compute | CTX | Unsloth PTQ GGUF | Google QAT GGUF | Ours |
|
| 79 |
+
|---------|----:|:-------:|:------:|:----:|
|
| 80 |
+
| CPU | 512 | 82.2 / 22.25 | 107.2 / 22.51 | 98.9 / **23.78** |
|
| 81 |
+
| CPU | 1024 | 76.6 / 21.65 | 97.8 / 21.94 | 90.8 / **23.16** |
|
| 82 |
+
| CPU | 4096 | 61.3 / 18.09 | 76.5 / 18.26 | 72.3 / **19.10** |
|
| 83 |
+
| HTP | 512 | 273.2 / 19.22 | 582.5 / 16.40 | 582.4 / **18.33** |
|
| 84 |
+
| HTP | 1024 | 266.9 / 18.93 | 552.3 / 16.45 | 552.2 / **18.17** |
|
| 85 |
+
| HTP | 4096 | 254.9 / 18.30 | 502.5 / 15.87 | 501.8 / **17.53** |
|
| 86 |
|
| 87 |
+
## Accuracy Metrics
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 88 |
|
| 89 |
+
The MMLU-Pro is measured:
|
| 90 |
|
| 91 |
+
| Subject | Unsloth PTQ GGUF | Google QAT GGUF | Ours |
|
| 92 |
| --- | ---: | ---: | ---: |
|
| 93 |
+
| **mmlu_pro** | **0.4084** | **0.4178** | **0.4131** |
|
| 94 |
+
| biology | 0.6248 ± 0.0181 | 0.6262 | 0.6248 ± 0.0181 |
|
| 95 |
+
| business | 0.4956 ± 0.0178 | 0.5019 | 0.4740 ± 0.0178 |
|
| 96 |
+
| chemistry | 0.3021 ± 0.0137 | 0.3127 | 0.3004 ± 0.0136 |
|
| 97 |
+
| computer_science | 0.4683 ± 0.0247 | 0.4854 | 0.4683 ± 0.0247 |
|
| 98 |
+
| economics | 0.4538 ± 0.0171 | 0.4976 | 0.4431 ± 0.0171 |
|
| 99 |
+
| engineering | 0.2859 ± 0.0145 | 0.3375 | 0.2931 ± 0.0146 |
|
| 100 |
+
| health | 0.3961 ± 0.0171 | 0.3423 | 0.3985 ± 0.0171 |
|
| 101 |
+
| history | 0.2730 ± 0.0229 | 0.2310 | 0.3097 ± 0.0237 |
|
| 102 |
+
| law | 0.2416 ± 0.0129 | 0.2289 | 0.2480 ± 0.0130 |
|
| 103 |
+
| math | 0.6366 ± 0.0131 | 0.6366 | 0.6373 ± 0.0131 |
|
| 104 |
+
| other | 0.3561 ± 0.0158 | 0.3528 | 0.3431 ± 0.0156 |
|
| 105 |
+
| philosophy | 0.3387 ± 0.0212 | 0.3848 | 0.3487 ± 0.0214 |
|
| 106 |
+
| physics | 0.3580 ± 0.0133 | 0.4003 | 0.3926 ± 0.0136 |
|
| 107 |
+
| psychology | 0.4875 ± 0.0177 | 0.5113 | 0.5013 ± 0.0177 |
|
| 108 |
|
| 109 |
## License
|
| 110 |
|