Instructions to use mlx-community/K2-Horizon-0.9B-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/K2-Horizon-0.9B-8bit with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("mlx-community/K2-Horizon-0.9B-8bit") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use mlx-community/K2-Horizon-0.9B-8bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/K2-Horizon-0.9B-8bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "mlx-community/K2-Horizon-0.9B-8bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use mlx-community/K2-Horizon-0.9B-8bit with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "mlx-community/K2-Horizon-0.9B-8bit"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "mlx-community/K2-Horizon-0.9B-8bit" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mlx-community/K2-Horizon-0.9B-8bit", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use mlx-community/K2-Horizon-0.9B-8bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/K2-Horizon-0.9B-8bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default mlx-community/K2-Horizon-0.9B-8bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use mlx-community/K2-Horizon-0.9B-8bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "mlx-community/K2-Horizon-0.9B-8bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "mlx-community/K2-Horizon-0.9B-8bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
K2-Horizon-0.9B (MLX, 8-bit)
8-bit MLX quantization of IFM/K2-Horizon-0.9B, converted from revision fa7f5de. 8.5 bits/weight effective, 1.2 GB on disk. For Apple silicon.
K2-Horizon-0.9B is IFM's compact dense K2-Horizon model: a 0.9B-class decoder-only model distilled from domain teachers, with a 128K (131,072-token) context window through YaRN RoPE scaling.
Requirements
mlx-lm doesn't support the k2_horizon architecture yet. There's an open request: mlx-lm#1876. Until support lands, this repo ships the MLX model code (k2_horizon.py), which mlx-lm loads through the model_file entry in config.json. So the released mlx-lm works as-is:
pip install -U mlx-lm
Pass --trust-remote-code (or trust_remote_code=True). mlx-lm 0.31.3 loads the file without it, but newer versions require it. The file runs on your machine, so read it first if you like; its comments describe how it differs from IFM's PyTorch code.
Once K2-Horizon support lands in a released mlx-lm, this repo will be updated to use it. If something breaks after that, open a discussion here asking for a re-upload.
How it was quantized
mlx_lm.convert -q --q-bits 8 --q-group-size 64 (integer affine quantization, bf16 scales and biases). All linear layers and the embeddings are 8-bit. WikiText-2 perplexity is within +0.3% of bf16, so it is effectively lossless.
Memory
Peak 1.3 GB for a short prompt; fits a 8 GB Mac. The KV cache adds about 56 KB per token in bf16 (7.0 GB at 128K tokens), so long contexts need more memory; --max-kv-size and KV-cache quantization (--kv-bits 8, where available) reduce it.
Conversion check
The MLX implementation was checked against IFM's PyTorch code (modeling_k2_horizon.py) in fp32 on the real weights of this model:
- Layer by layer, all 28 layers match to a relative error of 4e-6 or better, and the next-token predictions agree at every position.
- Token by token: the full model in fp32 greedily generated with the KV cache, and the PyTorch model picked the same token at every step (320/320 tokens across English, code, math and Chinese prompts). Cached and uncached outputs also agree at every step.
Smoke-tested after conversion with released mlx-lm 0.31.3: 17 * 23 → 391 and "capital of Australia" → Canberra, both ending normally. On a Mac Studio M4 Max 128GB: 306.9 tok/s generation, peak 1.3 GB (short prompt).
Benchmarks (all K2-Horizon-0.9B MLX variants)
WikiText-2 test perplexity (128 × 512 tokens, lower is better) and generation speed on an M4 Max 128GB, all measured the same way:
| bf16 | 8-bit | 4-bit | |
|---|---|---|---|
| Bits/weight | 16 | 8.5 | 6.4 |
| Disk | 2.2 GB | 1.2 GB | 0.9 GB |
| Peak memory | 2.2 GB | 1.3 GB | 1.0 GB |
| WikiText-2 perplexity | 18.654 | 18.713 (+0.3%) | 19.703 (+5.6%) |
| Generation | 195.7 tok/s | 306.9 tok/s | 379.8 tok/s |
Perplexity is a coarse signal. Test the versions on your own workload before picking one. Other K2-Horizon sizes: the K2-Horizon collection.
Usage
mlx_lm.generate --model mlx-community/K2-Horizon-0.9B-8bit --trust-remote-code --prompt "Explain mixture-of-experts in two sentences." --max-tokens 2048
from mlx_lm import load, generate
# Newer mlx-lm versions need trust_remote_code=True; on mlx-lm 0.31.3 use load("mlx-community/K2-Horizon-0.9B-8bit").
model, tokenizer = load("mlx-community/K2-Horizon-0.9B-8bit", trust_remote_code=True)
messages = [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt, max_tokens=2048))
mlx_lm.chat and mlx_lm.server take the same --trust-remote-code flag.
The model thinks before it answers, inside <ifm|think> … </ifm|think>. Leave room for that in max_tokens. The effort level is set with reasoning_effort in the chat template: "high" (default), "medium" or "low", e.g. apply_chat_template(messages, add_generation_prompt=True, reasoning_effort="low").
Notes:
- Server output: mlx-lm doesn't recognize the
<ifm|think>tags yet, somlx_lm.serverreturns the thinking text insidecontent, before</ifm|think>, rather than in a separatereasoningfield. K2-Horizon's tool-call format isn't parsed yet either. - End-of-turn token: the original
generation_config.jsonlists only<|endoftext|>(id 1) as EOS, but the chat template ends each turn with<|ifm|im_end|>(id 64019). This repo adds 64019 toeos_token_id, like the other K2-Horizon models, so generation stops at the end of the answer and<|ifm|im_end|>does not appear in the output. - dtype: the original config says
float32, but the checkpoint weights are bf16. The config here saysbfloat16, so no tool upcasts the weights. - YaRN: RoPE uses YaRN (factor 16, original 8K context). The MLX code checks that the YaRN settings match the reference and raises an error for settings it does not support.
- Recommended sampling (from IFM's model card):
temperature=0.6,top_p=0.95, and a large output budget for thinking.
License
Apache-2.0, inherited from the base model. Refer to the original model card for architecture, benchmarks and intended use. All credit for the model belongs to IFM.
- Downloads last month
- 37
8-bit
Model tree for mlx-community/K2-Horizon-0.9B-8bit
Base model
IFM/K2-Horizon-0.9B