K2-Horizon-0.9B (MLX, 8-bit)

8-bit MLX quantization of IFM/K2-Horizon-0.9B, converted from revision fa7f5de. 8.5 bits/weight effective, 1.2 GB on disk. For Apple silicon.

K2-Horizon-0.9B is IFM's compact dense K2-Horizon model: a 0.9B-class decoder-only model distilled from domain teachers, with a 128K (131,072-token) context window through YaRN RoPE scaling.

Requirements

mlx-lm doesn't support the k2_horizon architecture yet. There's an open request: mlx-lm#1876. Until support lands, this repo ships the MLX model code (k2_horizon.py), which mlx-lm loads through the model_file entry in config.json. So the released mlx-lm works as-is:

pip install -U mlx-lm

Pass --trust-remote-code (or trust_remote_code=True). mlx-lm 0.31.3 loads the file without it, but newer versions require it. The file runs on your machine, so read it first if you like; its comments describe how it differs from IFM's PyTorch code.

Once K2-Horizon support lands in a released mlx-lm, this repo will be updated to use it. If something breaks after that, open a discussion here asking for a re-upload.

How it was quantized

mlx_lm.convert -q --q-bits 8 --q-group-size 64 (integer affine quantization, bf16 scales and biases). All linear layers and the embeddings are 8-bit. WikiText-2 perplexity is within +0.3% of bf16, so it is effectively lossless.

Memory

Peak 1.3 GB for a short prompt; fits a 8 GB Mac. The KV cache adds about 56 KB per token in bf16 (7.0 GB at 128K tokens), so long contexts need more memory; --max-kv-size and KV-cache quantization (--kv-bits 8, where available) reduce it.

Conversion check

The MLX implementation was checked against IFM's PyTorch code (modeling_k2_horizon.py) in fp32 on the real weights of this model:

  • Layer by layer, all 28 layers match to a relative error of 4e-6 or better, and the next-token predictions agree at every position.
  • Token by token: the full model in fp32 greedily generated with the KV cache, and the PyTorch model picked the same token at every step (320/320 tokens across English, code, math and Chinese prompts). Cached and uncached outputs also agree at every step.

Smoke-tested after conversion with released mlx-lm 0.31.3: 17 * 23 → 391 and "capital of Australia" → Canberra, both ending normally. On a Mac Studio M4 Max 128GB: 306.9 tok/s generation, peak 1.3 GB (short prompt).

Benchmarks (all K2-Horizon-0.9B MLX variants)

WikiText-2 test perplexity (128 × 512 tokens, lower is better) and generation speed on an M4 Max 128GB, all measured the same way:

bf16 8-bit 4-bit
Bits/weight 16 8.5 6.4
Disk 2.2 GB 1.2 GB 0.9 GB
Peak memory 2.2 GB 1.3 GB 1.0 GB
WikiText-2 perplexity 18.654 18.713 (+0.3%) 19.703 (+5.6%)
Generation 195.7 tok/s 306.9 tok/s 379.8 tok/s

Perplexity is a coarse signal. Test the versions on your own workload before picking one. Other K2-Horizon sizes: the K2-Horizon collection.

Usage

mlx_lm.generate --model mlx-community/K2-Horizon-0.9B-8bit --trust-remote-code --prompt "Explain mixture-of-experts in two sentences." --max-tokens 2048
from mlx_lm import load, generate

# Newer mlx-lm versions need trust_remote_code=True; on mlx-lm 0.31.3 use load("mlx-community/K2-Horizon-0.9B-8bit").
model, tokenizer = load("mlx-community/K2-Horizon-0.9B-8bit", trust_remote_code=True)
messages = [{"role": "user", "content": "Explain mixture-of-experts in two sentences."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt, max_tokens=2048))

mlx_lm.chat and mlx_lm.server take the same --trust-remote-code flag.

The model thinks before it answers, inside <ifm|think> … </ifm|think>. Leave room for that in max_tokens. The effort level is set with reasoning_effort in the chat template: "high" (default), "medium" or "low", e.g. apply_chat_template(messages, add_generation_prompt=True, reasoning_effort="low").

Notes:

  • Server output: mlx-lm doesn't recognize the <ifm|think> tags yet, so mlx_lm.server returns the thinking text inside content, before </ifm|think>, rather than in a separate reasoning field. K2-Horizon's tool-call format isn't parsed yet either.
  • End-of-turn token: the original generation_config.json lists only <|endoftext|> (id 1) as EOS, but the chat template ends each turn with <|ifm|im_end|> (id 64019). This repo adds 64019 to eos_token_id, like the other K2-Horizon models, so generation stops at the end of the answer and <|ifm|im_end|> does not appear in the output.
  • dtype: the original config says float32, but the checkpoint weights are bf16. The config here says bfloat16, so no tool upcasts the weights.
  • YaRN: RoPE uses YaRN (factor 16, original 8K context). The MLX code checks that the YaRN settings match the reference and raises an error for settings it does not support.
  • Recommended sampling (from IFM's model card): temperature=0.6, top_p=0.95, and a large output budget for thinking.

License

Apache-2.0, inherited from the base model. Refer to the original model card for architecture, benchmarks and intended use. All credit for the model belongs to IFM.

Downloads last month
37
Safetensors
Model size
1B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/K2-Horizon-0.9B-8bit

Quantized
(11)
this model

Collection including mlx-community/K2-Horizon-0.9B-8bit