LFM2.5-2.6B-MLX / README.md
mlabonne's picture
Update README.md
b41f2b6 verified
|
Raw
History Blame Contribute Delete
4.05 kB
metadata
library_name: mlx
license: other
license_name: lfm1.0
license_link: LICENSE
language:
  - ar
  - zh
  - en
  - fr
  - de
  - hi
  - id
  - it
  - ja
  - ko
  - pl
  - pt
  - ru
  - es
  - th
  - vi
pipeline_tag: text-generation
tags:
  - liquid
  - lfm2.5
  - edge
  - mlx
base_model: LiquidAI/LFM2.5-2.6B
Liquid AI
Try LFMDocsLEAPDiscord

LFM2.5-2.6B-MLX

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Find more details in the original model card: https://huggingface.co/LiquidAI/LFM2.5-2.6B

Precisions

Each precision is available both as a standalone repo and as a subfolder of this repo.

Standalone repo Folder Precision Group Size Size
LiquidAI/LFM2.5-2.6B-MLX-bf16 bf16/ bf16 - 5.02 GB
LiquidAI/LFM2.5-2.6B-MLX-8bit 8bit/ 8-bit 64 2.67 GB
LiquidAI/LFM2.5-2.6B-MLX-6bit 6bit/ 6-bit 64 2.04 GB
LiquidAI/LFM2.5-2.6B-MLX-5bit 5bit/ 5-bit 64 1.76 GB
LiquidAI/LFM2.5-2.6B-MLX-4bit 4bit/ 4-bit 64 1.47 GB
LiquidAI/LFM2.5-2.6B-MLX-mxfp8 mxfp8/ MXFP8 32 2.59 GB
LiquidAI/LFM2.5-2.6B-MLX-mxfp4 mxfp4/ MXFP4 32 1.46 GB
LiquidAI/LFM2.5-2.6B-MLX-nvfp4 nvfp4/ NVFP4 16 1.53 GB

Use with mlx

pip install mlx-lm

The simplest option is to load a standalone repo directly:

from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
 
model, tokenizer = load("LiquidAI/LFM2.5-2.6B-MLX-4bit")
 
response = generate(
    model,
    tokenizer,
    prompt="The capital of France is",
    max_tokens=100,
    sampler=make_sampler(temp=0.7),
    verbose=True,
)

If you prefer this repo, note that mlx_lm.load does not resolve subfolders of a HuggingFace repo directly, so download the precision you want first:

from huggingface_hub import snapshot_download
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
 
path = snapshot_download("LiquidAI/LFM2.5-2.6B-MLX", allow_patterns=["4bit/*"])
model, tokenizer = load(f"{path}/4bit")
 
response = generate(
    model,
    tokenizer,
    prompt="The capital of France is",
    max_tokens=100,
    sampler=make_sampler(temp=0.7),
    verbose=True,
)