How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
Use Docker
docker model run hf.co/nicolasramos/odooclaw-medium-2.6b-ft:Q4_K_M
Quick Links

OdooClaw Medium 2.6B FT

The on-device agentic model for Odoo — the best quality-to-speed tradeoff in the OdooClaw family.

Fine-tuned LFM2.5-2.6B (Liquid AI) for tool calling inside Odoo (ERP) via MCP. Ask in natural language in the Odoo chat and the model picks the right Odoo tool. Current version: v18 (2026-08-19) — the definitive model of the Medium line.

Part of the OdooClaw collection. The MLX release (Apple Silicon) is odooclaw-medium-2.6b-ft-mlx.

Why this model

The Medium is the agentic sweet spot of the OdooClaw family:

  • LFM2.5-2.6B (this model, fine-tuned): the "on-device agentic" model of the LFM2.5 family — reasons before every tool call
  • Light 1.2B (odooclaw-light-1.2b-ft): faster and lighter, but lower accuracy on business/finance tasks
  • Medium 2.6B (this model): 94.2% conversation, 99.5% creation on 1000-case batteries — the most balanced model of the series

We deliberately chose the 2.6B for agentic workloads where the model reasons before every tool call — the extra accuracy is worth the small latency cost.

⭐ Evaluation results (v18, 1000-case batteries)

Battery (1000) v18
Conversation (990) 94.2%
Creation (1000) 99.5%
Business (1000) 74.0%
Invoices (1000) 71.6%

The v18 is the most balanced model of the series: top-2 in all 4 categories at once, no tradeoffs. Trained with balanced distribution (matches evaluation) and natural variety.

🚀 Hardware performance (v18, Q4_K_M)

Mac Mini M4 (MLX 4-bit)

Test Time Speed
Tool call (find partner) 0.98s —
Long text (200 tok) 3.24s 61.7 tok/s

📦 Formats

odooclaw-medium-2.6b-ft-Q4_K_M.gguf odooclaw-medium-2.6b-ft-mlx (Apple Silicon, 4-bit) Modelfile

🔧 Quick deploy

# llama.cpp
llama-server -m odooclaw-medium-2.6b-ft-Q4_K_M.gguf \
  --chat-template-file chat-template-medium-no-think.jinja \
  --temp 0.0 --jinja

# Ollama
ollama create odooclaw-medium -f Modelfile

📚 Training

LiquidAI/LFM2.5-2.6B train_lfm25_26b.py

Chat template

4-bit

License

Apache 2.0 — free for everyone, that's the whole point.

Downloads last month
187
GGUF
Model size
3B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nicolasramos/odooclaw-medium-2.6b-ft

Quantized
(98)
this model

Collection including nicolasramos/odooclaw-medium-2.6b-ft