Image-Text-to-Text
MLX
Safetensors
English
phi4mm
apple-silicon
vision-language-model
multimodal
phi-4
quantized
4bit
siglip
document-understanding
chart-understanding
ocr
conversational
Eval Results (legacy)
Instructions to use Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit") config = load_config("Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Ferox-AI/Phi-4-multimodal-instruct-mlx-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Upload Phi-4-multimodal MLX 4-bit quantized model
Browse files- .gitattributes +1 -0
- README.md +236 -0
- added_tokens.json +12 -0
- config.json +145 -0
- generation_config.json +11 -0
- merges.txt +0 -0
- model.safetensors +3 -0
- preprocessor_config.json +14 -0
- special_tokens_map.json +24 -0
- tokenizer.json +3 -0
- tokenizer_config.json +125 -0
- vocab.json +0 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,236 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
license_link: https://huggingface.co/microsoft/Phi-4-multimodal-instruct/resolve/main/LICENSE
|
| 4 |
+
language:
|
| 5 |
+
- en
|
| 6 |
+
tags:
|
| 7 |
+
- mlx
|
| 8 |
+
- apple-silicon
|
| 9 |
+
- vision-language-model
|
| 10 |
+
- multimodal
|
| 11 |
+
- phi-4
|
| 12 |
+
- quantized
|
| 13 |
+
- 4bit
|
| 14 |
+
- siglip
|
| 15 |
+
- document-understanding
|
| 16 |
+
- chart-understanding
|
| 17 |
+
- ocr
|
| 18 |
+
pipeline_tag: image-text-to-text
|
| 19 |
+
library_name: mlx
|
| 20 |
+
base_model: microsoft/Phi-4-multimodal-instruct
|
| 21 |
+
datasets:
|
| 22 |
+
- lmms-lab/DocVQA
|
| 23 |
+
- lmms-lab/ai2d
|
| 24 |
+
- MMMU/MMMU
|
| 25 |
+
- HuggingFaceM4/ChartQA
|
| 26 |
+
- lmms-lab/textvqa
|
| 27 |
+
- echo840/OCRBench
|
| 28 |
+
- derek-thomas/ScienceQA
|
| 29 |
+
- AI4Math/MathVista
|
| 30 |
+
model-index:
|
| 31 |
+
- name: Phi-4-multimodal-instruct-mlx-4bit
|
| 32 |
+
results:
|
| 33 |
+
- task:
|
| 34 |
+
type: image-text-to-text
|
| 35 |
+
dataset:
|
| 36 |
+
name: AI2D
|
| 37 |
+
type: lmms-lab/ai2d
|
| 38 |
+
split: test
|
| 39 |
+
metrics:
|
| 40 |
+
- type: accuracy
|
| 41 |
+
value: 83.0
|
| 42 |
+
name: Accuracy (n=100)
|
| 43 |
+
- task:
|
| 44 |
+
type: image-text-to-text
|
| 45 |
+
dataset:
|
| 46 |
+
name: ChartQA
|
| 47 |
+
type: HuggingFaceM4/ChartQA
|
| 48 |
+
split: test
|
| 49 |
+
metrics:
|
| 50 |
+
- type: relaxed_accuracy
|
| 51 |
+
value: 86.0
|
| 52 |
+
name: Relaxed Accuracy (n=100)
|
| 53 |
+
- task:
|
| 54 |
+
type: image-text-to-text
|
| 55 |
+
dataset:
|
| 56 |
+
name: DocVQA
|
| 57 |
+
type: lmms-lab/DocVQA
|
| 58 |
+
split: validation
|
| 59 |
+
metrics:
|
| 60 |
+
- type: anls
|
| 61 |
+
value: 82.8
|
| 62 |
+
name: ANLS (n=100)
|
| 63 |
+
- task:
|
| 64 |
+
type: image-text-to-text
|
| 65 |
+
dataset:
|
| 66 |
+
name: TextVQA
|
| 67 |
+
type: lmms-lab/textvqa
|
| 68 |
+
split: validation
|
| 69 |
+
metrics:
|
| 70 |
+
- type: accuracy
|
| 71 |
+
value: 80.0
|
| 72 |
+
name: Accuracy (n=100)
|
| 73 |
+
- task:
|
| 74 |
+
type: image-text-to-text
|
| 75 |
+
dataset:
|
| 76 |
+
name: OCRBench
|
| 77 |
+
type: echo840/OCRBench
|
| 78 |
+
split: test
|
| 79 |
+
metrics:
|
| 80 |
+
- type: score
|
| 81 |
+
value: 840
|
| 82 |
+
name: Score/1000 (n=100)
|
| 83 |
+
- task:
|
| 84 |
+
type: image-text-to-text
|
| 85 |
+
dataset:
|
| 86 |
+
name: ScienceQA
|
| 87 |
+
type: derek-thomas/ScienceQA
|
| 88 |
+
split: test
|
| 89 |
+
metrics:
|
| 90 |
+
- type: accuracy
|
| 91 |
+
value: 95.8
|
| 92 |
+
name: Accuracy (n=48, image-only)
|
| 93 |
+
- task:
|
| 94 |
+
type: image-text-to-text
|
| 95 |
+
dataset:
|
| 96 |
+
name: MathVista
|
| 97 |
+
type: AI4Math/MathVista
|
| 98 |
+
split: testmini
|
| 99 |
+
metrics:
|
| 100 |
+
- type: accuracy
|
| 101 |
+
value: 58.0
|
| 102 |
+
name: Accuracy (n=100)
|
| 103 |
+
---
|
| 104 |
+
|
| 105 |
+
# Phi-4-Multimodal-Instruct — MLX 4-bit
|
| 106 |
+
|
| 107 |
+
A 4-bit quantized [Apple MLX](https://github.com/ml-explore/mlx) conversion of [microsoft/Phi-4-multimodal-instruct](https://huggingface.co/microsoft/Phi-4-multimodal-instruct) for native inference on Apple Silicon.
|
| 108 |
+
|
| 109 |
+
**Converted by [Ferox AI](https://ferox.ca)** · Vision-language inference on MacBook / Mac Studio / Mac Pro without cloud dependencies.
|
| 110 |
+
|
| 111 |
+
| | |
|
| 112 |
+
|---|---|
|
| 113 |
+
| **Parameters** | 5.6B (pre-LoRA-fusion) |
|
| 114 |
+
| **Quantization** | 4-bit, group_size=64 (backbone only; SigLIP encoder remains FP16) |
|
| 115 |
+
| **Disk size** | ~3.9 GB |
|
| 116 |
+
| **Base model** | [microsoft/Phi-4-multimodal-instruct](https://huggingface.co/microsoft/Phi-4-multimodal-instruct) |
|
| 117 |
+
| **License** | MIT |
|
| 118 |
+
| **Modality** | Vision + Text (Phase 1; audio deferred) |
|
| 119 |
+
|
| 120 |
+
> **Other variants:** [bf16 (full precision)](https://huggingface.co/ferox-ai/Phi-4-multimodal-instruct-mlx-bf16) · [8-bit](https://huggingface.co/ferox-ai/Phi-4-multimodal-instruct-mlx-8bit)
|
| 121 |
+
|
| 122 |
+
## Quickstart
|
| 123 |
+
|
| 124 |
+
```python
|
| 125 |
+
from mlx_vlm import load, generate
|
| 126 |
+
|
| 127 |
+
model, processor = load("ferox-ai/Phi-4-multimodal-instruct-mlx-4bit")
|
| 128 |
+
|
| 129 |
+
output = generate(
|
| 130 |
+
model,
|
| 131 |
+
processor,
|
| 132 |
+
"Describe this image in detail.",
|
| 133 |
+
["path/to/image.jpg"],
|
| 134 |
+
max_tokens=512,
|
| 135 |
+
verbose=False,
|
| 136 |
+
)
|
| 137 |
+
print(output)
|
| 138 |
+
```
|
| 139 |
+
|
| 140 |
+
Requires `mlx-vlm >= 0.1.0` with Phi-4-MM architecture support. Install dependencies:
|
| 141 |
+
|
| 142 |
+
```bash
|
| 143 |
+
pip install mlx-vlm>=0.1.0 mlx>=0.22.0
|
| 144 |
+
```
|
| 145 |
+
|
| 146 |
+
## Benchmark Results
|
| 147 |
+
|
| 148 |
+
All evaluations conducted on a single Apple Silicon device using our [evaluation harness](https://github.com/ferox-ai/phi4mm-mlx). Scores are computed on a 100-sample subset of each benchmark. Microsoft's reference scores are reported on the full dataset using PyTorch FP16 — direct comparison should account for both the precision difference and sample-size variance.
|
| 149 |
+
|
| 150 |
+
| Benchmark | This Model (4-bit) | bf16 | Microsoft FP16 (full dataset) | Metric |
|
| 151 |
+
|-----------|:------------------:|:----:|:-----------------------------:|--------|
|
| 152 |
+
| **AI2D** | 83.0 | 90.0 | 82.3 | Accuracy |
|
| 153 |
+
| **ChartQA** | **86.0** | 85.0 | 81.4 | Relaxed Accuracy |
|
| 154 |
+
| **DocVQA** | 82.8 | 86.2 | 93.2 | ANLS |
|
| 155 |
+
| **MathVista** | 58.0 | 58.0 | 62.4 | Accuracy |
|
| 156 |
+
| **MMMU** | 24.0 | 31.0 | 55.1 | Accuracy |
|
| 157 |
+
| **OCRBench** | 840 | 840 | 844 | Score / 1000 |
|
| 158 |
+
| **ScienceQA** | 95.8† | 100.0† | 97.5 | Accuracy |
|
| 159 |
+
| **TextVQA** | **80.0** | 82.0 | 75.6 | Accuracy |
|
| 160 |
+
|
| 161 |
+
† ScienceQA: 48 of 100 samples scored (image-bearing questions only; 52 text-only questions excluded).
|
| 162 |
+
|
| 163 |
+
### Quantization impact
|
| 164 |
+
|
| 165 |
+
Across all benchmarks, 4-bit quantization produces a mean accuracy delta of −2.2 percentage points relative to bf16 — within the expected range for 4-bit group quantization on a model of this scale.
|
| 166 |
+
|
| 167 |
+
### Note on MMMU
|
| 168 |
+
|
| 169 |
+
The MMMU scores (24.0% 4-bit, 31.0% bf16) are significantly below Microsoft's reference (55.1%). Since the bf16 variant is lossless, this gap is not attributable to quantization or weight conversion. We attribute it to a combination of: (1) answer-extraction sensitivity in our evaluation harness for MMMU's multiple-choice format, and (2) variance inherent to a 100-sample evaluation subset. We are investigating the extraction logic and plan to re-evaluate on the full validation split. This does not reflect the model's actual MMMU capability.
|
| 170 |
+
|
| 171 |
+
## Architecture
|
| 172 |
+
|
| 173 |
+
| Component | Details |
|
| 174 |
+
|-----------|---------|
|
| 175 |
+
| **Backbone** | Phi-4-Mini (3.8B) — 32 transformer layers, hidden_size=3072, 24 query heads / 8 KV heads (GQA), head_dim=128, LongRoPE positional encoding (131K context) |
|
| 176 |
+
| **Vision encoder** | SigLIP-SO400M NaViT — 27 layers, 16 heads, head_dim=72, hidden_size=1152 |
|
| 177 |
+
| **Vision projection** | 2-layer MLP: Linear(4608→3072) → GELU → Linear(3072→3072). Input is a 2×2 spatial merge of SigLIP patch features |
|
| 178 |
+
| **Vision LoRA** | rank=256, alpha=512 (~370M parameters) — **pre-fused** into backbone weights before quantization |
|
| 179 |
+
| **Image preprocessing** | Dynamic HD tiling (deterministic grid, up to 8 crops at 448×448). PIL + NumPy only; zero PyTorch dependency at inference |
|
| 180 |
+
| **Quantization** | 4-bit with group_size=64. Applied to backbone linear layers only; SigLIP encoder weights remain in FP16 |
|
| 181 |
+
|
| 182 |
+
### Weight provenance
|
| 183 |
+
|
| 184 |
+
Weights are converted from `microsoft/Phi-4-multimodal-instruct` using a deterministic pipeline:
|
| 185 |
+
|
| 186 |
+
1. Download source checkpoint (PyTorch safetensors)
|
| 187 |
+
2. Fuse vision LoRA adapters into backbone weights (eliminates runtime adapter overhead)
|
| 188 |
+
3. Remap weight keys to MLX naming conventions
|
| 189 |
+
4. Transpose LoRA matrices (PEFT → MLX format)
|
| 190 |
+
5. Quantize backbone to 4-bit (SigLIP excluded)
|
| 191 |
+
6. Serialize as MLX safetensors
|
| 192 |
+
|
| 193 |
+
Full conversion and quantization scripts are available in the [project repository](https://github.com/ferox-ai/phi4mm-mlx).
|
| 194 |
+
|
| 195 |
+
## Intended Use
|
| 196 |
+
|
| 197 |
+
This model is designed for **local, on-device vision-language inference** on Apple Silicon hardware. Suitable applications include:
|
| 198 |
+
|
| 199 |
+
- Document understanding and extraction (invoices, forms, reports)
|
| 200 |
+
- Chart and diagram interpretation
|
| 201 |
+
- Visual question answering
|
| 202 |
+
- OCR and text recognition in images
|
| 203 |
+
- Educational content analysis
|
| 204 |
+
|
| 205 |
+
### Out of scope
|
| 206 |
+
|
| 207 |
+
- Audio processing (Phase 2, not included in this release)
|
| 208 |
+
- Production deployment without application-level safety filtering
|
| 209 |
+
- Use cases requiring guaranteed factual accuracy without human verification
|
| 210 |
+
|
| 211 |
+
## Limitations
|
| 212 |
+
|
| 213 |
+
- **100-sample evaluations.** Benchmark scores are computed on subsets, not full datasets. Expect variance relative to full-dataset evaluations.
|
| 214 |
+
- **Vision-only.** This is a Phase 1 release covering the vision modality. Audio support from the original Phi-4-multimodal architecture is not included.
|
| 215 |
+
- **No runtime LoRA switching.** Vision LoRA adapters are pre-fused; the model cannot dynamically swap adapters.
|
| 216 |
+
- **Apple Silicon required.** MLX is designed for Apple's unified memory architecture (M1/M2/M3/M4). This model will not run on CUDA or CPU-only systems.
|
| 217 |
+
|
| 218 |
+
## Citation
|
| 219 |
+
|
| 220 |
+
If you use this model in your work, please cite:
|
| 221 |
+
|
| 222 |
+
```bibtex
|
| 223 |
+
@misc{feroxai2025phi4mlx,
|
| 224 |
+
title={Phi-4-Multimodal-Instruct MLX Conversion},
|
| 225 |
+
author={Ferox AI},
|
| 226 |
+
year={2025},
|
| 227 |
+
url={https://huggingface.co/ferox-ai/Phi-4-multimodal-instruct-mlx-4bit},
|
| 228 |
+
note={4-bit quantized MLX port of microsoft/Phi-4-multimodal-instruct}
|
| 229 |
+
}
|
| 230 |
+
```
|
| 231 |
+
|
| 232 |
+
## Acknowledgments
|
| 233 |
+
|
| 234 |
+
- **Microsoft Research** for the [Phi-4-multimodal-instruct](https://huggingface.co/microsoft/Phi-4-multimodal-instruct) model and technical report
|
| 235 |
+
- **Apple MLX team** for the [MLX framework](https://github.com/ml-explore/mlx)
|
| 236 |
+
- **Prince Canuma** for [mlx-vlm](https://github.com/Blaizzy/mlx-vlm)
|
added_tokens.json
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"<|/tool_call|>": 200026,
|
| 3 |
+
"<|/tool|>": 200024,
|
| 4 |
+
"<|assistant|>": 200019,
|
| 5 |
+
"<|end|>": 200020,
|
| 6 |
+
"<|system|>": 200022,
|
| 7 |
+
"<|tag|>": 200028,
|
| 8 |
+
"<|tool_call|>": 200025,
|
| 9 |
+
"<|tool_response|>": 200027,
|
| 10 |
+
"<|tool|>": 200023,
|
| 11 |
+
"<|user|>": 200021
|
| 12 |
+
}
|
config.json
ADDED
|
@@ -0,0 +1,145 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"model_type": "phi4mm",
|
| 3 |
+
"hidden_size": 3072,
|
| 4 |
+
"intermediate_size": 8192,
|
| 5 |
+
"num_hidden_layers": 32,
|
| 6 |
+
"num_attention_heads": 24,
|
| 7 |
+
"num_key_value_heads": 8,
|
| 8 |
+
"vocab_size": 200064,
|
| 9 |
+
"rms_norm_eps": 1e-05,
|
| 10 |
+
"max_position_embeddings": 131072,
|
| 11 |
+
"original_max_position_embeddings": 4096,
|
| 12 |
+
"rope_theta": 10000.0,
|
| 13 |
+
"rope_scaling": {
|
| 14 |
+
"long_factor": [
|
| 15 |
+
1,
|
| 16 |
+
1.118320672,
|
| 17 |
+
1.250641126,
|
| 18 |
+
1.398617824,
|
| 19 |
+
1.564103225,
|
| 20 |
+
1.74916897,
|
| 21 |
+
1.956131817,
|
| 22 |
+
2.187582649,
|
| 23 |
+
2.446418898,
|
| 24 |
+
2.735880826,
|
| 25 |
+
3.059592084,
|
| 26 |
+
3.421605075,
|
| 27 |
+
3.826451687,
|
| 28 |
+
4.279200023,
|
| 29 |
+
4.785517845,
|
| 30 |
+
5.351743533,
|
| 31 |
+
5.984965424,
|
| 32 |
+
6.693110555,
|
| 33 |
+
7.485043894,
|
| 34 |
+
8.370679318,
|
| 35 |
+
9.36110372,
|
| 36 |
+
10.4687158,
|
| 37 |
+
11.70738129,
|
| 38 |
+
13.09260651,
|
| 39 |
+
14.64173252,
|
| 40 |
+
16.37415215,
|
| 41 |
+
18.31155283,
|
| 42 |
+
20.47818807,
|
| 43 |
+
22.90118105,
|
| 44 |
+
25.61086418,
|
| 45 |
+
28.64115884,
|
| 46 |
+
32.03,
|
| 47 |
+
32.1,
|
| 48 |
+
32.13,
|
| 49 |
+
32.23,
|
| 50 |
+
32.6,
|
| 51 |
+
32.61,
|
| 52 |
+
32.64,
|
| 53 |
+
32.66,
|
| 54 |
+
32.7,
|
| 55 |
+
32.71,
|
| 56 |
+
32.93,
|
| 57 |
+
32.97,
|
| 58 |
+
33.28,
|
| 59 |
+
33.49,
|
| 60 |
+
33.5,
|
| 61 |
+
44.16,
|
| 62 |
+
47.77
|
| 63 |
+
],
|
| 64 |
+
"short_factor": [
|
| 65 |
+
1.0,
|
| 66 |
+
1.0,
|
| 67 |
+
1.0,
|
| 68 |
+
1.0,
|
| 69 |
+
1.0,
|
| 70 |
+
1.0,
|
| 71 |
+
1.0,
|
| 72 |
+
1.0,
|
| 73 |
+
1.0,
|
| 74 |
+
1.0,
|
| 75 |
+
1.0,
|
| 76 |
+
1.0,
|
| 77 |
+
1.0,
|
| 78 |
+
1.0,
|
| 79 |
+
1.0,
|
| 80 |
+
1.0,
|
| 81 |
+
1.0,
|
| 82 |
+
1.0,
|
| 83 |
+
1.0,
|
| 84 |
+
1.0,
|
| 85 |
+
1.0,
|
| 86 |
+
1.0,
|
| 87 |
+
1.0,
|
| 88 |
+
1.0,
|
| 89 |
+
1.0,
|
| 90 |
+
1.0,
|
| 91 |
+
1.0,
|
| 92 |
+
1.0,
|
| 93 |
+
1.0,
|
| 94 |
+
1.0,
|
| 95 |
+
1.0,
|
| 96 |
+
1.0,
|
| 97 |
+
1.0,
|
| 98 |
+
1.0,
|
| 99 |
+
1.0,
|
| 100 |
+
1.0,
|
| 101 |
+
1.0,
|
| 102 |
+
1.0,
|
| 103 |
+
1.0,
|
| 104 |
+
1.0,
|
| 105 |
+
1.0,
|
| 106 |
+
1.0,
|
| 107 |
+
1.0,
|
| 108 |
+
1.0,
|
| 109 |
+
1.0,
|
| 110 |
+
1.0,
|
| 111 |
+
1.0,
|
| 112 |
+
1.0
|
| 113 |
+
],
|
| 114 |
+
"type": "longrope"
|
| 115 |
+
},
|
| 116 |
+
"partial_rotary_factor": 0.75,
|
| 117 |
+
"tie_word_embeddings": true,
|
| 118 |
+
"hidden_act": "silu",
|
| 119 |
+
"bos_token_id": 199999,
|
| 120 |
+
"eos_token_id": 199999,
|
| 121 |
+
"pad_token_id": 199999,
|
| 122 |
+
"image_token_id": 200010,
|
| 123 |
+
"vision_config": {
|
| 124 |
+
"hidden_size": 1152,
|
| 125 |
+
"intermediate_size": 4304,
|
| 126 |
+
"num_hidden_layers": 27,
|
| 127 |
+
"num_attention_heads": 16,
|
| 128 |
+
"head_dim": 72,
|
| 129 |
+
"image_size": 448,
|
| 130 |
+
"patch_size": 14,
|
| 131 |
+
"num_channels": 3,
|
| 132 |
+
"num_patches": 1024,
|
| 133 |
+
"hidden_act": "gelu_pytorch_tanh",
|
| 134 |
+
"layer_norm_eps": 1e-06,
|
| 135 |
+
"feature_layer": -2
|
| 136 |
+
},
|
| 137 |
+
"crop_size": 448,
|
| 138 |
+
"max_num_crops": 8,
|
| 139 |
+
"static_crop_dim": 9,
|
| 140 |
+
"vision_projection_input_size": 1152,
|
| 141 |
+
"quantization": {
|
| 142 |
+
"group_size": 64,
|
| 143 |
+
"bits": 4
|
| 144 |
+
}
|
| 145 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,11 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"_from_model_config": true,
|
| 3 |
+
"bos_token_id": 199999,
|
| 4 |
+
"eos_token_id": [
|
| 5 |
+
200020,
|
| 6 |
+
199999
|
| 7 |
+
],
|
| 8 |
+
"pad_token_id": 199999,
|
| 9 |
+
"transformers_version": "4.46.1",
|
| 10 |
+
"use_cache": true
|
| 11 |
+
}
|
merges.txt
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:badee6235dddd0d08d06cc509859304090adb6c655eb21cfa2bc49218716f05d
|
| 3 |
+
size 3894257630
|
preprocessor_config.json
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"auto_map": {
|
| 3 |
+
"AutoProcessor": "processing_phi4mm.Phi4MMProcessor",
|
| 4 |
+
"AutoImageProcessor": "processing_phi4mm.Phi4MMImageProcessor",
|
| 5 |
+
"AutoFeatureExtractor": "processing_phi4mm.Phi4MMAudioFeatureExtractor"
|
| 6 |
+
},
|
| 7 |
+
"image_processor_type": "Phi4MMImageProcessor",
|
| 8 |
+
"processor_class": "Phi4MMProcessor",
|
| 9 |
+
"feature_extractor_type": "Phi4MMAudioFeatureExtractor",
|
| 10 |
+
"audio_compression_rate": 8,
|
| 11 |
+
"audio_downsample_rate": 1,
|
| 12 |
+
"audio_feat_stride": 1,
|
| 13 |
+
"dynamic_hd": 36
|
| 14 |
+
}
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,24 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token": {
|
| 3 |
+
"content": "<|endoftext|>",
|
| 4 |
+
"lstrip": false,
|
| 5 |
+
"normalized": false,
|
| 6 |
+
"rstrip": false,
|
| 7 |
+
"single_word": false
|
| 8 |
+
},
|
| 9 |
+
"eos_token": {
|
| 10 |
+
"content": "<|endoftext|>",
|
| 11 |
+
"lstrip": false,
|
| 12 |
+
"normalized": false,
|
| 13 |
+
"rstrip": false,
|
| 14 |
+
"single_word": false
|
| 15 |
+
},
|
| 16 |
+
"pad_token": "<|endoftext|>",
|
| 17 |
+
"unk_token": {
|
| 18 |
+
"content": "<|endoftext|>",
|
| 19 |
+
"lstrip": false,
|
| 20 |
+
"normalized": false,
|
| 21 |
+
"rstrip": false,
|
| 22 |
+
"single_word": false
|
| 23 |
+
}
|
| 24 |
+
}
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:4c1b9f641d4f8b7247b8d5007dd3b6a9f6a87cb5123134fe0d326f14d10c0585
|
| 3 |
+
size 15524479
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,125 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"added_tokens_decoder": {
|
| 4 |
+
"200010": {
|
| 5 |
+
"content": "<|endoftext10|>",
|
| 6 |
+
"lstrip": false,
|
| 7 |
+
"normalized": false,
|
| 8 |
+
"rstrip": false,
|
| 9 |
+
"single_word": false,
|
| 10 |
+
"special": true
|
| 11 |
+
},
|
| 12 |
+
"200011": {
|
| 13 |
+
"content": "<|endoftext11|>",
|
| 14 |
+
"lstrip": false,
|
| 15 |
+
"normalized": false,
|
| 16 |
+
"rstrip": false,
|
| 17 |
+
"single_word": false,
|
| 18 |
+
"special": true
|
| 19 |
+
},
|
| 20 |
+
"199999": {
|
| 21 |
+
"content": "<|endoftext|>",
|
| 22 |
+
"lstrip": false,
|
| 23 |
+
"normalized": false,
|
| 24 |
+
"rstrip": false,
|
| 25 |
+
"single_word": false,
|
| 26 |
+
"special": true
|
| 27 |
+
},
|
| 28 |
+
"200018": {
|
| 29 |
+
"content": "<|endofprompt|>",
|
| 30 |
+
"lstrip": false,
|
| 31 |
+
"normalized": false,
|
| 32 |
+
"rstrip": false,
|
| 33 |
+
"single_word": false,
|
| 34 |
+
"special": true
|
| 35 |
+
},
|
| 36 |
+
"200019": {
|
| 37 |
+
"content": "<|assistant|>",
|
| 38 |
+
"lstrip": false,
|
| 39 |
+
"normalized": false,
|
| 40 |
+
"rstrip": true,
|
| 41 |
+
"single_word": false,
|
| 42 |
+
"special": true
|
| 43 |
+
},
|
| 44 |
+
"200020": {
|
| 45 |
+
"content": "<|end|>",
|
| 46 |
+
"lstrip": false,
|
| 47 |
+
"normalized": false,
|
| 48 |
+
"rstrip": true,
|
| 49 |
+
"single_word": false,
|
| 50 |
+
"special": true
|
| 51 |
+
},
|
| 52 |
+
"200021": {
|
| 53 |
+
"content": "<|user|>",
|
| 54 |
+
"lstrip": false,
|
| 55 |
+
"normalized": false,
|
| 56 |
+
"rstrip": true,
|
| 57 |
+
"single_word": false,
|
| 58 |
+
"special": true
|
| 59 |
+
},
|
| 60 |
+
"200022": {
|
| 61 |
+
"content": "<|system|>",
|
| 62 |
+
"lstrip": false,
|
| 63 |
+
"normalized": false,
|
| 64 |
+
"rstrip": true,
|
| 65 |
+
"single_word": false,
|
| 66 |
+
"special": true
|
| 67 |
+
},
|
| 68 |
+
"200023": {
|
| 69 |
+
"content": "<|tool|>",
|
| 70 |
+
"lstrip": false,
|
| 71 |
+
"normalized": false,
|
| 72 |
+
"rstrip": true,
|
| 73 |
+
"single_word": false,
|
| 74 |
+
"special": false
|
| 75 |
+
},
|
| 76 |
+
"200024": {
|
| 77 |
+
"content": "<|/tool|>",
|
| 78 |
+
"lstrip": false,
|
| 79 |
+
"normalized": false,
|
| 80 |
+
"rstrip": true,
|
| 81 |
+
"single_word": false,
|
| 82 |
+
"special": false
|
| 83 |
+
},
|
| 84 |
+
"200025": {
|
| 85 |
+
"content": "<|tool_call|>",
|
| 86 |
+
"lstrip": false,
|
| 87 |
+
"normalized": false,
|
| 88 |
+
"rstrip": true,
|
| 89 |
+
"single_word": false,
|
| 90 |
+
"special": false
|
| 91 |
+
},
|
| 92 |
+
"200026": {
|
| 93 |
+
"content": "<|/tool_call|>",
|
| 94 |
+
"lstrip": false,
|
| 95 |
+
"normalized": false,
|
| 96 |
+
"rstrip": true,
|
| 97 |
+
"single_word": false,
|
| 98 |
+
"special": false
|
| 99 |
+
},
|
| 100 |
+
"200027": {
|
| 101 |
+
"content": "<|tool_response|>",
|
| 102 |
+
"lstrip": false,
|
| 103 |
+
"normalized": false,
|
| 104 |
+
"rstrip": true,
|
| 105 |
+
"single_word": false,
|
| 106 |
+
"special": false
|
| 107 |
+
},
|
| 108 |
+
"200028": {
|
| 109 |
+
"content": "<|tag|>",
|
| 110 |
+
"lstrip": false,
|
| 111 |
+
"normalized": false,
|
| 112 |
+
"rstrip": true,
|
| 113 |
+
"single_word": false,
|
| 114 |
+
"special": true
|
| 115 |
+
}
|
| 116 |
+
},
|
| 117 |
+
"bos_token": "<|endoftext|>",
|
| 118 |
+
"chat_template": "{% for message in messages %}{% if message['role'] == 'system' and 'tools' in message and message['tools'] is not none %}{{ '<|' + message['role'] + '|>' + message['content'] + '<|tool|>' + message['tools'] + '<|/tool|>' + '<|end|>' }}{% else %}{{ '<|' + message['role'] + '|>' + message['content'] + '<|end|>' }}{% endif %}{% endfor %}{% if add_generation_prompt %}{{ '<|assistant|>' }}{% else %}{{ eos_token }}{% endif %}",
|
| 119 |
+
"clean_up_tokenization_spaces": false,
|
| 120 |
+
"eos_token": "<|endoftext|>",
|
| 121 |
+
"model_max_length": 131072,
|
| 122 |
+
"pad_token": "<|endoftext|>",
|
| 123 |
+
"tokenizer_class": "GPT2TokenizerFast",
|
| 124 |
+
"unk_token": "<|endoftext|>"
|
| 125 |
+
}
|
vocab.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|