Instructions to use Pluto-AI-Labs/Athena-1.0-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Pluto-AI-Labs/Athena-1.0-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Pluto-AI-Labs/Athena-1.0-4B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Pluto-AI-Labs/Athena-1.0-4B") model = AutoModelForMultimodalLM.from_pretrained("Pluto-AI-Labs/Athena-1.0-4B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Pluto-AI-Labs/Athena-1.0-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Pluto-AI-Labs/Athena-1.0-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pluto-AI-Labs/Athena-1.0-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Pluto-AI-Labs/Athena-1.0-4B
- SGLang
How to use Pluto-AI-Labs/Athena-1.0-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Pluto-AI-Labs/Athena-1.0-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pluto-AI-Labs/Athena-1.0-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Pluto-AI-Labs/Athena-1.0-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pluto-AI-Labs/Athena-1.0-4B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Pluto-AI-Labs/Athena-1.0-4B with Docker Model Runner:
docker model run hf.co/Pluto-AI-Labs/Athena-1.0-4B
Athena 1.0 — 4B
Athena 1.0 is a vision-language model developed by Pluto AI Labs through supervised adaptation of Qwen3.5-4B. The training corpus combines document understanding, charts, scene text, visual reasoning, conversations and grounding. This repository contains standalone BF16 Transformers weights: the trained LoRA updates are merged into the original language model, vision encoder and vision merger. Inference does not require PEFT or a separate adapter download.
Model at a glance
| Property | Value |
|---|---|
| Architecture | Qwen3.5 dense hybrid language model with vision encoder |
| Standalone parameters | 4,539,265,536 |
| Starting model | Qwen/Qwen3.5-4B, post-trained release |
| Adaptation | BF16 frozen base; 67,715,072 FP32 LoRA parameters |
| LoRA ranks | Language/merger 32; vision 16; alpha 2×rank; dropout 0 |
| Completed training | 250,000/250,000 rows; one pass; 15,625 optimizer steps |
| Supervised tokens | 28,933,709 |
| Training sequence limit | 4,096 tokens per window |
| Precision of this release | BF16 safetensors |
| Hardware | One NVIDIA H100-SXM5-80GB, across resumable Colab sessions |
The original architecture's native context configuration is retained. Context beyond the training window, inherited video capability and multilingual capability have not been established by the training loss alone. The full training recipe, environment and state are archived with the release.
Data and adaptation
| Block | Rows |
|---|---|
| Documents | 50,000 |
| OCR and scene text | 40,000 |
| Visual reasoning | 43,250 |
| Reasoning-family breadth | 53,000 |
| Charts and infographics | 30,000 |
| Visual conversations | 30,000 |
| Grounding | 750 |
| Identity examples | 3,000 |
| Total | 250,000 |
All 250,000 selected rows were processed. Long examples were handled with supervised windows; 10 prompts were shortened. Identity examples had a 1/6 loss weight. The training image policy allowed up to 512 visual tokens per image. This is supervised fine-tuning, not training a foundation model from scratch or reinforcement learning. There was no dedicated video training block.
Training completed in approximately 18.07 recorded wall hours with an estimated 122.34 Colab compute units. These are runtime estimates, not an account billing audit. The final 100-step mean training loss was 0.37682. On the fixed, small monitoring set, teacher-forced macro NLL changed from 2.17181 to 0.40028; the best recorded NLL was 0.37395 at step 10,556. An 81.57% reduction in this monitoring NLL is not an 81.57% increase in accuracy. The selected release is the final step-15,625 checkpoint, chosen before the benchmark run. The monitoring set contained 12 examples per ChartQA, DocVQA and TextVQA task; it is not a substitute for these eight benchmarks.
Quickstart
Use the tested Transformers 5.18.0 stack or a compatible later release. Qwen3.5's linear-attention kernels require compatible flash-linear-attention and causal-conv1d binaries; see the release notebook's environment checks.
import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
repo = "Pluto-AI-Labs/Athena-1.0-4B"
processor = AutoProcessor.from_pretrained(repo)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
repo, dtype=torch.bfloat16, device_map="auto", attn_implementation="sdpa"
).eval()
messages = [
{"role": "system", "content": "You are Athena 1.0, a vision-language model by Pluto AI Labs."},
{"role": "user", "content": [
{"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
{"type": "text", "text": "Describe the image."}
]}
]
inputs = processor.apply_chat_template(
messages, tokenize=True, add_generation_prompt=True, enable_thinking=False,
return_dict=True, return_tensors="pt"
).to(model.device)
with torch.inference_mode():
outputs = model.generate(**inputs, do_sample=False, max_new_tokens=256)
print(processor.decode(outputs[0, inputs.input_ids.shape[-1]:], skip_special_tokens=True))
The system prompt is recommended for deployment identity. Images should be supplied through the processor's multimodal chat template. The standalone model can be downloaded and loaded from this repository; no training-repository adapter path is required.
Intended use and limitations
Athena is intended for research and application development involving images, documents and conversation. It can hallucinate, misread small text, produce incorrect reasoning and inherit biases from its base model and training sources. Training loss, identity compliance and benchmark accuracy measure different properties. Evaluate the model on representative deployment data before relying on it. This release does not establish reliability for safety-critical decisions.
The base model is distributed under Apache 2.0. Training data originates from multiple upstream collections with their own licenses and conditions; consult the Athena dataset card and original data sources. This card does not certify that every source permits every downstream use. Benchmark media and gated test questions are not redistributed in this repository; only output journals, identifiers and reports are retained.
Reproducibility and provenance
- Starting weights:
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. - Training data revision:
67bc05bc4e35f5a1f29569b6aeea8114583424bf. - Final training snapshot:
eae70590cc5f5657ec64a8b35cfc4e9b21c05152. - Recovery tag:
training-final-15625; retains model adapters, optimizer, RNG and the processed-example manifest. - Standalone release provenance and hashes:
release/manifest.jsonandrelease/weights_sha256.json. - Evaluation run:
evaluations/athena-final-budget-v1/; raw generations and scored results are separate artifacts.
The main branch is cleaned for inference. Archived training history remains recoverable and still occupies Hub storage. Repository visibility is not changed by the release notebook.
Acknowledgements
We thank the Qwen team for the starting model and the contributors to the Athena data sources and evaluation benchmarks. Technical references: IFBench, GPQA, MathArena HMMT, MMMLU, MMMU-Pro, ERQA, OmniDocBench and Video-MME.
## Matched subset evaluation · athena-runall-chatstop-v3
Status: **PROVISIONAL: incomplete coverage, output cutoffs, or scorer failures**. Athena and its starting Qwen3.5-4B were evaluated on identical selected inputs using direct-answer, greedy decoding and verified native chat end tokens. Compatible saved v2 first-turn prefixes were rescored on CPU for both models, with source hashes retained; unchanged genuine length-limit failures were retained. Timeouts without a completed turn were regenerated. Completion and extraction diagnostics are explicit. Published Qwen leaderboard scores use different settings and are listed only in a separate reference file.
| Benchmark | Athena | Qwen4B_matched | Delta_pp | Paired | Selected | Athena_token_limits | Qwen_token_limits | Athena_format_failures | Qwen_format_failures | Status |
|:-------------------------|---------:|-----------------:|-----------:|---------:|-----------:|----------------------:|--------------------:|-------------------------:|-----------------------:|:-------------------------| | IFBench | 23.81 | 28.57 | -4.76 | 21 | 21 | 1 | 2 | 1 | 2 | PROVISIONAL / INCOMPLETE | | GPQA Diamond | 38.1 | 47.62 | -9.52 | 21 | 21 | 0 | 0 | 0 | 0 | Completed subset | | HMMT Feb 2025 | 0 | 0 | 0 | 5 | 5 | 3 | 5 | 3 | 5 | PROVISIONAL / INCOMPLETE | | MMMLU | 47.62 | 66.67 | -19.05 | 42 | 42 | 0 | 0 | 0 | 0 | Completed subset | | MMMU-Pro vision | 37.21 | 41.86 | -4.65 | 43 | 43 | 0 | 0 | 1 | 0 | Completed subset | | ERQA | 43.1 | 41.38 | 1.72 | 58 | 58 | 0 | 0 | 0 | 0 | Completed subset | | OmniDocBench v1.5 | 51.11 | 92.85 | -41.74 | 10 | 10 | 0 | 0 | 1 | 0 | PROVISIONAL / INCOMPLETE | | Video-MME with subtitles | 44.44 | 22.22 | 22.22 | 9 | 9 | 0 | 0 | 0 | 4 | Completed subset |

[Protocol](evaluations/athena-runall-chatstop-v3/protocol.json) · [Raw responses](evaluations/athena-runall-chatstop-v3/responses.jsonl.gz) · [Results ZIP](evaluations/athena-runall-chatstop-v3/Athena_Benchmark_Results_v3_stopfix.zip) · [Machine-readable results](evaluations/athena-runall-chatstop-v3/report.json)
The report includes paired bootstrap intervals where applicable, all fourteen MMMLU languages, video-cluster uncertainty, identity screening, exact training-image overlap exclusions, and the original v1.5 OmniDoc metrics when all required components score successfully. Scores do not establish general superiority over 4B, 7B or 9B models.
<!-- ATHENA_RUNALL_EVAL_END -->
- Downloads last month
- 5