Athena 1.0 — 4B

Athena 1.0 is a vision-language model developed by Pluto AI Labs through supervised adaptation of Qwen3.5-4B. The training corpus combines document understanding, charts, scene text, visual reasoning, conversations and grounding. This repository contains standalone BF16 Transformers weights: the trained LoRA updates are merged into the original language model, vision encoder and vision merger. Inference does not require PEFT or a separate adapter download.

Model at a glance

Property Value
Architecture Qwen3.5 dense hybrid language model with vision encoder
Standalone parameters 4,539,265,536
Starting model Qwen/Qwen3.5-4B, post-trained release
Adaptation BF16 frozen base; 67,715,072 FP32 LoRA parameters
LoRA ranks Language/merger 32; vision 16; alpha 2×rank; dropout 0
Completed training 250,000/250,000 rows; one pass; 15,625 optimizer steps
Supervised tokens 28,933,709
Training sequence limit 4,096 tokens per window
Precision of this release BF16 safetensors
Hardware One NVIDIA H100-SXM5-80GB, across resumable Colab sessions

The original architecture's native context configuration is retained. Context beyond the training window, inherited video capability and multilingual capability have not been established by the training loss alone. The full training recipe, environment and state are archived with the release.

Data and adaptation

Block Rows
Documents 50,000
OCR and scene text 40,000
Visual reasoning 43,250
Reasoning-family breadth 53,000
Charts and infographics 30,000
Visual conversations 30,000
Grounding 750
Identity examples 3,000
Total 250,000

All 250,000 selected rows were processed. Long examples were handled with supervised windows; 10 prompts were shortened. Identity examples had a 1/6 loss weight. The training image policy allowed up to 512 visual tokens per image. This is supervised fine-tuning, not training a foundation model from scratch or reinforcement learning. There was no dedicated video training block.

Training completed in approximately 18.07 recorded wall hours with an estimated 122.34 Colab compute units. These are runtime estimates, not an account billing audit. The final 100-step mean training loss was 0.37682. On the fixed, small monitoring set, teacher-forced macro NLL changed from 2.17181 to 0.40028; the best recorded NLL was 0.37395 at step 10,556. An 81.57% reduction in this monitoring NLL is not an 81.57% increase in accuracy. The selected release is the final step-15,625 checkpoint, chosen before the benchmark run. The monitoring set contained 12 examples per ChartQA, DocVQA and TextVQA task; it is not a substitute for these eight benchmarks.

Quickstart

Use the tested Transformers 5.18.0 stack or a compatible later release. Qwen3.5's linear-attention kernels require compatible flash-linear-attention and causal-conv1d binaries; see the release notebook's environment checks.

import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration

repo = "Pluto-AI-Labs/Athena-1.0-4B"
processor = AutoProcessor.from_pretrained(repo)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    repo, dtype=torch.bfloat16, device_map="auto", attn_implementation="sdpa"
).eval()
messages = [
    {"role": "system", "content": "You are Athena 1.0, a vision-language model by Pluto AI Labs."},
    {"role": "user", "content": [
        {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
        {"type": "text", "text": "Describe the image."}
    ]}
]
inputs = processor.apply_chat_template(
    messages, tokenize=True, add_generation_prompt=True, enable_thinking=False,
    return_dict=True, return_tensors="pt"
).to(model.device)
with torch.inference_mode():
    outputs = model.generate(**inputs, do_sample=False, max_new_tokens=256)
print(processor.decode(outputs[0, inputs.input_ids.shape[-1]:], skip_special_tokens=True))

The system prompt is recommended for deployment identity. Images should be supplied through the processor's multimodal chat template. The standalone model can be downloaded and loaded from this repository; no training-repository adapter path is required.

Intended use and limitations

Athena is intended for research and application development involving images, documents and conversation. It can hallucinate, misread small text, produce incorrect reasoning and inherit biases from its base model and training sources. Training loss, identity compliance and benchmark accuracy measure different properties. Evaluate the model on representative deployment data before relying on it. This release does not establish reliability for safety-critical decisions.

The base model is distributed under Apache 2.0. Training data originates from multiple upstream collections with their own licenses and conditions; consult the Athena dataset card and original data sources. This card does not certify that every source permits every downstream use. Benchmark media and gated test questions are not redistributed in this repository; only output journals, identifiers and reports are retained.

Reproducibility and provenance

The main branch is cleaned for inference. Archived training history remains recoverable and still occupies Hub storage. Repository visibility is not changed by the release notebook.

Acknowledgements

We thank the Qwen team for the starting model and the contributors to the Athena data sources and evaluation benchmarks. Technical references: IFBench, GPQA, MathArena HMMT, MMMLU, MMMU-Pro, ERQA, OmniDocBench and Video-MME.

## Matched subset evaluation · athena-runall-chatstop-v3

Status: **PROVISIONAL: incomplete coverage, output cutoffs, or scorer failures**. Athena and its starting Qwen3.5-4B were evaluated on identical selected inputs using direct-answer, greedy decoding and verified native chat end tokens. Compatible saved v2 first-turn prefixes were rescored on CPU for both models, with source hashes retained; unchanged genuine length-limit failures were retained. Timeouts without a completed turn were regenerated. Completion and extraction diagnostics are explicit. Published Qwen leaderboard scores use different settings and are listed only in a separate reference file.

| Benchmark                |   Athena |   Qwen4B_matched |   Delta_pp |   Paired |   Selected |   Athena_token_limits |   Qwen_token_limits |   Athena_format_failures |   Qwen_format_failures | Status                   |

|:-------------------------|---------:|-----------------:|-----------:|---------:|-----------:|----------------------:|--------------------:|-------------------------:|-----------------------:|:-------------------------| | IFBench | 23.81 | 28.57 | -4.76 | 21 | 21 | 1 | 2 | 1 | 2 | PROVISIONAL / INCOMPLETE | | GPQA Diamond | 38.1 | 47.62 | -9.52 | 21 | 21 | 0 | 0 | 0 | 0 | Completed subset | | HMMT Feb 2025 | 0 | 0 | 0 | 5 | 5 | 3 | 5 | 3 | 5 | PROVISIONAL / INCOMPLETE | | MMMLU | 47.62 | 66.67 | -19.05 | 42 | 42 | 0 | 0 | 0 | 0 | Completed subset | | MMMU-Pro vision | 37.21 | 41.86 | -4.65 | 43 | 43 | 0 | 0 | 1 | 0 | Completed subset | | ERQA | 43.1 | 41.38 | 1.72 | 58 | 58 | 0 | 0 | 0 | 0 | Completed subset | | OmniDocBench v1.5 | 51.11 | 92.85 | -41.74 | 10 | 10 | 0 | 0 | 1 | 0 | PROVISIONAL / INCOMPLETE | | Video-MME with subtitles | 44.44 | 22.22 | 22.22 | 9 | 9 | 0 | 0 | 0 | 4 | Completed subset |

![Matched subset comparison](evaluations/athena-runall-chatstop-v3/comparison.png)

    [Protocol](evaluations/athena-runall-chatstop-v3/protocol.json) · [Raw responses](evaluations/athena-runall-chatstop-v3/responses.jsonl.gz) · [Results ZIP](evaluations/athena-runall-chatstop-v3/Athena_Benchmark_Results_v3_stopfix.zip) · [Machine-readable results](evaluations/athena-runall-chatstop-v3/report.json)

The report includes paired bootstrap intervals where applicable, all fourteen MMMLU languages, video-cluster uncertainty, identity screening, exact training-image overlap exclusions, and the original v1.5 OmniDoc metrics when all required components score successfully. Scores do not establish general superiority over 4B, 7B or 9B models.

<!-- ATHENA_RUNALL_EVAL_END -->
Downloads last month
5
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Pluto-AI-Labs/Athena-1.0-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(908)
this model

Dataset used to train Pluto-AI-Labs/Athena-1.0-4B