Falcon-OCR · Poneglyph

Recognition of French manga speech bubbles, fine-tuned on validated Poneglyph transcriptions. Input: an already cropped bubble. Output: text. No layout detector is trained.

Which model do the numbers describe?

The root model was retrained from the base checkpoint on 100% of the exported corpus. It has no independent held-out score. All test metrics and analysis images below describe the separate evaluated/ checkpoint, before full-data retraining.

Selection uses validation CER only, with early stopping. Test pages are used once after selection, for both the base model and selected checkpoint. Full-data retraining uses the selected number of epochs and restarts from the base, with no further test-based selection. This procedure does not guarantee that full-data retraining beats the evaluated checkpoint.

Evaluated checkpoint · held-out test Value
CER 0.5261%
Strict CER 0.5261%
WER 1.8267%
Exact match 90.27%
Empty outputs 0.00%
Generation cap reached 0.00%
Test bubbles 1501

CER is corpus edit distance divided by corpus reference length (not average sample CER). Normalization is NFC and whitespace collapse; accents, punctuation and case are preserved. benchmark_test.json includes a 95% bootstrap interval resampling whole pages, length/shape slices, per-bubble predictions, and measured generation time. The split is by page, with exact duplicate crops/pages grouped to avoid leakage. Near-duplicate scans and different pages from the same series can remain related; these results do not measure generalization to completely unseen series.

Inference

import torch
from PIL import Image
from transformers import AutoModelForCausalLM

model = AutoModelForCausalLM.from_pretrained(
    "Remidesbois/Falcon-OCR-Poneglyph", trust_remote_code=True,
    torch_dtype=torch.bfloat16,
).to("cuda").eval()
# Use subfolder="evaluated" to load the independently evaluated checkpoint.
texts = model.generate(Image.open("bubble.png").convert("RGB"),
                       category="plain", min_dimension=64,
                       max_dimension=896, max_new_tokens=256)
print(texts[0])

The upstream generation method can round the token budget up to a block boundary. The benchmark uses an exact cap. inference.py provides the benchmark-compatible preprocessing (white padding for pathological aspect ratios), FP32 master weights with BF16 autocast, and exact-cap decoding. Loading all weights directly in BF16 can introduce additional rounding differences. Remote code is the pinned upstream revision 42ec56b72a23984ac059e7c8a6d397a8529423fe; review it before use.

Training

RTX 5090 profile: PyTorch 2.11 / CUDA 13, full parameter fine-tuning, FP32 master weights and AdamW states, BF16 autocast, SDPA with differentiable attention sinks, gradient checkpointing, cosine schedule, warmup and gradient clipping. Only target text and its stop token are supervised; images and prompts are masked. Loss is averaged per bubble. Conservative photometric augmentation applies only in training. The inference-only Triton MLP is replaced by differentiable PyTorch during training. Published weights retain the original Falcon architecture and inference code.

Exported corpus: 9707 bubbles. Selected training duration: 6 epochs. See run_summary.json, run_config.json, dataset_report.json, environment.txt, and training_code/ for provenance. Training data are validated annotations; empty references and invalid/missing crops are listed in the export accounting. Download errors fail the run instead of silently dropping data. Corpus snapshots and their assignment fingerprints are fixed on resume.

Analysis of the evaluated checkpoint

Training curves Base versus selected checkpoint on the same test Error distribution

Representative examples (largest absolute errors, followed by exact matches):

OCR example

OCR example

OCR example

OCR example

OCR example

OCR example

OCR example

OCR example

Model weights derive from Falcon-OCR. Manga examples retain their respective owners' rights; the model license does not license the source images.

Downloads last month
37
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Remidesbois/Falcon-OCR-Poneglyph

Quantized
(1)
this model