GLM-OCR Field + BBox Extraction (LoRA, full-scale)

Fine-tuned from zai-org/GLM-OCR to jointly extract structured fields and their bounding boxes from prescription document images. Given a prescription photo, the model returns a single JSON object with every field below, each with its on-page location.

This is a LoRA adapter merged into the base weights (rank 64, alpha 128, dropout 0.05, targeting the vision encoder attention, vision-language merger, and LM attention/MLP layers). A sibling full fine-tune of the same task is at KeraCare/glm-ocr-field-bbox-full-ft — in our evaluation this LoRA checkpoint outperformed it on every metric, at both pilot and full scale.

Training data

17,192 prescription images, Kimi-model-generated field+bbox annotations (partially manually verified), 3 epochs. Trained on an NVIDIA GB10 (DGX Spark).

Output format

A single JSON object:

  • is_prescription (bool), declared_med_count (int)
  • structure_name, structure_phone, structure_address, structure_email: clinic/pharmacy header fields
  • prescriber_name, prescriber_specialty
  • patient_name, patient_age_dob, date
  • stamp, signature: region only, text is always null
  • header_logo, other
  • medications: list of {name, dosage, frequency, quantity, duration, drug_type}

Every field except is_prescription/declared_med_count is a list of {"text": <string or null>, "bbox": [x1, y1, x2, y2]} (empty list if the field doesn't appear on the document). bbox is normalized to image width/height and scaled to the 0-1000 integer range (not 0-1 floats, not pixels) — x_pixel = x_bbox * image_width / 1000.

Evaluation (n=476, held-out, zero overlap with training data)

Metric This model (LoRA) Full-FT sibling
is_prescription accuracy 99.8% 99.4%
declared_med_count exact match 94.3% 90.5%
Micro text-F1 @ fuzzy≥0.7 79.2% 75.2%
Micro geometry-F1 @ IoU≥0.5 65.6% 57.8%
Macro mean IoU 0.597 0.536

Inference

import json
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

MODEL_PATH = "KeraCare/glm-ocr-field-bbox-lora"
IMAGE_PATH = "prescription.jpg"

processor = AutoProcessor.from_pretrained(MODEL_PATH, use_fast=False)
model = AutoModelForImageTextToText.from_pretrained(
    MODEL_PATH, torch_dtype=torch.bfloat16, device_map="cuda",
)

# Match the resolution used during training.
processor.image_processor.size = {"longest_edge": 1_048_576, "shortest_edge": 1024}

PROMPT = """Extract every field below from this prescription image, including the on-page location
(bounding box) of each piece of text you find.

Return a single JSON object with these top-level keys:
- is_prescription: boolean, whether the image is a prescription.
- declared_med_count: integer, the number of medications listed on the prescription.
- structure_name, structure_phone, structure_address, structure_email: lists of
  {"text": <string>, "bbox": [x1, y1, x2, y2]} for the clinic/pharmacy header.
- prescriber_name, prescriber_specialty: lists of {"text": <string>, "bbox": [...]}.
- patient_name, patient_age_dob, date: lists of {"text": <string>, "bbox": [...]}.
- stamp, signature: lists of {"text": null, "bbox": [...]} marking the region only (text is always null).
- header_logo, other: lists of {"text": <string or null>, "bbox": [...]} for any logo or unclassified element.
- medications: a list of objects, each with keys name, dosage, frequency, quantity, duration
  (each {"text": <string or null>, "bbox": [x1, y1, x2, y2] or null}) and drug_type (string or null).

bbox is [x1, y1, x2, y2]: the top-left and bottom-right corners, normalized to the image width/height
and scaled to the 0-1000 integer range. If a field does not appear on the prescription, return an
empty list for it. Output ONLY the JSON object, no other text.
"""

image = Image.open(IMAGE_PATH).convert("RGB")
messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": IMAGE_PATH},
            {"type": "text", "text": PROMPT},
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
inputs = processor(text=text, images=[[image]], return_tensors="pt").to(model.device)

generated_ids = model.generate(**inputs, max_new_tokens=2048)
output_text = processor.decode(
    generated_ids[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True
)
result = json.loads(output_text)
print(json.dumps(result, indent=2, ensure_ascii=False))

No system prompt is used — only the single user turn (image + the instruction above) and the model's JSON completion, matching the exact format used during training.

Downloads last month
123
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for KeraCare/glm-ocr-field-bbox-lora

Base model

zai-org/GLM-OCR
Adapter
(13)
this model