GLM-OCR Field + BBox Extraction (LoRA, full-scale)
Fine-tuned from zai-org/GLM-OCR to jointly extract structured fields and their bounding boxes from prescription document images. Given a prescription photo, the model returns a single JSON object with every field below, each with its on-page location.
This is a LoRA adapter merged into the base weights (rank 64, alpha 128, dropout 0.05, targeting the vision encoder attention, vision-language merger, and LM attention/MLP layers). A sibling full fine-tune of the same task is at KeraCare/glm-ocr-field-bbox-full-ft — in our evaluation this LoRA checkpoint outperformed it on every metric, at both pilot and full scale.
Training data
17,192 prescription images, Kimi-model-generated field+bbox annotations (partially manually verified), 3 epochs. Trained on an NVIDIA GB10 (DGX Spark).
Output format
A single JSON object:
is_prescription(bool),declared_med_count(int)structure_name,structure_phone,structure_address,structure_email: clinic/pharmacy header fieldsprescriber_name,prescriber_specialtypatient_name,patient_age_dob,datestamp,signature: region only, text is alwaysnullheader_logo,othermedications: list of{name, dosage, frequency, quantity, duration, drug_type}
Every field except is_prescription/declared_med_count is a list of {"text": <string or null>, "bbox": [x1, y1, x2, y2]} (empty list if the field doesn't appear on the document). bbox is normalized to image width/height and scaled to the 0-1000 integer range (not 0-1 floats, not pixels) — x_pixel = x_bbox * image_width / 1000.
Evaluation (n=476, held-out, zero overlap with training data)
| Metric | This model (LoRA) | Full-FT sibling |
|---|---|---|
is_prescription accuracy |
99.8% | 99.4% |
declared_med_count exact match |
94.3% | 90.5% |
| Micro text-F1 @ fuzzy≥0.7 | 79.2% | 75.2% |
| Micro geometry-F1 @ IoU≥0.5 | 65.6% | 57.8% |
| Macro mean IoU | 0.597 | 0.536 |
Inference
import json
from PIL import Image
import torch
from transformers import AutoProcessor, AutoModelForImageTextToText
MODEL_PATH = "KeraCare/glm-ocr-field-bbox-lora"
IMAGE_PATH = "prescription.jpg"
processor = AutoProcessor.from_pretrained(MODEL_PATH, use_fast=False)
model = AutoModelForImageTextToText.from_pretrained(
MODEL_PATH, torch_dtype=torch.bfloat16, device_map="cuda",
)
# Match the resolution used during training.
processor.image_processor.size = {"longest_edge": 1_048_576, "shortest_edge": 1024}
PROMPT = """Extract every field below from this prescription image, including the on-page location
(bounding box) of each piece of text you find.
Return a single JSON object with these top-level keys:
- is_prescription: boolean, whether the image is a prescription.
- declared_med_count: integer, the number of medications listed on the prescription.
- structure_name, structure_phone, structure_address, structure_email: lists of
{"text": <string>, "bbox": [x1, y1, x2, y2]} for the clinic/pharmacy header.
- prescriber_name, prescriber_specialty: lists of {"text": <string>, "bbox": [...]}.
- patient_name, patient_age_dob, date: lists of {"text": <string>, "bbox": [...]}.
- stamp, signature: lists of {"text": null, "bbox": [...]} marking the region only (text is always null).
- header_logo, other: lists of {"text": <string or null>, "bbox": [...]} for any logo or unclassified element.
- medications: a list of objects, each with keys name, dosage, frequency, quantity, duration
(each {"text": <string or null>, "bbox": [x1, y1, x2, y2] or null}) and drug_type (string or null).
bbox is [x1, y1, x2, y2]: the top-left and bottom-right corners, normalized to the image width/height
and scaled to the 0-1000 integer range. If a field does not appear on the prescription, return an
empty list for it. Output ONLY the JSON object, no other text.
"""
image = Image.open(IMAGE_PATH).convert("RGB")
messages = [
{
"role": "user",
"content": [
{"type": "image", "url": IMAGE_PATH},
{"type": "text", "text": PROMPT},
],
}
]
text = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True, enable_thinking=False,
)
inputs = processor(text=text, images=[[image]], return_tensors="pt").to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=2048)
output_text = processor.decode(
generated_ids[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True
)
result = json.loads(output_text)
print(json.dumps(result, indent=2, ensure_ascii=False))
No system prompt is used — only the single user turn (image + the instruction above) and the model's JSON completion, matching the exact format used during training.
- Downloads last month
- 123
Model tree for KeraCare/glm-ocr-field-bbox-lora
Base model
zai-org/GLM-OCR