Instructions to use yerim0210/REPAIR_Scientific_Retrievers_7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use yerim0210/REPAIR_Scientific_Retrievers_7B with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen2.5-7B") model = PeftModel.from_pretrained(base_model, "yerim0210/REPAIR_Scientific_Retrievers_7B") - Notebooks
- Google Colab
- Kaggle
REPAIR-7B: Scientific Dense Retriever
REPAIR (REtriever via Epistemic API-Guided Iterative Refinement) is a self-evolving data-augmentation framework for scientific dense retrievers. It iteratively (1) diagnoses the long-tail scientific concepts the retriever confuses, (2) expands them with fact-verified evidence from scientific APIs (Semantic Scholar, Elsevier, PubChem, Materials Project), and (3) differentiates fine-grained factual distinctions with model-aware hard negatives.
This repository holds the 7B REPAIR retriever as a LoRA adapter on top of
Qwen/Qwen2.5-7B (r=16, alpha=32, all attention + MLP projections).
Other sizes: 0.5B, 1.5B.
- Paper: REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement
- Code (training / data pipeline / eval): https://github.com/yerimoh/REPAIR
- Authors: Yerim Oh, Gunhee Kim (Seoul National University)
- Venue: EMNLP 2026
Results (nDCG@10; BIOSSES: Spearman)
Numbers are from the paper's main table. The BMRetriever row is the closest-size baseline.
| Model | NFCorpus | SciFact | SciDocs | TREC-COVID | IR avg. | BIOSSES | Avg. |
|---|---|---|---|---|---|---|---|
| BMRetriever-7B | 0.364 | 0.778 | 0.201 | 0.861 | 0.551 | 0.847 | 0.610 |
| REPAIR-7B (this repo) | 0.413 | 0.789 | 0.227 | 0.842 | 0.568 | 0.846 | 0.623 |
How to use
The model is a decoder-only embedder with EOS (last-token) pooling. Queries carry a one-line task instruction; passages carry a fixed prefix. Embeddings are L2-normalised and scored by dot product (= cosine).
import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel
ADAPTER = "yerim0210/REPAIR_Scientific_Retrievers_7B"
BASE = "Qwen/Qwen2.5-7B"
tok = AutoTokenizer.from_pretrained(ADAPTER)
model = PeftModel.from_pretrained(AutoModel.from_pretrained(BASE, torch_dtype=torch.bfloat16), ADAPTER)
model = model.eval().cuda()
if tok.pad_token is None:
tok.pad_token = tok.eos_token
tok.padding_side = "left"
def last_token_pool(h, mask):
if mask[:, -1].sum() == mask.shape[0]: # right-padded or no padding
return h[:, -1]
idx = mask.sum(dim=1) - 1 # left padding
return h[torch.arange(h.shape[0], device=h.device), idx]
@torch.no_grad()
def encode(texts, max_length=512):
batch = tok(texts, max_length=max_length - 1, truncation=True, padding=False,
return_attention_mask=False, return_token_type_ids=False)
batch["input_ids"] = [ids + [tok.eos_token_id] for ids in batch["input_ids"]] # append EOS
batch = tok.pad(batch, padding=True, return_attention_mask=True, return_tensors="pt").to("cuda")
out = model(**batch, use_cache=False)
return F.normalize(last_token_pool(out.last_hidden_state, batch["attention_mask"]), dim=-1)
task = "Given a query, retrieve passages that are relevant to the query."
queries = [f"{task}\nQuery: {q}" for q in ["Which protein does venetoclax inhibit?"]]
passages = [f"Represent this passage.\nPassage: {p}" for p in [
"Venetoclax is a selective BCL-2 inhibitor approved for chronic lymphocytic leukemia.",
"BCL1 (cyclin D1) rearrangement t(11;14) is the hallmark of mantle cell lymphoma.",
]]
scores = encode(queries) @ encode(passages).T
print(scores)
Task instructions used in the paper's evaluation, for example:
| Benchmark | Query instruction |
|---|---|
| default / BEIR | Given a query, retrieve passages that are relevant to the query. |
| TREC-COVID | Given a query on COVID-19, retrieve documents that answer the query |
| PubMedQA | Given a question, retrieve relevant PubMed passages that answer the question |
| BIOSSES | Given a sentence, retrieve sentences with the same meaning |
Training
- Backbone: Qwen/Qwen2.5-7B with LoRA (
peft),task_type=FEATURE_EXTRACTION, dropout 0.05 - Objective: contrastive learning with EOS pooling over (query, positive, hard negative) triplets
- Data: a compact seed set of public scientific corpora (4M pairs in total), iteratively augmented by the REPAIR Diagnosis → Expansion → Differentiation loop; every synthesized pair is grounded in documents returned by external scientific APIs. See the GitHub repository for the full pipeline.
Intended use and limitations
Dense retrieval and sentence similarity over English scientific text (biomedicine, chemistry, materials science). The retriever is not a generator and should not be used to produce text. Performance on general-domain or non-English queries has not been evaluated.
Citation
@misc{oh2026repairresolvinglongtailconfusion,
title={REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement},
author={Yerim Oh and Gunhee Kim},
year={2026},
booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
}
- Downloads last month
- 21
Model tree for yerim0210/REPAIR_Scientific_Retrievers_7B
Base model
Qwen/Qwen2.5-7B