REPAIR-7B: Scientific Dense Retriever

REPAIR (REtriever via Epistemic API-Guided Iterative Refinement) is a self-evolving data-augmentation framework for scientific dense retrievers. It iteratively (1) diagnoses the long-tail scientific concepts the retriever confuses, (2) expands them with fact-verified evidence from scientific APIs (Semantic Scholar, Elsevier, PubChem, Materials Project), and (3) differentiates fine-grained factual distinctions with model-aware hard negatives.

This repository holds the 7B REPAIR retriever as a LoRA adapter on top of Qwen/Qwen2.5-7B (r=16, alpha=32, all attention + MLP projections). Other sizes: 0.5B, 1.5B.

  • Paper: REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement
  • Code (training / data pipeline / eval): https://github.com/yerimoh/REPAIR
  • Authors: Yerim Oh, Gunhee Kim (Seoul National University)
  • Venue: EMNLP 2026

Results (nDCG@10; BIOSSES: Spearman)

Numbers are from the paper's main table. The BMRetriever row is the closest-size baseline.

Model NFCorpus SciFact SciDocs TREC-COVID IR avg. BIOSSES Avg.
BMRetriever-7B 0.364 0.778 0.201 0.861 0.551 0.847 0.610
REPAIR-7B (this repo) 0.413 0.789 0.227 0.842 0.568 0.846 0.623

How to use

The model is a decoder-only embedder with EOS (last-token) pooling. Queries carry a one-line task instruction; passages carry a fixed prefix. Embeddings are L2-normalised and scored by dot product (= cosine).

import torch, torch.nn.functional as F
from transformers import AutoModel, AutoTokenizer
from peft import PeftModel

ADAPTER = "yerim0210/REPAIR_Scientific_Retrievers_7B"
BASE = "Qwen/Qwen2.5-7B"

tok = AutoTokenizer.from_pretrained(ADAPTER)
model = PeftModel.from_pretrained(AutoModel.from_pretrained(BASE, torch_dtype=torch.bfloat16), ADAPTER)
model = model.eval().cuda()
if tok.pad_token is None:
    tok.pad_token = tok.eos_token
    tok.padding_side = "left"

def last_token_pool(h, mask):
    if mask[:, -1].sum() == mask.shape[0]:          # right-padded or no padding
        return h[:, -1]
    idx = mask.sum(dim=1) - 1                       # left padding
    return h[torch.arange(h.shape[0], device=h.device), idx]

@torch.no_grad()
def encode(texts, max_length=512):
    batch = tok(texts, max_length=max_length - 1, truncation=True, padding=False,
                return_attention_mask=False, return_token_type_ids=False)
    batch["input_ids"] = [ids + [tok.eos_token_id] for ids in batch["input_ids"]]   # append EOS
    batch = tok.pad(batch, padding=True, return_attention_mask=True, return_tensors="pt").to("cuda")
    out = model(**batch, use_cache=False)
    return F.normalize(last_token_pool(out.last_hidden_state, batch["attention_mask"]), dim=-1)

task = "Given a query, retrieve passages that are relevant to the query."
queries  = [f"{task}\nQuery: {q}" for q in ["Which protein does venetoclax inhibit?"]]
passages = [f"Represent this passage.\nPassage: {p}" for p in [
    "Venetoclax is a selective BCL-2 inhibitor approved for chronic lymphocytic leukemia.",
    "BCL1 (cyclin D1) rearrangement t(11;14) is the hallmark of mantle cell lymphoma.",
]]
scores = encode(queries) @ encode(passages).T
print(scores)

Task instructions used in the paper's evaluation, for example:

Benchmark Query instruction
default / BEIR Given a query, retrieve passages that are relevant to the query.
TREC-COVID Given a query on COVID-19, retrieve documents that answer the query
PubMedQA Given a question, retrieve relevant PubMed passages that answer the question
BIOSSES Given a sentence, retrieve sentences with the same meaning

Training

  • Backbone: Qwen/Qwen2.5-7B with LoRA (peft), task_type=FEATURE_EXTRACTION, dropout 0.05
  • Objective: contrastive learning with EOS pooling over (query, positive, hard negative) triplets
  • Data: a compact seed set of public scientific corpora (4M pairs in total), iteratively augmented by the REPAIR Diagnosis → Expansion → Differentiation loop; every synthesized pair is grounded in documents returned by external scientific APIs. See the GitHub repository for the full pipeline.

Intended use and limitations

Dense retrieval and sentence similarity over English scientific text (biomedicine, chemistry, materials science). The retriever is not a generator and should not be used to produce text. Performance on general-domain or non-English queries has not been evaluated.

Citation

@misc{oh2026repairresolvinglongtailconfusion,
      title={REPAIR: Resolving Long-Tail Confusion in Scientific Retrievers via Fact-Verified Iterative Refinement},
      author={Yerim Oh and Gunhee Kim},
      year={2026},
      booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
}
Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yerim0210/REPAIR_Scientific_Retrievers_7B

Base model

Qwen/Qwen2.5-7B
Adapter
(870)
this model

Datasets used to train yerim0210/REPAIR_Scientific_Retrievers_7B