Instructions to use ratishsp/progen2-base-bidirectional-llm2vec with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ratishsp/progen2-base-bidirectional-llm2vec with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("hugohrban/progen2-base") model = PeftModel.from_pretrained(base_model, "ratishsp/progen2-base-bidirectional-llm2vec") - Notebooks
- Google Colab
- Kaggle
ProGen2-base β Bidirectional Protein Encoder (LLM2Vec recipe)
A LoRA adapter that converts the generative (decoder-only) protein language model
hugohrban/progen2-base into a
bidirectional sequence encoder for producing fixed-length protein embeddings,
following the LLM2Vec recipe adapted to proteins.
The adaptation has three ingredients:
- Bidirectional attention: the causal triangular attention mask is removed so every residue attends to the whole sequence.
- Two-stage training: masked next-token prediction (MNTP) first, then a separate SimCSE contrastive stage (same sequence encoded twice under dropout = a positive pair), resuming the MNTP adapter.
- LoRA: only low-rank adapters are trained; the base weights are frozen.
Results: full benchmark (frozen encoder + linear probe)
Evaluated on all 9 tasks from the TAPE (Rao et al. 2019) and ProteinBERT (Brandes et al. 2022) benchmark suites. Protocol: freeze the encoder, take a fixed representation (mean-pooled for sequence-level tasks, per-residue for token-level), fit a closed-form linear probe on the train split, and score the test split. "Baseline" = the same bidirectional ProGen2 without this adapter.
Validation vs. held-out. One of the nine tasks, Stability (a protein-stability regression benchmark from TAPE), was used to pick the recipe: across a 10-config sweep, the hyperparameters scoring highest on Stability were kept. The other 8 tasks were then evaluated once on that fixed configuration, so their numbers are held-out generalization, not values the recipe was tuned to maximize.
Sequence-level
| Task | metric | baseline | this adapter | Ξ |
|---|---|---|---|---|
| Stability | Spearman Ο | 0.403 | 0.661 | +0.258 |
| Fluorescence | Spearman Ο | 0.157 | 0.367 | +0.210 |
| Remote homology | accuracy | 0.063 | 0.102 | +0.039 |
| Fold class. | accuracy | 0.099 | 0.210 | +0.112 |
| Signal peptide | accuracy | 0.838 | 0.931 | +0.093 |
| Neuropeptide | accuracy | 0.691 | 0.932 | +0.240 |
Token-level (per-residue)
| Task | metric | baseline | this adapter | Ξ |
|---|---|---|---|---|
| Secondary structure (SS3) | Q3 accuracy | 0.535 | 0.590 | +0.056 |
| PTM (phosphosite) | ROC-AUC | 0.908 | 0.920 | +0.012 |
| Disorder | ROC-AUC | 0.745 | 0.822 | +0.077 |
The adapter improves the frozen representation on every task, at both the sequence and per-residue level. PTM and disorder are reported as ROC-AUC because they are heavily imbalanced (majority class β 0.98), making raw accuracy uninformative. Absolute values are those of a frozen linear probe (an embedding-quality measurement), not task-specific fine-tuning.
Usage
import torch, torch.nn.functional as F
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
BASE = "hugohrban/progen2-base"
ADAPTER = "ratishsp/progen2-base-bidirectional-llm2vec" # this repo
tok = AutoTokenizer.from_pretrained(BASE, trust_remote_code=True)
if tok.pad_token_id is None:
tok.pad_token = tok.eos_token
model = AutoModelForCausalLM.from_pretrained(
BASE, trust_remote_code=True, torch_dtype=torch.bfloat16,
attn_implementation="eager", # required: the bidirectional mask edit needs eager attention
)
# 1) make attention bidirectional: flip each attention module's causal `bias` buffer to all-True
for m in model.modules():
b = getattr(m, "bias", None)
if isinstance(b, torch.Tensor) and b.dtype == torch.bool:
b.fill_(True)
# 2) load the LoRA adapter
model = PeftModel.from_pretrained(model, ADAPTER)
model.eval().cuda()
@torch.no_grad()
def embed(seqs):
enc = tok(seqs, padding=True, truncation=True, max_length=512, return_tensors="pt").to("cuda")
h = model(input_ids=enc.input_ids, attention_mask=enc.attention_mask,
output_hidden_states=True).hidden_states[-1]
mask = enc.attention_mask.unsqueeze(-1).float()
pooled = (h * mask).sum(1) / mask.sum(1).clamp(min=1) # mean pool
return F.normalize(pooled.float(), dim=-1)
emb = embed(["MKTAYIAKQR", "MVLSPADKTNVKAAW"])
The full, tested implementation is bundled in this repo under code/: training
(pretrain.py), the bidirectional conversion and objectives (src/bidir_progen.py:
make_bidirectional, mean_pool), and all benchmark evals (eval_protein.py, eval_token.py).
Training details
| Base model | hugohrban/progen2-base (764M, BSD-3-Clause) |
| Pretraining data | UniRef50 (agemagician/uniref50), streamed |
| Stage 1 (MNTP) | 2,000 steps, max-len 512 |
| Stage 2 (SimCSE) | 2,000 steps, temperature 0.10, dropout 0.1, max-len 256 |
| LoRA | r=16, Ξ±=32, targets: qkv_proj, out_proj, fc_in, fc_out |
| Optimizer | AdamW, lr 1e-4, warmup 50 |
| Precision | bf16 |
This was the best of a 10-configuration sweep. Key findings: the MNTP stage is necessary (SimCSE-only transfers far worse), SimCSE temperature 0.10 beat 0.05/0.02, and the short recipe outperformed longer / larger-data runs.
Limitations
- Numbers come from a frozen linear probe, not task fine-tuning; they measure embedding quality, so absolute values trail task-specific SOTA (e.g. SS3 Q3 β 0.59 vs ~0.8+ for dedicated predictors).
- Sequence-level uses a single mean-pooled vector; no specialised pooling head.
- No per-task hyperparameter tuning: the recipe was tuned once, on Stability (see Results).
Training data summary
This section follows the structure of the European Commission's template for the public summary of training content (AI Act, Article 53(1)(d)). This adapter is a narrow protein-embedding model, not a general-purpose AI model.
Provider. Ratish Puduppully, IT University of Copenhagen, rapu@itu.dk.
Model covered. ratishsp/progen2-base-bidirectional-llm2vec, a LoRA adapter of about 10M trainable parameters.
Model dependencies. The adapter modifies hugohrban/progen2-base, a Hugging Face port of ProGen2-base (764M parameters, Nijkamp et al. 2023, Salesforce). ProGen2 is a protein language model trained on UniRef90 and BFD30 protein sequences. The base weights are frozen and are not redistributed in this repository. No public training-content summary exists for ProGen2, which predates the AI Act.
Modality and size. Other modality, protein amino-acid sequences written as single-letter strings, with no natural language. Both training stages read the first 150,000 sequences of the UniRef50 training split, under 1 percent of the dataset. Each stage runs 2,000 steps. The masked-prediction stage uses batch 32 with sequences truncated to 512 residues, and the SimCSE stage uses batch 64 truncated to 256 residues. The data was streamed from Hugging Face at training time in June 2026. The adapter was trained once and has not been updated since.
Data sources. One publicly available dataset, agemagician/uniref50 on Hugging Face. It is a redistribution of UniRef50 from the UniProt consortium under CC BY 4.0. No licensed data, no private third-party data, no crawled or scraped content, no user data and no AI-generated data were used. The benchmark datasets named in the results section were used only for evaluation.
Data processing. Rights reservation, not applicable. UniProt publishes UniRef under CC BY 4.0, and protein sequences are not works in the copyright sense. Illegal content, not applicable. The data contains only amino-acid strings. Preprocessing is limited to truncation and tokenisation with the ProGen2 tokenizer.
License & provenance
- Adapter weights: BSD-3-Clause, matching the
progen2-basebase model. - Bundled code (
code/): MIT; seecode/LICENSE.
Research artifact produced on the DCAI Gefion cluster. The base model and all benchmark datasets (TAPE via AI4Protein/GleghornLab mirrors; ProteinBERT signal-peptide/neuropeptide via GrimSqueaker mirrors; PTM/disorder from the ProteinBERT data repo) are publicly available, so the results are reproducible from public sources.
Framework versions
- PEFT 0.12.0
- transformers 4.44.2
- Downloads last month
- 16
Model tree for ratishsp/progen2-base-bidirectional-llm2vec
Base model
hugohrban/progen2-base