SentenceTransformer based on dmis-lab/biobert-base-cased-v1.2

This is a sentence-transformers model finetuned from dmis-lab/biobert-base-cased-v1.2. It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.

Model Details

Model Description

  • Model Type: Sentence Transformer
  • Base model: dmis-lab/biobert-base-cased-v1.2
  • Maximum Sequence Length: 512 tokens
  • Output Dimensionality: 768 dimensions
  • Similarity Function: Cosine Similarity
  • Supported Modality: Text

Model Sources

Full Model Architecture

SentenceTransformer(
  (0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'unpad_inputs': True, 'architecture': 'BertModel'})
  (1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)

Usage

Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

pip install -U sentence-transformers

Then you can load this model and run inference.

from sentence_transformers import SentenceTransformer

# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
queries = [
    'collagen biosynthetic process',
]
documents = [
    'Cleaves the propeptides of type II collagen prior to fibril assembly. Does not act on types I and III collagens.',
    'Catalyzes the decarboxylation of L-3,4-dihydroxyphenylalanine (DOPA) to dopamine and L-5-hydroxytryptophan to serotonin. .',
    'Participates in the reverse transport of cholesterol from tissues to the liver for excretion by promoting cholesterol efflux from tissues and by acting as a cofactor for the lecithin cholesterol acyltransferase (LCAT). As part of the SPAP complex, activates spermatozoa motility. .',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]

# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.5905, 0.1095, 0.1499]])

Evaluation

Metrics

Information Retrieval

Metric Value
cosine_accuracy@1 0.2257
cosine_accuracy@3 0.3944
cosine_accuracy@5 0.483
cosine_accuracy@10 0.5934
cosine_precision@1 0.2257
cosine_precision@3 0.2071
cosine_precision@5 0.1947
cosine_precision@10 0.1653
cosine_recall@1 0.0257
cosine_recall@3 0.0685
cosine_recall@5 0.1033
cosine_recall@10 0.1625
cosine_ndcg@10 0.2135
cosine_mrr@10 0.3313
cosine_map@100 0.142

Training Details

How to read the training samples

Negatives in this dataset are hard by construction. Pairs that look almost identical are the training signal, not a data error: the model is being taught to separate proteins whose descriptions differ in one clause.

source triplets share what it is
sibling term 50,953 85.2% a protein annotated to a sibling GO term — one sharing an is_a parent with the anchor. These share pathway boilerplate and differ only where the label depends on it.
curated NOT 6,858 11.5% a GO NOT annotation, propagated to descendant terms. Documented absence of the function, not inferred.
random 1,981 3.3% drawn from the unblocked pool. Easy negatives.
total 59,792

A worked example

For the anchor alpha-1,6-mannosyltransferase activity:

protein catalyses
positive Q9BV10 ALG12 adds the 8th mannose in alpha-1,6 linkage
negative Q9H6U8 ALG9 adds the 7th and 9th in alpha-1,2 linkage

The two UniProt descriptions share 81% of their characters — the dolichol-linked oligosaccharide pathway is described identically for both — and diverge only in the linkage specificity. The anchor names alpha-1,6, so ALG12 belongs and ALG9 does not. Telling these apart is the task.

What is excluded

A protein with any evidence linking it to the anchor is excluded from all three negative pools, including evidence too weak to use as a positive (94,891 annotations were dropped as weak but still block). Where a protein carried both a trusted positive and a contradicting NOT, the positive wins and the NOT is discarded (417 cases).

Limitation

Sibling negatives assume GO's annotation is complete for the anchor. Where a protein has the function but has not been annotated with it, a true positive can be drawn as a negative, and nothing in the data distinguishes that from a correct hard negative. This is inherent to using an ontology as supervision.

Training Dataset

Unnamed Dataset

  • Size: 59,792 training samples
  • Columns: anchor, positive, and negative
  • Approximate statistics based on the first 100 samples:
    anchor positive negative
    type string string string
    modality text text text
    details
    • min: 26 characters
    • mean: 31.83 characters
    • max: 41 characters
    • min: 92 characters
    • mean: 1680.79 characters
    • max: 13660 characters
    • min: 73 characters
    • mean: 1254.81 characters
    • max: 6393 characters
  • Samples:
    anchor positive negative
    alpha-1,6-mannosyltransferase activity Mannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. In the lumen of the endoplasmic reticulum, adds the eighth mannose residue in an alpha-1,6 linkage onto Man(7)GlcNAc(2)-PP-dolichol to produce Man(8)GlcNAc(2)-PP-dolichol. . Mannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. In the lumen of the endoplasmic reticulum, catalyzes the addition of the seventh and ninth alpha-1,2-linked mannose residues to Man(6)GlcNAc(2)-PP-dolichol and Man(8)GlcNAc(2)-PP-dolichol respectively. .
    alpha-1,6-mannosyltransferase activity Alpha-1,6-mannosyltransferase that catalyzes the transfer of the second mannose, via an alpha-1,6 bond, from a dolichol-phosphate-mannose (Dol-P-Man) to a 2-acyl-6-(alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl)-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H2) intermediate to generate a 2-acyl-6-[alpha-D-mannosyl-(1->6)-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H3) and participates in the seventh step of the glycosylphosphatidylinositol-anchor biosynthesis (Probable) (PubMed:15623507). May also transfer the second mannose on a 2-acyl-6-[2-phosphoethanolamine-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H5), but less effectively (Probable). . Dol-P-Man:Man(5)GlcNAc(2)-PP-Dol alpha-1,3-mannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. In the lumen of the endoplasmic reticulum, adds the first dolichyl beta-D-mannosyl phosphate derived mannose in an alpha-1,3 linkage to Man(5)GlcNAc(2)-PP-dolichol to produce Man(6)GlcNAc(2)-PP-dolichol (PubMed:10581255). Man(6)GlcNAc(2)-PP-dolichol is a substrate for ALG9, the following enzyme in the biosynthetic pathway (PubMed:1058125...
    alpha-1,6-mannosyltransferase activity Mannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. Catalyzes, on the cytoplasmic face of the endoplasmic reticulum, the addition of the second and third mannose residues to the dolichol-linked oligosaccharide chain, to produce Man3GlcNAc(2)-PP-dolichol core oligosaccharide. Man3GlcNAc(2)-PP-dolichol is a substrate for ALG11, the following enzyme in the biosynthetic pathway (PubMed:12684507, PubMed:35136180). While both alpha 1,3 ... Alpha-1,2-mannosyltransferase that catalyzes the transfer of the third mannose, via an alpha-1,2 bond, from a dolichol-phosphate-mannose (Dol-P-Man) to a 2-acyl-6-[alpha-D-mannosyl-(1->6)-2-phosphoethanolamine-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol intermediate to generate a 2-acyl-6-[alpha-D-mannosyl-(1->2)-alpha-D-mannosyl-(1->6)-2-phosphoethanolamine-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H6) and participates in the nineth step of the glycosylphosphatidylinositol-anchor biosynthesis (Probable). May also add the third mannose to a 2-acyl-6-[alpha-D-mannosyl-(1->6)-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H3) intermediate generating a 2-acyl-6-(alpha-D-mannosyl-(1->2)-alpha-D-mannosyl-(1->6)-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl)-1-(1-radyl,2-acyl-sn-glycero-3-phospho)...
  • Loss: CachedMultipleNegativesRankingLoss with these parameters:
    {
        "scale": 20.0,
        "similarity_fct": "cos_sim",
        "mini_batch_size": 32,
        "mini_batch_num_tokens": 65536,
        "gather_across_devices": false,
        "directions": [
            "query_to_doc"
        ],
        "partition_mode": "joint",
        "hardness_mode": null,
        "hardness_strength": 0.0
    }
    

Training Hyperparameters

Non-Default Hyperparameters

  • per_device_train_batch_size: 1024
  • num_train_epochs: 10
  • learning_rate: 0.0001
  • warmup_steps: 0.1
  • bf16: True
  • load_best_model_at_end: True
  • batch_sampler: no_duplicates

All Hyperparameters

Click to expand
  • per_device_train_batch_size: 1024
  • num_train_epochs: 10
  • max_steps: -1
  • learning_rate: 0.0001
  • lr_scheduler_type: linear
  • lr_scheduler_kwargs: None
  • warmup_steps: 0.1
  • optim: adamw_torch_fused
  • optim_args: None
  • weight_decay: 0.0
  • adam_beta1: 0.9
  • adam_beta2: 0.999
  • adam_epsilon: 1e-08
  • optim_target_modules: None
  • gradient_accumulation_steps: 1
  • average_tokens_across_devices: True
  • max_grad_norm: 1.0
  • label_smoothing_factor: 0.0
  • bf16: True
  • fp16: False
  • bf16_full_eval: False
  • fp16_full_eval: False
  • tf32: None
  • gradient_checkpointing: False
  • gradient_checkpointing_kwargs: None
  • torch_compile: False
  • torch_compile_backend: None
  • torch_compile_mode: None
  • use_liger_kernel: False
  • liger_kernel_config: None
  • use_cache: False
  • neftune_noise_alpha: None
  • torch_empty_cache_steps: None
  • auto_find_batch_size: False
  • log_on_each_node: True
  • logging_nan_inf_filter: True
  • include_num_input_tokens_seen: no
  • log_level: passive
  • log_level_replica: warning
  • disable_tqdm: False
  • project: huggingface
  • trackio_space_id: None
  • trackio_bucket_id: None
  • trackio_static_space_id: None
  • per_device_eval_batch_size: 8
  • prediction_loss_only: True
  • eval_on_start: False
  • eval_do_concat_batches: True
  • eval_use_gather_object: False
  • eval_accumulation_steps: None
  • include_for_metrics: []
  • batch_eval_metrics: False
  • save_only_model: False
  • save_on_each_node: False
  • enable_jit_checkpoint: False
  • push_to_hub: False
  • hub_private_repo: None
  • hub_model_id: None
  • hub_strategy: every_save
  • hub_always_push: False
  • hub_revision: None
  • load_best_model_at_end: True
  • ignore_data_skip: False
  • restore_callback_states_from_checkpoint: False
  • full_determinism: False
  • seed: 42
  • data_seed: None
  • use_cpu: False
  • accelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}
  • parallelism_config: None
  • dataloader_drop_last: False
  • dataloader_num_workers: 0
  • dataloader_pin_memory: True
  • dataloader_persistent_workers: False
  • dataloader_prefetch_factor: None
  • dataloader_multiprocessing_context: None
  • dataloader_in_order: True
  • remove_unused_columns: True
  • label_names: None
  • train_sampling_strategy: random
  • length_column_name: length
  • ddp_find_unused_parameters: None
  • ddp_bucket_cap_mb: None
  • ddp_broadcast_buffers: False
  • ddp_static_graph: None
  • ddp_backend: None
  • ddp_timeout: 1800
  • fsdp: None
  • fsdp_config: None
  • deepspeed: None
  • debug: []
  • skip_memory_metrics: True
  • do_predict: False
  • resume_from_checkpoint: None
  • local_rank: -1
  • prompts: None
  • batch_sampler: no_duplicates
  • multi_dataset_batch_sampler: proportional
  • router_mapping: {}
  • learning_rate_mapping: {}
  • warmup_ratio: None

Training Logs

Epoch Step Training Loss heldout_go_cosine_ndcg@10
-1 -1 - 0.0285
0.3390 20 7.3693 -
0.6780 40 6.2354 -
1.0 59 - 0.1549
1.0169 60 5.3803 -
1.3559 80 5.3655 -
1.6949 100 5.1811 -
2.0 118 - 0.1886
2.0339 120 4.7409 -
2.3729 140 4.7735 -
2.7119 160 4.7011 -
3.0 177 - 0.2067
3.0508 180 4.4012 -
3.3898 200 4.4621 -
3.7288 220 4.4667 -
4.0 236 - 0.2075
4.0678 240 4.1357 -
4.4068 260 4.2363 -
4.7458 280 4.2808 -
5.0 295 - 0.2086
5.0847 300 3.9821 -
5.4237 320 4.1327 -
5.7627 340 4.1406 -
6.0 354 - 0.2106
6.1017 360 3.8678 -
6.4407 380 4.0266 -
6.7797 400 4.0569 -
7.0 413 - 0.2114
7.1186 420 3.7898 -
7.4576 440 3.9818 -
7.7966 460 4.0110 -
8.0 472 - 0.2131
8.1356 480 3.7468 -
8.4746 500 3.9718 -
8.8136 520 3.9436 -
9.0 531 - 0.2130
9.1525 540 3.7193 -
9.4915 560 3.9474 -
9.8305 580 3.9401 -
10.0 590 - 0.2134
-1 -1 - 0.2135
  • The bold row denotes the saved checkpoint.

Training Time

  • Training: 41.7 minutes
  • Evaluation: 1.3 minutes
  • Total: 43.0 minutes

Framework Versions

  • Python: 3.13.15
  • Sentence Transformers: 6.1.0
  • Transformers: 5.18.0
  • PyTorch: 2.11.0+cu130
  • Accelerate: 1.15.0
  • Datasets: 4.8.5
  • Tokenizers: 0.23.2

Additional Resources

Citation

BibTeX

Sentence Transformers

@inproceedings{reimers-2019-sentence-bert,
    title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
    author = "Reimers, Nils and Gurevych, Iryna",
    booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
    month = "11",
    year = "2019",
    publisher = "Association for Computational Linguistics",
    url = "https://arxiv.org/abs/1908.10084",
}

CachedMultipleNegativesRankingLoss

@misc{gao2021scaling,
    title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
    author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
    year={2021},
    eprint={2101.06983},
    archivePrefix={arXiv},
    primaryClass={cs.LG}
}

MultipleNegativesRankingLoss

@misc{oord2019representationlearningcontrastivepredictive,
      title={Representation Learning with Contrastive Predictive Coding},
      author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
      year={2019},
      eprint={1807.03748},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/1807.03748},
}
Downloads last month
27
Safetensors
Model size
0.1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for yoyo458/BioBERT_GOUniProt_FineTuned

Finetuned
(36)
this model

Papers for yoyo458/BioBERT_GOUniProt_FineTuned

Evaluation results