Instructions to use yoyo458/BioBERT_GOUniProt_FineTuned with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use yoyo458/BioBERT_GOUniProt_FineTuned with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("yoyo458/BioBERT_GOUniProt_FineTuned") sentences = [ "metalloendopeptidase activity", "Cysteine protease. May have an important role in corneal physiology. .", "Metalloprotease (PubMed:16585064, PubMed:39672391). Was previously shown to degrade COMP (PubMed:16585064). However, a later study found no activity against COMP (PubMed:39672391). .", "Degrades casein, gelatins of types I, III, IV, and V, and fibronectin. Activates procollagenase. ." ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
SentenceTransformer based on dmis-lab/biobert-base-cased-v1.2
This is a sentence-transformers model finetuned from dmis-lab/biobert-base-cased-v1.2. It maps inputs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, classification, clustering, and more.
Model Details
Model Description
- Model Type: Sentence Transformer
- Base model: dmis-lab/biobert-base-cased-v1.2
- Maximum Sequence Length: 512 tokens
- Output Dimensionality: 768 dimensions
- Similarity Function: Cosine Similarity
- Supported Modality: Text
Model Sources
- Documentation: Sentence Transformers Documentation
- Repository: Sentence Transformers on GitHub
- Hugging Face: Sentence Transformers on Hugging Face
Full Model Architecture
SentenceTransformer(
(0): Transformer({'transformer_task': 'feature-extraction', 'modality_config': {'text': {'method': 'forward', 'method_output_name': 'last_hidden_state'}}, 'module_output_name': 'token_embeddings', 'unpad_inputs': True, 'architecture': 'BertModel'})
(1): Pooling({'embedding_dimension': 768, 'pooling_mode': 'mean', 'include_prompt': True})
)
Usage
Direct Usage (Sentence Transformers)
First install the Sentence Transformers library:
pip install -U sentence-transformers
Then you can load this model and run inference.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("sentence_transformers_model_id")
# Run inference
queries = [
'collagen biosynthetic process',
]
documents = [
'Cleaves the propeptides of type II collagen prior to fibril assembly. Does not act on types I and III collagens.',
'Catalyzes the decarboxylation of L-3,4-dihydroxyphenylalanine (DOPA) to dopamine and L-5-hydroxytryptophan to serotonin. .',
'Participates in the reverse transport of cholesterol from tissues to the liver for excretion by promoting cholesterol efflux from tissues and by acting as a cofactor for the lecithin cholesterol acyltransferase (LCAT). As part of the SPAP complex, activates spermatozoa motility. .',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]
# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.5905, 0.1095, 0.1499]])
Evaluation
Metrics
Information Retrieval
- Dataset:
heldout_go - Evaluated with
InformationRetrievalEvaluator
| Metric | Value |
|---|---|
| cosine_accuracy@1 | 0.2257 |
| cosine_accuracy@3 | 0.3944 |
| cosine_accuracy@5 | 0.483 |
| cosine_accuracy@10 | 0.5934 |
| cosine_precision@1 | 0.2257 |
| cosine_precision@3 | 0.2071 |
| cosine_precision@5 | 0.1947 |
| cosine_precision@10 | 0.1653 |
| cosine_recall@1 | 0.0257 |
| cosine_recall@3 | 0.0685 |
| cosine_recall@5 | 0.1033 |
| cosine_recall@10 | 0.1625 |
| cosine_ndcg@10 | 0.2135 |
| cosine_mrr@10 | 0.3313 |
| cosine_map@100 | 0.142 |
Training Details
How to read the training samples
Negatives in this dataset are hard by construction. Pairs that look almost identical are the training signal, not a data error: the model is being taught to separate proteins whose descriptions differ in one clause.
| source | triplets | share | what it is |
|---|---|---|---|
| sibling term | 50,953 | 85.2% | a protein annotated to a sibling GO term — one sharing an is_a parent with the anchor. These share pathway boilerplate and differ only where the label depends on it. |
curated NOT |
6,858 | 11.5% | a GO NOT annotation, propagated to descendant terms. Documented absence of the function, not inferred. |
| random | 1,981 | 3.3% | drawn from the unblocked pool. Easy negatives. |
| total | 59,792 |
A worked example
For the anchor alpha-1,6-mannosyltransferase activity:
| protein | catalyses | |
|---|---|---|
| positive | Q9BV10 ALG12 |
adds the 8th mannose in alpha-1,6 linkage |
| negative | Q9H6U8 ALG9 |
adds the 7th and 9th in alpha-1,2 linkage |
The two UniProt descriptions share 81% of their characters — the dolichol-linked oligosaccharide pathway is described identically for both — and diverge only in the linkage specificity. The anchor names alpha-1,6, so ALG12 belongs and ALG9 does not. Telling these apart is the task.
What is excluded
A protein with any evidence linking it to the anchor is excluded from all
three negative pools, including evidence too weak to use as a positive
(94,891 annotations were dropped as weak but still block). Where a protein
carried both a trusted positive and a contradicting NOT, the positive wins
and the NOT is discarded (417 cases).
Limitation
Sibling negatives assume GO's annotation is complete for the anchor. Where a protein has the function but has not been annotated with it, a true positive can be drawn as a negative, and nothing in the data distinguishes that from a correct hard negative. This is inherent to using an ontology as supervision.
Training Dataset
Unnamed Dataset
- Size: 59,792 training samples
- Columns:
anchor,positive, andnegative - Approximate statistics based on the first 100 samples:
anchor positive negative type string string string modality text text text details - min: 26 characters
- mean: 31.83 characters
- max: 41 characters
- min: 92 characters
- mean: 1680.79 characters
- max: 13660 characters
- min: 73 characters
- mean: 1254.81 characters
- max: 6393 characters
- Samples:
anchor positive negative alpha-1,6-mannosyltransferase activityMannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. In the lumen of the endoplasmic reticulum, adds the eighth mannose residue in an alpha-1,6 linkage onto Man(7)GlcNAc(2)-PP-dolichol to produce Man(8)GlcNAc(2)-PP-dolichol. .Mannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. In the lumen of the endoplasmic reticulum, catalyzes the addition of the seventh and ninth alpha-1,2-linked mannose residues to Man(6)GlcNAc(2)-PP-dolichol and Man(8)GlcNAc(2)-PP-dolichol respectively. .alpha-1,6-mannosyltransferase activityAlpha-1,6-mannosyltransferase that catalyzes the transfer of the second mannose, via an alpha-1,6 bond, from a dolichol-phosphate-mannose (Dol-P-Man) to a 2-acyl-6-(alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl)-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H2) intermediate to generate a 2-acyl-6-[alpha-D-mannosyl-(1->6)-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H3) and participates in the seventh step of the glycosylphosphatidylinositol-anchor biosynthesis (Probable) (PubMed:15623507). May also transfer the second mannose on a 2-acyl-6-[2-phosphoethanolamine-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H5), but less effectively (Probable). .Dol-P-Man:Man(5)GlcNAc(2)-PP-Dol alpha-1,3-mannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. In the lumen of the endoplasmic reticulum, adds the first dolichyl beta-D-mannosyl phosphate derived mannose in an alpha-1,3 linkage to Man(5)GlcNAc(2)-PP-dolichol to produce Man(6)GlcNAc(2)-PP-dolichol (PubMed:10581255). Man(6)GlcNAc(2)-PP-dolichol is a substrate for ALG9, the following enzyme in the biosynthetic pathway (PubMed:1058125...alpha-1,6-mannosyltransferase activityMannosyltransferase that operates in the biosynthetic pathway of dolichol-linked oligosaccharides, the glycan precursors employed in protein asparagine (N)-glycosylation. The assembly of dolichol-linked oligosaccharides begins on the cytosolic side of the endoplasmic reticulum membrane and finishes in its lumen. The sequential addition of sugars to dolichol pyrophosphate produces dolichol-linked oligosaccharides containing fourteen sugars, including two GlcNAcs, nine mannoses and three glucoses. Once assembled, the oligosaccharide is transferred from the lipid to nascent proteins by oligosaccharyltransferases. Catalyzes, on the cytoplasmic face of the endoplasmic reticulum, the addition of the second and third mannose residues to the dolichol-linked oligosaccharide chain, to produce Man3GlcNAc(2)-PP-dolichol core oligosaccharide. Man3GlcNAc(2)-PP-dolichol is a substrate for ALG11, the following enzyme in the biosynthetic pathway (PubMed:12684507, PubMed:35136180). While both alpha 1,3 ...Alpha-1,2-mannosyltransferase that catalyzes the transfer of the third mannose, via an alpha-1,2 bond, from a dolichol-phosphate-mannose (Dol-P-Man) to a 2-acyl-6-[alpha-D-mannosyl-(1->6)-2-phosphoethanolamine-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol intermediate to generate a 2-acyl-6-[alpha-D-mannosyl-(1->2)-alpha-D-mannosyl-(1->6)-2-phosphoethanolamine-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H6) and participates in the nineth step of the glycosylphosphatidylinositol-anchor biosynthesis (Probable). May also add the third mannose to a 2-acyl-6-[alpha-D-mannosyl-(1->6)-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl]-1-(1-radyl,2-acyl-sn-glycero-3-phospho)-1D-myo-inositol (also termed H3) intermediate generating a 2-acyl-6-(alpha-D-mannosyl-(1->2)-alpha-D-mannosyl-(1->6)-alpha-D-mannosyl-(1->4)-alpha-D-glucosaminyl)-1-(1-radyl,2-acyl-sn-glycero-3-phospho)... - Loss:
CachedMultipleNegativesRankingLosswith these parameters:{ "scale": 20.0, "similarity_fct": "cos_sim", "mini_batch_size": 32, "mini_batch_num_tokens": 65536, "gather_across_devices": false, "directions": [ "query_to_doc" ], "partition_mode": "joint", "hardness_mode": null, "hardness_strength": 0.0 }
Training Hyperparameters
Non-Default Hyperparameters
per_device_train_batch_size: 1024num_train_epochs: 10learning_rate: 0.0001warmup_steps: 0.1bf16: Trueload_best_model_at_end: Truebatch_sampler: no_duplicates
All Hyperparameters
Click to expand
per_device_train_batch_size: 1024num_train_epochs: 10max_steps: -1learning_rate: 0.0001lr_scheduler_type: linearlr_scheduler_kwargs: Nonewarmup_steps: 0.1optim: adamw_torch_fusedoptim_args: Noneweight_decay: 0.0adam_beta1: 0.9adam_beta2: 0.999adam_epsilon: 1e-08optim_target_modules: Nonegradient_accumulation_steps: 1average_tokens_across_devices: Truemax_grad_norm: 1.0label_smoothing_factor: 0.0bf16: Truefp16: Falsebf16_full_eval: Falsefp16_full_eval: Falsetf32: Nonegradient_checkpointing: Falsegradient_checkpointing_kwargs: Nonetorch_compile: Falsetorch_compile_backend: Nonetorch_compile_mode: Noneuse_liger_kernel: Falseliger_kernel_config: Noneuse_cache: Falseneftune_noise_alpha: Nonetorch_empty_cache_steps: Noneauto_find_batch_size: Falselog_on_each_node: Truelogging_nan_inf_filter: Trueinclude_num_input_tokens_seen: nolog_level: passivelog_level_replica: warningdisable_tqdm: Falseproject: huggingfacetrackio_space_id: Nonetrackio_bucket_id: Nonetrackio_static_space_id: Noneper_device_eval_batch_size: 8prediction_loss_only: Trueeval_on_start: Falseeval_do_concat_batches: Trueeval_use_gather_object: Falseeval_accumulation_steps: Noneinclude_for_metrics: []batch_eval_metrics: Falsesave_only_model: Falsesave_on_each_node: Falseenable_jit_checkpoint: Falsepush_to_hub: Falsehub_private_repo: Nonehub_model_id: Nonehub_strategy: every_savehub_always_push: Falsehub_revision: Noneload_best_model_at_end: Trueignore_data_skip: Falserestore_callback_states_from_checkpoint: Falsefull_determinism: Falseseed: 42data_seed: Noneuse_cpu: Falseaccelerator_config: {'split_batches': False, 'dispatch_batches': None, 'even_batches': True, 'use_seedable_sampler': True, 'non_blocking': False, 'gradient_accumulation_kwargs': None}parallelism_config: Nonedataloader_drop_last: Falsedataloader_num_workers: 0dataloader_pin_memory: Truedataloader_persistent_workers: Falsedataloader_prefetch_factor: Nonedataloader_multiprocessing_context: Nonedataloader_in_order: Trueremove_unused_columns: Truelabel_names: Nonetrain_sampling_strategy: randomlength_column_name: lengthddp_find_unused_parameters: Noneddp_bucket_cap_mb: Noneddp_broadcast_buffers: Falseddp_static_graph: Noneddp_backend: Noneddp_timeout: 1800fsdp: Nonefsdp_config: Nonedeepspeed: Nonedebug: []skip_memory_metrics: Truedo_predict: Falseresume_from_checkpoint: Nonelocal_rank: -1prompts: Nonebatch_sampler: no_duplicatesmulti_dataset_batch_sampler: proportionalrouter_mapping: {}learning_rate_mapping: {}warmup_ratio: None
Training Logs
| Epoch | Step | Training Loss | heldout_go_cosine_ndcg@10 |
|---|---|---|---|
| -1 | -1 | - | 0.0285 |
| 0.3390 | 20 | 7.3693 | - |
| 0.6780 | 40 | 6.2354 | - |
| 1.0 | 59 | - | 0.1549 |
| 1.0169 | 60 | 5.3803 | - |
| 1.3559 | 80 | 5.3655 | - |
| 1.6949 | 100 | 5.1811 | - |
| 2.0 | 118 | - | 0.1886 |
| 2.0339 | 120 | 4.7409 | - |
| 2.3729 | 140 | 4.7735 | - |
| 2.7119 | 160 | 4.7011 | - |
| 3.0 | 177 | - | 0.2067 |
| 3.0508 | 180 | 4.4012 | - |
| 3.3898 | 200 | 4.4621 | - |
| 3.7288 | 220 | 4.4667 | - |
| 4.0 | 236 | - | 0.2075 |
| 4.0678 | 240 | 4.1357 | - |
| 4.4068 | 260 | 4.2363 | - |
| 4.7458 | 280 | 4.2808 | - |
| 5.0 | 295 | - | 0.2086 |
| 5.0847 | 300 | 3.9821 | - |
| 5.4237 | 320 | 4.1327 | - |
| 5.7627 | 340 | 4.1406 | - |
| 6.0 | 354 | - | 0.2106 |
| 6.1017 | 360 | 3.8678 | - |
| 6.4407 | 380 | 4.0266 | - |
| 6.7797 | 400 | 4.0569 | - |
| 7.0 | 413 | - | 0.2114 |
| 7.1186 | 420 | 3.7898 | - |
| 7.4576 | 440 | 3.9818 | - |
| 7.7966 | 460 | 4.0110 | - |
| 8.0 | 472 | - | 0.2131 |
| 8.1356 | 480 | 3.7468 | - |
| 8.4746 | 500 | 3.9718 | - |
| 8.8136 | 520 | 3.9436 | - |
| 9.0 | 531 | - | 0.2130 |
| 9.1525 | 540 | 3.7193 | - |
| 9.4915 | 560 | 3.9474 | - |
| 9.8305 | 580 | 3.9401 | - |
| 10.0 | 590 | - | 0.2134 |
| -1 | -1 | - | 0.2135 |
- The bold row denotes the saved checkpoint.
Training Time
- Training: 41.7 minutes
- Evaluation: 1.3 minutes
- Total: 43.0 minutes
Framework Versions
- Python: 3.13.15
- Sentence Transformers: 6.1.0
- Transformers: 5.18.0
- PyTorch: 2.11.0+cu130
- Accelerate: 1.15.0
- Datasets: 4.8.5
- Tokenizers: 0.23.2
Additional Resources
- Training and Finetuning Embedding Models with Sentence Transformers: the end-to-end guide for training or finetuning Sentence Transformer models.
- Introduction to Matryoshka Embedding Models: variable-size embeddings that can be truncated with minimal quality loss.
- Binary and Scalar Embedding Quantization for Significantly Faster & Cheaper Retrieval: post-training compression of embedding vectors.
- Multimodal Embedding & Reranker Models with Sentence Transformers: use text, image, audio, and video models through the same API.
- Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers: train multimodal embedding models, with a Visual Document Retrieval walkthrough.
Citation
BibTeX
Sentence Transformers
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
CachedMultipleNegativesRankingLoss
@misc{gao2021scaling,
title={Scaling Deep Contrastive Learning Batch Size under Memory Limited Setup},
author={Luyu Gao and Yunyi Zhang and Jiawei Han and Jamie Callan},
year={2021},
eprint={2101.06983},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
MultipleNegativesRankingLoss
@misc{oord2019representationlearningcontrastivepredictive,
title={Representation Learning with Contrastive Predictive Coding},
author={Aaron van den Oord and Yazhe Li and Oriol Vinyals},
year={2019},
eprint={1807.03748},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/1807.03748},
}
- Downloads last month
- 27
Model tree for yoyo458/BioBERT_GOUniProt_FineTuned
Base model
dmis-lab/biobert-base-cased-v1.2Papers for yoyo458/BioBERT_GOUniProt_FineTuned
Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Representation Learning with Contrastive Predictive Coding
Evaluation results
- Cosine Accuracy@1 on heldout goself-reported0.226
- Cosine Accuracy@3 on heldout goself-reported0.394
- Cosine Accuracy@5 on heldout goself-reported0.483
- Cosine Accuracy@10 on heldout goself-reported0.593
- Cosine Precision@1 on heldout goself-reported0.226
- Cosine Precision@3 on heldout goself-reported0.207
- Cosine Precision@5 on heldout goself-reported0.195
- Cosine Precision@10 on heldout goself-reported0.165