TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Paper • 2604.12012 • Published Apr 13 • 16
Gemini Embedding 2: A Native Multimodal Embedding Model from Gemini Paper • 2605.27295 • Published May 26 • 21
Theano: A Python framework for fast computation of mathematical expressions Paper • 1605.02688 • Published May 9, 2016 • 2
Gemma 2: Improving Open Language Models at a Practical Size Paper • 2408.00118 • Published Jul 31, 2024 • 80
EmbeddingGemma: Powerful and Lightweight Text Representations Paper • 2509.20354 • Published Sep 24, 2025 • 51
EmbeddingGemma: Powerful and Lightweight Text Representations Paper • 2509.20354 • Published Sep 24, 2025 • 51
Gemini: A Family of Highly Capable Multimodal Models Paper • 2312.11805 • Published Dec 19, 2023 • 51
Gemini Embedding: Generalizable Embeddings from Gemini Paper • 2503.07891 • Published Mar 10, 2025 • 49
Improved Long-Form Speech Recognition by Jointly Modeling the Primary and Non-primary Speakers Paper • 2312.11123 • Published Dec 18, 2023
CVSS Corpus and Massively Multilingual Speech-to-Speech Translation Paper • 2201.03713 • Published Jan 11, 2022
SpeakerStew: Scaling to Many Languages with a Triaged Multilingual Text-Dependent and Text-Independent Speaker Verification System Paper • 2104.02125 • Published Apr 5, 2021
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech Paper • 2202.12163 • Published Feb 24, 2022
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models Paper • 2401.03506 • Published Jan 7, 2024 • 17
ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech Paper • 1911.01601 • Published Nov 5, 2019
Transfer Learning from Speaker Verification to Multispeaker Text-To-Speech Synthesis Paper • 1806.04558 • Published Jun 12, 2018