Rename Termite references to Antfly Inference
Browse files
README.md
CHANGED
|
@@ -12,7 +12,7 @@ tags:
|
|
| 12 |
- embeddings
|
| 13 |
- feature-extraction
|
| 14 |
- antfly
|
| 15 |
-
-
|
| 16 |
pipeline_tag: feature-extraction
|
| 17 |
datasets:
|
| 18 |
- OpenSound/AudioCaps
|
|
@@ -22,7 +22,7 @@ datasets:
|
|
| 22 |
|
| 23 |
CLIPCLAP is a unified multimodal embedding model that maps **text**, **images**, and **audio** into a shared 512-dimensional vector space. It combines OpenAI's [CLIP](https://huggingface.co/openai/clip-vit-base-patch32) (text + image) with LAION's [CLAP](https://huggingface.co/laion/larger_clap_music_and_speech) (audio) through a trained linear projection.
|
| 24 |
|
| 25 |
-
Built by [antflydb](https://github.com/antflydb) for use with [
|
| 26 |
|
| 27 |
## Architecture
|
| 28 |
|
|
@@ -44,12 +44,12 @@ All three modalities produce **512-dimensional L2-normalized embeddings** that a
|
|
| 44 |
- Cross-modal retrieval (find images from audio queries, audio from text, etc.)
|
| 45 |
- Audio-visual content discovery
|
| 46 |
|
| 47 |
-
## How to Use with
|
| 48 |
|
| 49 |
```bash
|
| 50 |
# Pull and run the model
|
| 51 |
-
|
| 52 |
-
|
| 53 |
|
| 54 |
# Embed text
|
| 55 |
curl -X POST http://localhost:8082/embed \
|
|
|
|
| 12 |
- embeddings
|
| 13 |
- feature-extraction
|
| 14 |
- antfly
|
| 15 |
+
- antfly-inference
|
| 16 |
pipeline_tag: feature-extraction
|
| 17 |
datasets:
|
| 18 |
- OpenSound/AudioCaps
|
|
|
|
| 22 |
|
| 23 |
CLIPCLAP is a unified multimodal embedding model that maps **text**, **images**, and **audio** into a shared 512-dimensional vector space. It combines OpenAI's [CLIP](https://huggingface.co/openai/clip-vit-base-patch32) (text + image) with LAION's [CLAP](https://huggingface.co/laion/larger_clap_music_and_speech) (audio) through a trained linear projection.
|
| 24 |
|
| 25 |
+
Built by [antflydb](https://github.com/antflydb) for use with [Antfly Inference](https://github.com/antflydb/antfly/tree/main/zig/pkg/inference), a standalone ML inference service for embeddings, chunking, reranking, and local model serving.
|
| 26 |
|
| 27 |
## Architecture
|
| 28 |
|
|
|
|
| 44 |
- Cross-modal retrieval (find images from audio queries, audio from text, etc.)
|
| 45 |
- Audio-visual content discovery
|
| 46 |
|
| 47 |
+
## How to Use with Antfly Inference
|
| 48 |
|
| 49 |
```bash
|
| 50 |
# Pull and run the model
|
| 51 |
+
antfly inference pull antflydb/clipclap:gguf:Q4_K
|
| 52 |
+
antfly inference run
|
| 53 |
|
| 54 |
# Embed text
|
| 55 |
curl -X POST http://localhost:8082/embed \
|