Escarda-VE

Vision encoder for the Escarda family (~39.6M, 448px / patch16 / 784 tokens). Byrne-VE plus a JEPA head next to the HRM refine block - that's the Escarda trait.

  • ViT-style: RMSNorm, 2D-axial RoPE, QK-Norm, SwiGLU, HRM refine.
  • JEPA head (auxiliary): from a patch token, predict the raster-order neighbour patch's representation (stop-gradient target, 1-cosine loss). Zero inference cost.
  • Distilled from frozen DINOv2-base (50k steps, CLS + patch), then 20k steps of DINO-style self-distillation (EMA teacher, cosine prototypes, no collapse).

Escarda-VE vs Byrne-VE (DINOv2 teacher-alignment, n=1024 held-out)

Byrne-VE Escarda-VE
Params 39.34M 39.60M (+JEPA head)
CLS cosine 0.776 0.771
PATCH cosine 0.600 0.584
JEPA self-consistency - 0.040

Escarda trades ~1-3% teacher-alignment (JEPA pulls a bit of capacity off pure mimicry) for a self-supervised spatial neighbour-prediction signal that costs nothing at inference. Same size class as Byrne-VE.

Related

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including Quazim0t0/Escarda-VE