VidEoMT-S on YouTube-VIS 2019

This repository contains the Hugging Face Transformers conversion of the official VidEoMT checkpoint yt_2019_vit_small_52.8.pth from tue-mps/VidEoMT.

Model details

Reported metrics

Metric Value
AP 52.8
AR@10 62.2
FPS 294

The metrics above are the numbers reported by the authors in the official model zoo.

Usage

from transformers import AutoModelForUniversalSegmentation, AutoVideoProcessor

model_id = "tue-mps/videomt-dinov2-small-ytvis2019"
processor = AutoVideoProcessor.from_pretrained(model_id)
model = AutoModelForUniversalSegmentation.from_pretrained(model_id)

Use processor.post_process_instance_segmentation, processor.post_process_panoptic_segmentation, or processor.post_process_semantic_segmentation depending on the target task.

Downloads last month
12,109
Safetensors
Model size
24.1M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Spaces using tue-mps/videomt-dinov2-small-ytvis2019 2

Paper for tue-mps/videomt-dinov2-small-ytvis2019