timm Qwen3 ViT Encoders
Collection
39 items • Updated • 4
How to use timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b with timm:
import timm
model = timm.create_model("hf_hub:timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b", pretrained=True)How to use timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("image-feature-extraction", model="timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b") # Load model directly
from transformers import AutoModel
model = AutoModel.from_pretrained("timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b", device_map="auto")A Qwen ViT image feature model extracted from Qwen-Drive-1.0-4B. This is the native vision encoder, including the spatial merger and projection to the source LLM width.
NOTE: This checkpoint is a native timm remap of the original vision weights, with no additional training. It contains no language-model weights or trained image-classification head.
mean=(0.5, 0.5, 0.5) and std=(0.5, 0.5, 0.5). Rectangular inputs are supported. Each image dimension must be divisible by 16; any variant using the 2×2 merger requires divisibility by 32.forward_features() returns raw, unnormalized NHWC backbone features. The _enc variant returns spatially merged tokens from forward(); the classifier variant returns pooled image embeddings until a classification head is added.import torch
import timm
from PIL import Image
model = timm.create_model('hf-hub:timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b', pretrained=True).eval()
data_config = timm.data.resolve_model_data_config(model)
transform = timm.data.create_transform(**data_config, is_training=False)
image = Image.open('image.jpg').convert('RGB')
x = transform(image).unsqueeze(0)
with torch.inference_mode():
output = model(x) # (1, 576, 2560): projected spatial tokens
features = model.forward_features(x) # (1, 48, 48, 1024): raw backbone features (NHWC)
with torch.inference_mode():
maps = model.forward_intermediates(
x, indices=3, output_fmt='NCHW', intermediates_only=True,
)
for feature_map in maps:
print(feature_map.shape) # (1, 1024, 48, 48)
@misc{zhou2026qwendrive10initialstepvisionlanguage,
title={Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving},
author={Xin Zhou and Zongchuang Zhao and Zhibo Yang and Mingsheng Li and Humen Zhong and Shuai Bai and Du Chu and Ruizhe Chen and Zhaohai Li and Jun Tang and Qiuyue Wang and Mingkun Yang and Jiazhao Zhang and Dayiheng Liu and Dingkang Liang and Xiang Bai},
year={2026},
eprint={2609.00111},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.00111}
}
@misc{rw2019timm,
author = {Ross Wightman},
title = {PyTorch Image Models},
year = {2019},
publisher = {GitHub},
journal = {GitHub repository},
doi = {10.5281/zenodo.4414861},
howpublished = {\url{https://github.com/huggingface/pytorch-image-models}}
}