Model card for qwen3_vit_306m_enc.qwen_drive_1_0_4b

A Qwen ViT image feature model extracted from Qwen-Drive-1.0-4B. This is the native vision encoder, including the spatial merger and projection to the source LLM width.

NOTE: This checkpoint is a native timm remap of the original vision weights, with no additional training. It contains no language-model weights or trained image-classification head.

Model Notes

  • Image inputs repeat one frame across the original temporal patch kernel. The temporal Conv3d weights are summed into a Conv2d for this image-only implementation.
  • The backbone uses GELU-tanh MLPs, learned absolute positions and axial 2D RoPE. Absolute positions are interpolated for the input grid; RoPE is regenerated at each size.
  • The timm transforms normalize RGB pixels using mean=(0.5, 0.5, 0.5) and std=(0.5, 0.5, 0.5). Rectangular inputs are supported. Each image dimension must be divisible by 16; any variant using the 2×2 merger requires divisibility by 32.
  • forward_features() returns raw, unnormalized NHWC backbone features. The _enc variant returns spatially merged tokens from forward(); the classifier variant returns pooled image embeddings until a classification head is added.

Model Details

Model Usage

Image Features

import torch
import timm
from PIL import Image

model = timm.create_model('hf-hub:timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b', pretrained=True).eval()
data_config = timm.data.resolve_model_data_config(model)
transform = timm.data.create_transform(**data_config, is_training=False)
image = Image.open('image.jpg').convert('RGB')
x = transform(image).unsqueeze(0)

with torch.inference_mode():
    output = model(x)  # (1, 576, 2560): projected spatial tokens
    features = model.forward_features(x)  # (1, 48, 48, 1024): raw backbone features (NHWC)

Intermediate Feature Maps

with torch.inference_mode():
    maps = model.forward_intermediates(
        x, indices=3, output_fmt='NCHW', intermediates_only=True,
    )
for feature_map in maps:
    print(feature_map.shape)  # (1, 1024, 48, 48)

Citation

@misc{zhou2026qwendrive10initialstepvisionlanguage,
  title={Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving},
  author={Xin Zhou and Zongchuang Zhao and Zhibo Yang and Mingsheng Li and Humen Zhong and Shuai Bai and Du Chu and Ruizhe Chen and Zhaohai Li and Jun Tang and Qiuyue Wang and Mingkun Yang and Jiazhao Zhang and Dayiheng Liu and Dingkang Liang and Xiang Bai},
  year={2026},
  eprint={2609.00111},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.00111}
}
@misc{rw2019timm,
  author = {Ross Wightman},
  title = {PyTorch Image Models},
  year = {2019},
  publisher = {GitHub},
  journal = {GitHub repository},
  doi = {10.5281/zenodo.4414861},
  howpublished = {\url{https://github.com/huggingface/pytorch-image-models}}
}
Downloads last month
89
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(5)
this model

Collection including timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b

Paper for timm/qwen3_vit_306m_enc.qwen_drive_1_0_4b