yanggangu commited on
Commit
0da4c22
路
verified 路
1 Parent(s): d0f44de

Release Table 2 vision expert

Browse files
Files changed (4) hide show
  1. LICENSE +22 -0
  2. README.md +16 -0
  3. encoder.pt +3 -0
  4. metadata.json +41 -0
LICENSE ADDED
@@ -0,0 +1,22 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ MIT License
2
+
3
+ Copyright (c) 2021 OpenAI
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
22
+
README.md ADDED
@@ -0,0 +1,16 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ base_model: openai/clip-vit-base-patch32
4
+ tags:
5
+ - smat
6
+ - model-merging
7
+ - vision
8
+ - arxiv:2609.33437
9
+ ---
10
+ # CLIP ViT base-patch32 路 SMAT 路 DTD
11
+
12
+ A SMAT expert from Table 2 of **SMAT: Simple and Efficient Merge-Aware Training** (seed 42). SMAT trains experts with model merging in mind.
13
+
14
+ [Paper](https://arxiv.org/abs/2609.33437) 路 [GitHub & usage](https://github.com/egangu/smat/blob/main/docs/HUGGINGFACE.md)
15
+
16
+ `encoder.pt` is the original FP32 vision-encoder state dictionary; load it with the SMAT code and the matching OpenAI CLIP base model. Dataset terms apply separately.
encoder.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:144dfca93eb0dfb08b5ea437e5f6d18cf6b768a5b63029ac74ede37fc6a9d95d
3
+ size 349896625
metadata.json ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "experiment": "clip8_vit",
3
+ "task": "DTD",
4
+ "seed": 42,
5
+ "steps": 4000,
6
+ "epochs": null,
7
+ "train": {
8
+ "optimizer": "adam",
9
+ "method": "smat",
10
+ "batch_size": 128,
11
+ "learning_rate": 1e-05,
12
+ "weight_decay": 0.0,
13
+ "max_steps": 4000,
14
+ "schedule": "cosine",
15
+ "warmup_steps": 0,
16
+ "num_workers": 8,
17
+ "fused": true,
18
+ "regularizer": {
19
+ "name": "none",
20
+ "lambda": 0.1
21
+ },
22
+ "smat": {
23
+ "mode": "joint",
24
+ "interval": 4
25
+ },
26
+ "backend": "fast",
27
+ "scale": {
28
+ "alpha_min": 0.1,
29
+ "scope": "global"
30
+ },
31
+ "mask": {
32
+ "probability": 0.5,
33
+ "scope": "block-linear",
34
+ "embedding_probability": 0
35
+ },
36
+ "perturb": {
37
+ "rms": 0.001
38
+ },
39
+ "log_every": 20
40
+ }
41
+ }