wwhyyyyyy commited on 20 days ago

Commit

7f42d7b

1 Parent(s): b381465

Fix GeneralVLA asset repository layout

Browse files

Files changed (37) hide show

README.md +10 -134
added_tokens.json +0 -6
config.json +0 -171
generation_config.json +0 -7
merges.txt +0 -0
preprocessor_config.json +0 -19
pytorch_model-00001-of-00002.bin +0 -3
pytorch_model-00002-of-00002.bin +0 -3
pytorch_model.bin +0 -3
pytorch_model.bin.index.json +0 -930
special_tokens_map.json +0 -1
tokenizer.json +0 -0
tokenizer.model +0 -3
tokenizer_config.json +0 -34
vocab.json +0 -0
zzzmmz/SegAgent-Model/.mdl +0 -0
zzzmmz/SegAgent-Model/.msc +0 -0
zzzmmz/SegAgent-Model/.mv +0 -1
zzzmmz/SegAgent-Model/README.md +0 -50
zzzmmz/SegAgent-Model/config.json +0 -49
zzzmmz/SegAgent-Model/configuration.json +0 -1
zzzmmz/SegAgent-Model/configuration_qwen.py +0 -65
zzzmmz/SegAgent-Model/generation_config.json +0 -11
zzzmmz/SegAgent-Model/model-00001-of-00004.safetensors +0 -3
zzzmmz/SegAgent-Model/model-00002-of-00004.safetensors +0 -3
zzzmmz/SegAgent-Model/model-00003-of-00004.safetensors +0 -3
zzzmmz/SegAgent-Model/model-00004-of-00004.safetensors +0 -3
zzzmmz/SegAgent-Model/model.safetensors.index.json +0 -860
zzzmmz/SegAgent-Model/modeling_qwen.py +0 -1172
zzzmmz/SegAgent-Model/qwen.tiktoken +0 -0
zzzmmz/SegAgent-Model/qwen_generation_utils.py +0 -420
zzzmmz/SegAgent-Model/special_tokens_map.json +0 -3
zzzmmz/SegAgent-Model/tokenization_qwen.py +0 -598
zzzmmz/SegAgent-Model/tokenizer_config.json +0 -14
zzzmmz/SegAgent-Model/trainer_state.json +0 -0
zzzmmz/SegAgent-Model/visual.py +0 -472
zzzmmz/SegAgent-Model/zero_to_fp32.py +0 -587

README.md CHANGED Viewed

@@ -1,145 +1,21 @@
----
-tags:
-- vision
-widget:
-- src: https://huggingface.co/datasets/mishig/sample_images/resolve/main/cat-dog-music.png
-  candidate_labels: playing music, playing sports
-  example_title: Cat & Dog
----
-# Model Card: CLIP
-Disclaimer: The model card is taken and modified from the official CLIP repository, it can be found [here](https://github.com/openai/CLIP/blob/main/model-card.md).
-## Model Details
-The CLIP model was developed by researchers at OpenAI to learn about what contributes to robustness in computer vision tasks. The model was also developed to test the ability of models to generalize to arbitrary image classification tasks in a zero-shot manner. It was not developed for general model deployment - to deploy models like CLIP, researchers will first need to carefully study their capabilities in relation to the specific context they’re being deployed within.
-### Model Date
-January 2021
-### Model Type
-The base model uses a ViT-L/14 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained to maximize the similarity of (image, text) pairs via a contrastive loss.
-The original implementation had two variants: one using a ResNet image encoder and the other using a Vision Transformer. This repository has the variant with the Vision Transformer.
-### Documents
-- [Blog Post](https://openai.com/blog/clip/)
-- [CLIP Paper](https://arxiv.org/abs/2103.00020)
-### Use with Transformers
-```python
-from PIL import Image
-import requests
-from transformers import CLIPProcessor, CLIPModel
-model = CLIPModel.from_pretrained("openai/clip-vit-large-patch14")
-processor = CLIPProcessor.from_pretrained("openai/clip-vit-large-patch14")
-url = "http://images.cocodataset.org/val2017/000000039769.jpg"
-image = Image.open(requests.get(url, stream=True).raw)
-inputs = processor(text=["a photo of a cat", "a photo of a dog"], images=image, return_tensors="pt", padding=True)
-outputs = model(**inputs)
-logits_per_image = outputs.logits_per_image # this is the image-text similarity score
-probs = logits_per_image.softmax(dim=1) # we can take the softmax to get the label probabilities
-```
-## Model Use
-### Intended Use
-The model is intended as a research output for research communities. We hope that this model will enable researchers to better understand and explore zero-shot, arbitrary image classification. We also hope it can be used for interdisciplinary studies of the potential impact of such models - the CLIP paper includes a discussion of potential downstream impacts to provide an example for this sort of analysis.
-#### Primary intended uses
-The primary intended users of these models are AI researchers.
-We primarily imagine the model will be used by researchers to better understand robustness, generalization, and other capabilities, biases, and constraints of computer vision models.
-### Out-of-Scope Use Cases
-**Any** deployed use case of the model - whether commercial or not - is currently out of scope. Non-deployed use cases such as image search in a constrained environment, are also not recommended unless there is thorough in-domain testing of the model with a specific, fixed class taxonomy. This is because our safety assessment demonstrated a high need for task specific testing especially given the variability of CLIP’s performance with different class taxonomies. This makes untested and unconstrained deployment of the model in any use case currently potentially harmful.
-Certain use cases which would fall under the domain of surveillance and facial recognition are always out-of-scope regardless of performance of the model. This is because the use of artificial intelligence for tasks such as these can be premature currently given the lack of testing norms and checks to ensure its fair use.
-Since the model has not been purposefully trained in or evaluated on any languages other than English, its use should be limited to English language use cases.
-## Data
-The model was trained on publicly available image-caption data. This was done through a combination of crawling a handful of websites and using commonly-used pre-existing image datasets such as [YFCC100M](http://projects.dfki.uni-kl.de/yfcc100m/). A large portion of the data comes from our crawling of the internet. This means that the data is more representative of people and societies most connected to the internet which tend to skew towards more developed nations, and younger, male users.
-### Data Mission Statement
-Our goal with building this dataset was to test out robustness and generalizability in computer vision tasks. As a result, the focus was on gathering large quantities of data from different publicly-available internet data sources. The data was gathered in a mostly non-interventionist manner. However, we only crawled websites that had policies against excessively violent and adult images and allowed us to filter out such content. We do not intend for this dataset to be used as the basis for any commercial or deployed model and will not be releasing the dataset.
-## Performance and Limitations
-### Performance
-We have evaluated the performance of CLIP on a wide range of benchmarks across a variety of computer vision datasets such as OCR to texture recognition to fine-grained classification. The paper describes model performance on the following datasets:
-- Food101
-- CIFAR10
-- CIFAR100
-- Birdsnap
-- SUN397
-- Stanford Cars
-- FGVC Aircraft
-- VOC2007
-- DTD
-- Oxford-IIIT Pet dataset
-- Caltech101
-- Flowers102
-- MNIST
-- SVHN
-- IIIT5K
-- Hateful Memes
-- SST-2
-- UCF101
-- Kinetics700
-- Country211
-- CLEVR Counting
-- KITTI Distance
-- STL-10
-- RareAct
-- Flickr30
-- MSCOCO
-- ImageNet
-- ImageNet-A
-- ImageNet-R
-- ImageNet Sketch
-- ObjectNet (ImageNet Overlap)
-- Youtube-BB
-- ImageNet-Vid
-## Limitations
-CLIP and our analysis of it have a number of limitations. CLIP currently struggles with respect to certain tasks such as fine grained classification and counting objects. CLIP also poses issues with regards to fairness and bias which we discuss in the paper and briefly in the next section. Additionally, our approach to testing CLIP also has an important limitation- in many cases we have used linear probes to evaluate the performance of CLIP and there is evidence suggesting that linear probes can underestimate model performance.
-### Bias and Fairness
-We find that the performance of CLIP - and the specific biases it exhibits - can depend significantly on class design and the choices one makes for categories to include and exclude. We tested the risk of certain kinds of denigration with CLIP by classifying images of people from [Fairface](https://arxiv.org/abs/1908.04913) into crime-related and non-human animal categories. We found significant disparities with respect to race and gender. Additionally, we found that these disparities could shift based on how the classes were constructed. (Details captured in the Broader Impacts Section in the paper).
-We also tested the performance of CLIP on gender, race and age classification using the Fairface dataset (We default to using race categories as they are constructed in the Fairface dataset.) in order to assess quality of performance across different demographics. We found accuracy >96% across all races for gender classification with ‘Middle Eastern’ having the highest accuracy (98.4%) and ‘White’ having the lowest (96.5%). Additionally, CLIP averaged ~93% for racial classification and ~63% for age classification. Our use of evaluations to test for gender, race and age classification as well as denigration harms is simply to evaluate performance of the model across people and surface potential risks and not to demonstrate an endorsement/enthusiasm for such tasks.
-## Feedback
-### Where to send questions or comments about the model
-Please use [this Google Form](https://forms.gle/Uv7afRH5dvY34ZEs9)

+   # GeneralVLA Model Assets
+   This repository stores pretrained model assets and checkpoints for GeneralVLA.
+   ## Layout
+   - `LISA-7B-v1-explanatory/`
+  - `clip-vit-large-patch14/`
+  - `segagent/zzzmmz/SegAgent-Model/`
+ - `sam_vit_h_4b8939.pth`
+- `checkpoints/v1/checkpoint-rs.tar`

added_tokens.json DELETED Viewed

@@ -1,6 +0,0 @@
-{
-  "<im_end>": 32002,
-  "<im_patch>": 32000,
-  "<im_start>": 32001,
-  "[SEG]": 32003
-}

config.json DELETED Viewed

@@ -1,171 +0,0 @@
-{
-  "_name_or_path": "clip-vit-large-patch14/",
-  "architectures": [
-    "CLIPModel"
-  ],
-  "initializer_factor": 1.0,
-  "logit_scale_init_value": 2.6592,
-  "model_type": "clip",
-  "projection_dim": 768,
-  "text_config": {
-    "_name_or_path": "",
-    "add_cross_attention": false,
-    "architectures": null,
-    "attention_dropout": 0.0,
-    "bad_words_ids": null,
-    "bos_token_id": 0,
-    "chunk_size_feed_forward": 0,
-    "cross_attention_hidden_size": null,
-    "decoder_start_token_id": null,
-    "diversity_penalty": 0.0,
-    "do_sample": false,
-    "dropout": 0.0,
-    "early_stopping": false,
-    "encoder_no_repeat_ngram_size": 0,
-    "eos_token_id": 2,
-    "finetuning_task": null,
-    "forced_bos_token_id": null,
-    "forced_eos_token_id": null,
-    "hidden_act": "quick_gelu",
-    "hidden_size": 768,
-    "id2label": {
-      "0": "LABEL_0",
-      "1": "LABEL_1"
-    },
-    "initializer_factor": 1.0,
-    "initializer_range": 0.02,
-    "intermediate_size": 3072,
-    "is_decoder": false,
-    "is_encoder_decoder": false,
-    "label2id": {
-      "LABEL_0": 0,
-      "LABEL_1": 1
-    },
-    "layer_norm_eps": 1e-05,
-    "length_penalty": 1.0,
-    "max_length": 20,
-    "max_position_embeddings": 77,
-    "min_length": 0,
-    "model_type": "clip_text_model",
-    "no_repeat_ngram_size": 0,
-    "num_attention_heads": 12,
-    "num_beam_groups": 1,
-    "num_beams": 1,
-    "num_hidden_layers": 12,
-    "num_return_sequences": 1,
-    "output_attentions": false,
-    "output_hidden_states": false,
-    "output_scores": false,
-    "pad_token_id": 1,
-    "prefix": null,
-    "problem_type": null,
-    "projection_dim" : 768,
-    "pruned_heads": {},
-    "remove_invalid_values": false,
-    "repetition_penalty": 1.0,
-    "return_dict": true,
-    "return_dict_in_generate": false,
-    "sep_token_id": null,
-    "task_specific_params": null,
-    "temperature": 1.0,
-    "tie_encoder_decoder": false,
-    "tie_word_embeddings": true,
-    "tokenizer_class": null,
-    "top_k": 50,
-    "top_p": 1.0,
-    "torch_dtype": null,
-    "torchscript": false,
-    "transformers_version": "4.16.0.dev0",
-    "use_bfloat16": false,
-    "vocab_size": 49408
-  },
-  "text_config_dict": {
-    "hidden_size": 768,
-    "intermediate_size": 3072,
-    "num_attention_heads": 12,
-    "num_hidden_layers": 12,
-    "projection_dim": 768
-  },
-  "torch_dtype": "float32",
-  "transformers_version": null,
-  "vision_config": {
-    "_name_or_path": "",
-    "add_cross_attention": false,
-    "architectures": null,
-    "attention_dropout": 0.0,
-    "bad_words_ids": null,
-    "bos_token_id": null,
-    "chunk_size_feed_forward": 0,
-    "cross_attention_hidden_size": null,
-    "decoder_start_token_id": null,
-    "diversity_penalty": 0.0,
-    "do_sample": false,
-    "dropout": 0.0,
-    "early_stopping": false,
-    "encoder_no_repeat_ngram_size": 0,
-    "eos_token_id": null,
-    "finetuning_task": null,
-    "forced_bos_token_id": null,
-    "forced_eos_token_id": null,
-    "hidden_act": "quick_gelu",
-    "hidden_size": 1024,
-    "id2label": {
-      "0": "LABEL_0",
-      "1": "LABEL_1"
-    },
-    "image_size": 224,
-    "initializer_factor": 1.0,
-    "initializer_range": 0.02,
-    "intermediate_size": 4096,
-    "is_decoder": false,
-    "is_encoder_decoder": false,
-    "label2id": {
-      "LABEL_0": 0,
-      "LABEL_1": 1
-    },
-    "layer_norm_eps": 1e-05,
-    "length_penalty": 1.0,
-    "max_length": 20,
-    "min_length": 0,
-    "model_type": "clip_vision_model",
-    "no_repeat_ngram_size": 0,
-    "num_attention_heads": 16,
-    "num_beam_groups": 1,
-    "num_beams": 1,
-    "num_hidden_layers": 24,
-    "num_return_sequences": 1,
-    "output_attentions": false,
-    "output_hidden_states": false,
-    "output_scores": false,
-    "pad_token_id": null,
-    "patch_size": 14,
-    "prefix": null,
-    "problem_type": null,
-    "projection_dim" : 768,
-    "pruned_heads": {},
-    "remove_invalid_values": false,
-    "repetition_penalty": 1.0,
-    "return_dict": true,
-    "return_dict_in_generate": false,
-    "sep_token_id": null,
-    "task_specific_params": null,
-    "temperature": 1.0,
-    "tie_encoder_decoder": false,
-    "tie_word_embeddings": true,
-    "tokenizer_class": null,
-    "top_k": 50,
-    "top_p": 1.0,
-    "torch_dtype": null,
-    "torchscript": false,
-    "transformers_version": "4.16.0.dev0",
-    "use_bfloat16": false
-  },
-  "vision_config_dict": {
-    "hidden_size": 1024,
-    "intermediate_size": 4096,
-    "num_attention_heads": 16,
-    "num_hidden_layers": 24,
-    "patch_size": 14,
-    "projection_dim": 768
-  }
-}

generation_config.json DELETED Viewed

@@ -1,7 +0,0 @@
-{
-  "_from_model_config": true,
-  "bos_token_id": 0,
-  "eos_token_id": 1,
-  "pad_token_id": 0,
-  "transformers_version": "4.31.0"
-}

merges.txt DELETED Viewed

The diff for this file is too large to render. See raw diff

preprocessor_config.json DELETED Viewed

@@ -1,19 +0,0 @@
-{
-  "crop_size": 224,
-  "do_center_crop": true,
-  "do_normalize": true,
-  "do_resize": true,
-  "feature_extractor_type": "CLIPFeatureExtractor",
-  "image_mean": [
-    0.48145466,
-    0.4578275,
-    0.40821073
-  ],
-  "image_std": [
-    0.26862954,
-    0.26130258,
-    0.27577711
-  ],
-  "resample": 3,
-  "size": 224
-}

pytorch_model-00001-of-00002.bin DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:38ac5624a8f7b982deb37ca66c6713047b08c2c7cf828241a7e2571d7271144a
-size 9976667326

pytorch_model-00002-of-00002.bin DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:8db6d703169c626e2f83b133ac5788adfb35a05024d200e1cb687d7780c1a75d
-size 6144646041

pytorch_model.bin DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:f1a17cdbe0f36fec524f5cafb1c261ea3bbbc13e346e0f74fc9eb0460dedd0d3
-size 1710671599

pytorch_model.bin.index.json DELETED Viewed

@@ -1,930 +0,0 @@
-{
-  "metadata": {
-    "total_size": 16120985792
-  },
-  "weight_map": {
-    "lm_head.weight": "pytorch_model-00002-of-00002.bin",
-    "model.embed_tokens.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.0.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.1.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.10.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.11.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.12.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.13.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.14.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.15.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.16.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.17.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.18.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.19.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.2.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.20.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.21.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.22.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.23.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.24.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.24.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.25.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.26.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.27.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.28.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.29.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.3.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.3.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.30.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.30.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.input_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.mlp.down_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.mlp.gate_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.mlp.up_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.post_attention_layernorm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.self_attn.o_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.self_attn.rotary_emb.inv_freq": "pytorch_model-00002-of-00002.bin",
-    "model.layers.31.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.layers.4.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.4.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.5.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.6.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.7.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.8.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.input_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.mlp.down_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.mlp.gate_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.mlp.up_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.post_attention_layernorm.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.self_attn.k_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.self_attn.o_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.self_attn.q_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.self_attn.rotary_emb.inv_freq": "pytorch_model-00001-of-00002.bin",
-    "model.layers.9.self_attn.v_proj.weight": "pytorch_model-00001-of-00002.bin",
-    "model.mm_projector.bias": "pytorch_model-00002-of-00002.bin",
-    "model.mm_projector.weight": "pytorch_model-00002-of-00002.bin",
-    "model.norm.weight": "pytorch_model-00002-of-00002.bin",
-    "model.text_hidden_fcs.0.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.text_hidden_fcs.0.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.text_hidden_fcs.0.2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.text_hidden_fcs.0.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.0.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.1.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.10.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.11.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.12.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.13.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.14.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.15.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.16.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.17.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.18.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.19.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.2.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.20.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.21.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.22.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.23.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.24.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.25.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.26.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.27.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.28.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.29.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.3.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.30.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.31.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.4.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.5.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.6.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.7.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.8.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.attn.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.attn.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.attn.qkv.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.attn.qkv.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.attn.rel_pos_h": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.attn.rel_pos_w": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.blocks.9.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.neck.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.neck.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.neck.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.neck.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.neck.3.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.neck.3.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.patch_embed.proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.patch_embed.proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.image_encoder.pos_embed": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_prediction_head.layers.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_prediction_head.layers.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_prediction_head.layers.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_prediction_head.layers.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_prediction_head.layers.2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_prediction_head.layers.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.iou_token.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.mask_tokens.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.0.layers.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.1.layers.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.2.layers.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_hypernetworks_mlps.3.layers.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_upscaling.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_upscaling.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_upscaling.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_upscaling.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_upscaling.3.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.output_upscaling.3.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.final_attn_token_to_image.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_image_to_token.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.cross_attn_token_to_image.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm3.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm3.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm4.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.norm4.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.0.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_image_to_token.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.cross_attn_token_to_image.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.mlp.lin1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.mlp.lin2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm2.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm3.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm3.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm4.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.norm4.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.k_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.out_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.q_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.layers.1.self_attn.v_proj.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.norm_final_attn.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.mask_decoder.transformer.norm_final_attn.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.0.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.1.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.3.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.3.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.4.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.4.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.6.bias": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.mask_downscaling.6.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.no_mask_embed.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.not_a_point_embed.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.pe_layer.positional_encoding_gaussian_matrix": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.point_embeddings.0.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.point_embeddings.1.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.point_embeddings.2.weight": "pytorch_model-00002-of-00002.bin",
-    "model.visual_model.prompt_encoder.point_embeddings.3.weight": "pytorch_model-00002-of-00002.bin"
-  }
-}

special_tokens_map.json DELETED Viewed

	@@ -1 +0,0 @@
1	- {"bos_token": {"content": "<\|startoftext\|>", "single_word": false, "lstrip": false, "rstrip": false, "normalized": true}, "eos_token": {"content": "<\|endoftext\|>", "single_word": false, "lstrip": false, "rstrip": false, "normalized": true}, "unk_token": {"content": "<\|endoftext\|>", "single_word": false, "lstrip": false, "rstrip": false, "normalized": true}, "pad_token": "<\|endoftext\|>"}

tokenizer.json DELETED Viewed

The diff for this file is too large to render. See raw diff

tokenizer.model DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:9e556afd44213b6bd1be2b850ebbbd98f5481437a8021afaf58ee7fb1818d347
-size 499723

tokenizer_config.json DELETED Viewed

@@ -1,34 +0,0 @@
-{
-    "unk_token": {
-        "content": "<|endoftext|>",
-        "single_word": false,
-        "lstrip": false,
-        "rstrip": false,
-        "normalized": true,
-        "__type": "AddedToken"
-    },
-    "bos_token": {
-        "content": "<|startoftext|>",
-        "single_word": false,
-        "lstrip": false,
-        "rstrip": false,
-        "normalized": true,
-        "__type": "AddedToken"
-    },
-    "eos_token": {
-        "content": "<|endoftext|>",
-        "single_word": false,
-        "lstrip": false,
-        "rstrip": false,
-        "normalized": true,
-        "__type": "AddedToken"
-    },
-    "pad_token": "<|endoftext|>",
-    "add_prefix_space": false,
-    "errors": "replace",
-    "do_lower_case": true,
-    "name_or_path": "openai/clip-vit-base-patch32",
-    "model_max_length": 77,
-    "special_tokens_map_file": "./special_tokens_map.json",
-    "tokenizer_class": "CLIPTokenizer"
-}

vocab.json DELETED Viewed

The diff for this file is too large to render. See raw diff

zzzmmz/SegAgent-Model/.mdl DELETED Viewed

Binary file (44 Bytes)

zzzmmz/SegAgent-Model/.msc DELETED Viewed

Binary file (1.45 kB)

zzzmmz/SegAgent-Model/.mv DELETED Viewed

	@@ -1 +0,0 @@
1	- Revision:master,CreatedAt:1754641259

zzzmmz/SegAgent-Model/README.md DELETED Viewed

@@ -1,50 +0,0 @@
----
-frameworks:
-- Pytorch
-license: Apache License 2.0
-tasks:
-- image-captioning
-#model-type:
-##such as  gpt、phi、llama、chatglm、baichuan, etc.
-#- gpt
-#domain:
-##such as  nlp、cv、audio、multi-modal, etc.
-#- nlp
-#language:
-##language code list https://help.aliyun.com/document_detail/215387.html?spm=a2c4g.11186623.0.0.9f8d7467kni6Aa
-#- cn
-#metrics:
-##such as  CIDEr、Blue、ROUGE, etc.
-#- CIDEr
-#tags:
-##various custom tags, including pretrained, fine-tuned, instruction-tuned, RL-tuned, and others
-#- pretrained
-#tools:
-##such as  vllm、fastchat、llamacpp、AdaSeq, etc.
-#- vllm
----
-### You are viewing the default Readme template as no detailed model-card was provided by the model’s contributors. You can access the model files in the "Files and versions" tab.
-#### Model files may be downloaded with ModelScope SDK or through git clone directly.
-Download with ModelScope’s Python SDK
-```bash
-#Install ModelScope
-pip install modelscope
-```
-```python
-#Download with ModelScope’s Python SDK
-from modelscope import snapshot_download
-model_dir = snapshot_download('zzzmmz/SegAgent-Model')
-```
-Download with Git clone
-```
-git clone https://www.modelscope.cn/zzzmmz/SegAgent-Model.git
-```
-<p style="color: lightgrey;">If you are a contributor to this model, we invite you to promptly update the model card content according to <a href="https://modelscope.cn/docs/ModelScope%E6%A8%A1%E5%9E%8B%E6%8E%A5%E5%85%A5%E6%B5%81%E7%A8%8B%E6%A6%82%E8%A7%88" style="color: lightgrey; text-decoration: underline;">the model contribution documentation</a>.</p>

zzzmmz/SegAgent-Model/config.json DELETED Viewed

@@ -1,49 +0,0 @@
-{
-  "_name_or_path": "/mnt/input/zhumuzhi.zmz/weight/models--Qwen--Qwen-VL-Chat",
-  "architectures": [
-    "QWenLMHeadModel"
-  ],
-  "attn_dropout_prob": 0.0,
-  "auto_map": {
-    "AutoConfig": "configuration_qwen.QWenConfig",
-    "AutoModelForCausalLM": "modeling_qwen.QWenLMHeadModel"
-  },
-  "bf16": true,
-  "emb_dropout_prob": 0.0,
-  "fp16": false,
-  "fp32": false,
-  "hidden_size": 4096,
-  "initializer_range": 0.02,
-  "intermediate_size": 22016,
-  "kv_channels": 128,
-  "layer_norm_epsilon": 1e-06,
-  "max_position_embeddings": 8192,
-  "model_type": "qwen",
-  "no_bias": true,
-  "num_attention_heads": 32,
-  "num_hidden_layers": 32,
-  "onnx_safe": null,
-  "rotary_emb_base": 10000,
-  "rotary_pct": 1.0,
-  "scale_attn_weights": true,
-  "seq_length": 2048,
-  "tie_word_embeddings": false,
-  "tokenizer_type": "QWenTokenizer",
-  "torch_dtype": "bfloat16",
-  "transformers_version": "4.37.2",
-  "use_cache": false,
-  "use_dynamic_ntk": true,
-  "use_flash_attn": false,
-  "use_logn_attn": true,
-  "visual": {
-    "heads": 16,
-    "image_size": 448,
-    "image_start_id": 151857,
-    "layers": 48,
-    "mlp_ratio": 4.9231,
-    "output_dim": 4096,
-    "patch_size": 14,
-    "width": 1664
-  },
-  "vocab_size": 151936
-}

zzzmmz/SegAgent-Model/configuration.json DELETED Viewed

	@@ -1 +0,0 @@
1	- {"framework":"Pytorch","task":"image-captioning"}

zzzmmz/SegAgent-Model/configuration_qwen.py DELETED Viewed

@@ -1,65 +0,0 @@
-# Copyright (c) Alibaba Cloud.
-#
-# This source code is licensed under the license found in the
-# LICENSE file in the root directory of this source tree.
-from transformers import PretrainedConfig
-class QWenConfig(PretrainedConfig):
-    model_type = "qwen"
-    keys_to_ignore_at_inference = ["past_key_values"]
-    def __init__(
-        self,
-        vocab_size=151936,
-        hidden_size=4096,
-        num_hidden_layers=32,
-        num_attention_heads=32,
-        emb_dropout_prob=0.0,
-        attn_dropout_prob=0.0,
-        layer_norm_epsilon=1e-6,
-        initializer_range=0.02,
-        max_position_embeddings=8192,
-        scale_attn_weights=True,
-        use_cache=True,
-        bf16=False,
-        fp16=False,
-        fp32=False,
-        kv_channels=128,
-        rotary_pct=1.0,
-        rotary_emb_base=10000,
-        use_dynamic_ntk=True,
-        use_logn_attn=True,
-        use_flash_attn="auto",
-        intermediate_size=22016,
-        no_bias=True,
-        tie_word_embeddings=False,
-        **kwargs,
-    ):
-        self.vocab_size = vocab_size
-        self.hidden_size = hidden_size
-        self.intermediate_size = intermediate_size
-        self.num_hidden_layers = num_hidden_layers
-        self.num_attention_heads = num_attention_heads
-        self.emb_dropout_prob = emb_dropout_prob
-        self.attn_dropout_prob = attn_dropout_prob
-        self.layer_norm_epsilon = layer_norm_epsilon
-        self.initializer_range = initializer_range
-        self.scale_attn_weights = scale_attn_weights
-        self.use_cache = use_cache
-        self.max_position_embeddings = max_position_embeddings
-        self.bf16 = bf16
-        self.fp16 = fp16
-        self.fp32 = fp32
-        self.kv_channels = kv_channels
-        self.rotary_pct = rotary_pct
-        self.rotary_emb_base = rotary_emb_base
-        self.use_dynamic_ntk = use_dynamic_ntk
-        self.use_logn_attn = use_logn_attn
-        self.use_flash_attn = use_flash_attn
-        self.no_bias = no_bias
-        super().__init__(
-            tie_word_embeddings=tie_word_embeddings,
-            **kwargs
-        )

zzzmmz/SegAgent-Model/generation_config.json DELETED Viewed

@@ -1,11 +0,0 @@
-{
-  "chat_format": "chatml",
-  "do_sample": true,
-  "eos_token_id": 151643,
-  "max_new_tokens": 512,
-  "max_window_size": 6144,
-  "pad_token_id": 151643,
-  "top_k": 0,
-  "top_p": 0.3,
-  "transformers_version": "4.37.2"
-}

zzzmmz/SegAgent-Model/model-00001-of-00004.safetensors DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:470ade4cefc50a771662a41077ea08f0d166c701eacb6a399652ccc383a5f386
-size 4988485656

zzzmmz/SegAgent-Model/model-00002-of-00004.safetensors DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:d2ee9a55ff2501ed625d914fa21e16ff6751058d74f96e4d408cb7fc9718f348
-size 4981246520

zzzmmz/SegAgent-Model/model-00003-of-00004.safetensors DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:d665efb246e1fe8a80f2a26e6f4af2256c2a78f94b532486a751d5ee99b70599
-size 4977360088

zzzmmz/SegAgent-Model/model-00004-of-00004.safetensors DELETED Viewed

@@ -1,3 +0,0 @@
-version https://git-lfs.github.com/spec/v1
-oid sha256:8b975b9b3b9dcbf70e093f81684cf36b3c2e3c1aacf19546e6cb838addaa3446
-size 4366885504

zzzmmz/SegAgent-Model/model.safetensors.index.json DELETED Viewed

@@ -1,860 +0,0 @@
-{
-  "metadata": {
-    "total_size": 19313870336
-  },
-  "weight_map": {
-    "lm_head.weight": "model-00004-of-00004.safetensors",
-    "transformer.h.0.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.0.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.0.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.0.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.0.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.0.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.0.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.0.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.1.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.1.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.10.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.10.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.10.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.10.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.10.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.10.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.10.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.10.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.11.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.11.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.12.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.12.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.13.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.13.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.14.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.14.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.15.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.15.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.16.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.16.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.17.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.17.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.18.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.18.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.19.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.19.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.2.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.2.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.2.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.2.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.2.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.2.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.2.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.2.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.20.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.20.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.20.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.20.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.20.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.20.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.20.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.20.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.21.attn.c_attn.bias": "model-00002-of-00004.safetensors",
-    "transformer.h.21.attn.c_attn.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.21.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.21.ln_1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.21.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.21.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.21.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.21.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.22.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.22.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.23.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.23.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.24.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.24.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.25.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.25.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.26.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.26.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.27.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.27.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.28.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.28.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.29.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.29.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.3.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.3.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.3.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.3.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.3.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.3.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.3.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.3.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.30.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.30.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.30.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.30.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.30.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.30.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.30.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.30.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.attn.c_attn.bias": "model-00003-of-00004.safetensors",
-    "transformer.h.31.attn.c_attn.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.attn.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.mlp.w1.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.31.mlp.w2.weight": "model-00003-of-00004.safetensors",
-    "transformer.h.4.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.4.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.4.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.4.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.4.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.4.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.4.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.4.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.5.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.5.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.6.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.6.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.7.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.7.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.8.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.attn.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.ln_2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.mlp.c_proj.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.mlp.w1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.8.mlp.w2.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.9.attn.c_attn.bias": "model-00001-of-00004.safetensors",
-    "transformer.h.9.attn.c_attn.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.9.attn.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.9.ln_1.weight": "model-00001-of-00004.safetensors",
-    "transformer.h.9.ln_2.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.9.mlp.c_proj.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.9.mlp.w1.weight": "model-00002-of-00004.safetensors",
-    "transformer.h.9.mlp.w2.weight": "model-00002-of-00004.safetensors",
-    "transformer.ln_f.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.attn_pool.attn.in_proj_bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.attn.in_proj_weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.kv_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.ln_kv.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.ln_kv.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.ln_q.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.ln_q.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.pos_embed": "model-00004-of-00004.safetensors",
-    "transformer.visual.attn_pool.query": "model-00004-of-00004.safetensors",
-    "transformer.visual.conv1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.ln_post.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.ln_post.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.ln_pre.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.ln_pre.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.positional_embedding": "model-00003-of-00004.safetensors",
-    "transformer.visual.proj": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.0.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.1.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.10.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.11.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.12.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.13.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.14.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.15.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.16.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.17.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.18.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.19.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.2.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.20.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.21.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.22.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.23.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.24.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.25.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.26.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.27.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.28.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.29.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.3.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.30.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.31.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.32.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.33.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.34.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.35.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.36.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.37.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.38.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.39.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.4.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.40.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.41.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.42.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.43.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.44.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.45.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.46.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.attn.in_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.attn.in_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.attn.out_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.attn.out_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.ln_1.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.ln_1.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.ln_2.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.ln_2.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.47.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.5.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.6.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.7.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.mlp.c_fc.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.mlp.c_fc.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.mlp.c_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.8.mlp.c_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.attn.in_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.attn.in_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.attn.out_proj.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.attn.out_proj.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.ln_1.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.ln_1.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.ln_2.bias": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.ln_2.weight": "model-00003-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.mlp.c_fc.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.mlp.c_fc.weight": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.mlp.c_proj.bias": "model-00004-of-00004.safetensors",
-    "transformer.visual.transformer.resblocks.9.mlp.c_proj.weight": "model-00004-of-00004.safetensors",
-    "transformer.wte.weight": "model-00001-of-00004.safetensors"
-  }
-}

zzzmmz/SegAgent-Model/modeling_qwen.py DELETED Viewed

@@ -1,1172 +0,0 @@
-# Copyright (c) Alibaba Cloud.
-#
-# This source code is licensed under the license found in the
-# LICENSE file in the root directory of this source tree.
-import importlib
-import math
-from typing import TYPE_CHECKING, Optional, Tuple, Union, Callable, List, Any, Generator
-import torch
-import torch.nn.functional as F
-import torch.utils.checkpoint
-from torch.cuda.amp import autocast
-from torch.nn import CrossEntropyLoss
-from transformers import PreTrainedTokenizer, GenerationConfig, StoppingCriteriaList
-from transformers.generation.logits_process import LogitsProcessorList
-if TYPE_CHECKING:
-    from transformers.generation.streamers import BaseStreamer
-from transformers.generation.utils import GenerateOutput
-from transformers.modeling_outputs import (
-    BaseModelOutputWithPast,
-    CausalLMOutputWithPast,
-)
-from transformers.modeling_utils import PreTrainedModel
-from transformers.utils import logging
-try:
-    from einops import rearrange
-except ImportError:
-    rearrange = None
-from torch import nn
-SUPPORT_CUDA = torch.cuda.is_available()
-SUPPORT_BF16 = SUPPORT_CUDA and torch.cuda.is_bf16_supported()
-SUPPORT_FP16 = SUPPORT_CUDA and torch.cuda.get_device_capability(0)[0] >= 7
-from .configuration_qwen import QWenConfig
-from .qwen_generation_utils import (
-    HistoryType,
-    make_context,
-    decode_tokens,
-    get_stop_words_ids,
-    StopWordsLogitsProcessor,
-)
-from .visual import VisionTransformer
-logger = logging.get_logger(__name__)
-_CHECKPOINT_FOR_DOC = "qwen"
-_CONFIG_FOR_DOC = "QWenConfig"
-QWen_PRETRAINED_MODEL_ARCHIVE_LIST = ["qwen-7b"]
-_ERROR_BAD_CHAT_FORMAT = """\
-We detect you are probably using the pretrained model (rather than chat model) for chatting, since the chat_format in generation_config is not "chatml".
-If you are directly using the model downloaded from Huggingface, please make sure you are using our "Qwen/Qwen-7B-Chat" Huggingface model (rather than "Qwen/Qwen-7B") when you call model.chat().
-我们检测到您可能在使用预训练模型（而非chat模型）进行多轮chat，因为您当前在generation_config指定的chat_format，并未设置为我们在对话中所支持的"chatml"格式。
-如果您在直接使用我们从Huggingface提供的模型，请确保您在调用model.chat()时，使用的是"Qwen/Qwen-7B-Chat"模型（而非"Qwen/Qwen-7B"预训练模型）。
-"""
-_SENTINEL = object()
-_ERROR_STREAM_IN_CHAT = """\
-Pass argument `stream` to model.chat() is buggy, deprecated, and marked for removal. Please use model.chat_stream(...) instead of model.chat(..., stream=True).
-向model.chat()传入参数stream的用法可能存在Bug，该用法已被废弃，将在未来被移除。请使用model.chat_stream(...)代替model.chat(..., stream=True)。
-"""
-apply_rotary_emb_func = None
-rms_norm = None
-def int_list_to_str(lst: List[int]) -> str:
-    """将整数列表转换回字符串"""
-    return ''.join(chr(i) for i in lst)
-# Copied from transformers.models.bart.modeling_bart._make_causal_mask
-def _make_causal_mask(
-    input_ids_shape: torch.Size, dtype: torch.dtype, device: torch.device, past_key_values_length: int = 0
-):
-    """
-    Make causal mask used for bi-directional self-attention.
-    """
-    bsz, tgt_len = input_ids_shape
-    mask = torch.full((tgt_len, tgt_len), torch.finfo(dtype).min, device=device)
-    mask_cond = torch.arange(mask.size(-1), device=device)
-    mask.masked_fill_(mask_cond < (mask_cond + 1).view(mask.size(-1), 1), 0)
-    mask = mask.to(dtype)
-    if past_key_values_length > 0:
-        mask = torch.cat([torch.zeros(tgt_len, past_key_values_length, dtype=dtype, device=device), mask], dim=-1)
-    return mask[None, None, :, :].expand(bsz, 1, tgt_len, tgt_len + past_key_values_length)
-# Copied from transformers.models.bart.modeling_bart._expand_mask
-def _expand_mask(mask: torch.Tensor, dtype: torch.dtype, tgt_len: Optional[int] = None):
-    """
-    Expands attention_mask from `[bsz, seq_len]` to `[bsz, 1, tgt_seq_len, src_seq_len]`.
-    """
-    bsz, src_len = mask.size()
-    tgt_len = tgt_len if tgt_len is not None else src_len
-    expanded_mask = mask[:, None, None, :].expand(bsz, 1, tgt_len, src_len).to(dtype)
-    inverted_mask = 1.0 - expanded_mask
-    return inverted_mask.masked_fill(inverted_mask.to(torch.bool), torch.finfo(dtype).min)
-class QWenAttention(nn.Module):
-    def __init__(self, config):
-        super().__init__()
-        self.register_buffer("masked_bias", torch.tensor(-1e4), persistent=False)
-        self.seq_length = config.seq_length
-        self.hidden_size = config.hidden_size
-        self.split_size = config.hidden_size
-        self.num_heads = config.num_attention_heads
-        self.head_dim = self.hidden_size // self.num_heads
-        self.scale_attn_weights = True
-        self.projection_size = config.kv_channels * config.num_attention_heads
-        assert self.projection_size % config.num_attention_heads == 0
-        self.hidden_size_per_attention_head = (
-            self.projection_size // config.num_attention_heads
-        )
-        self.c_attn = nn.Linear(config.hidden_size, 3 * self.projection_size)
-        self.c_proj = nn.Linear(
-            config.hidden_size, self.projection_size, bias=not config.no_bias
-        )
-        self.is_fp32 = not (config.bf16 or config.fp16)
-        self.bf16 = config.bf16
-        self.use_dynamic_ntk = config.use_dynamic_ntk
-        self.use_logn_attn = config.use_logn_attn
-        logn_list = [
-            math.log(i, self.seq_length) if i > self.seq_length else 1
-            for i in range(1, 32768)
-        ]
-        self.logn_tensor = torch.tensor(logn_list)[None, :, None, None]
-        self.attn_dropout = nn.Dropout(config.attn_dropout_prob)
-    def _attn(self, query, key, value, registered_causal_mask, attention_mask=None, head_mask=None):
-        attn_weights = torch.matmul(query, key.transpose(-1, -2))
-        if self.scale_attn_weights:
-            attn_weights = attn_weights / torch.full(
-                [],
-                value.size(-1) ** 0.5,
-                dtype=attn_weights.dtype,
-                device=attn_weights.device,
-            )
-        query_length, key_length = query.size(-2), key.size(-2)
-        # causal_mask = self.bias[
-        #     :, :, key_length - query_length : key_length, :key_length
-        # ]
-        # mask_value = torch.finfo(attn_weights.dtype).min
-        # mask_value = torch.full([], mask_value, dtype=attn_weights.dtype).to(
-        #     attn_weights.device
-        # )
-        # attn_weights = torch.where(
-        #     causal_mask, attn_weights.to(attn_weights.dtype), mask_value
-        # )
-        attn_weights = attn_weights + attention_mask
-        attn_weights = nn.functional.softmax(attn_weights, dim=-1)
-        attn_weights = attn_weights.type(value.dtype)
-        attn_weights = self.attn_dropout(attn_weights)
-        if head_mask is not None:
-            attn_weights = attn_weights * head_mask
-        attn_output = torch.matmul(attn_weights, value)
-        attn_output = attn_output.transpose(1, 2)
-        return attn_output, attn_weights
-    def _upcast_and_reordered_attn(
-        self, query, key, value, registered_causal_mask, attention_mask=None, head_mask=None
-    ):
-        bsz, num_heads, q_seq_len, dk = query.size()
-        _, _, k_seq_len, _ = key.size()
-        attn_weights = torch.empty(
-            bsz * num_heads,
-            q_seq_len,
-            k_seq_len,
-            dtype=torch.float32,
-            device=query.device,
-        )
-        scale_factor = 1.0
-        if self.scale_attn_weights:
-            scale_factor /= float(value.size(-1)) ** 0.5
-        with autocast(enabled=False):
-            q, k = query.reshape(-1, q_seq_len, dk), key.transpose(-1, -2).reshape(
-                -1, dk, k_seq_len
-            )
-            attn_weights = torch.baddbmm(
-                attn_weights, q.float(), k.float(), beta=0, alpha=scale_factor
-            )
-            attn_weights = attn_weights.reshape(bsz, num_heads, q_seq_len, k_seq_len)
-        query_length, key_length = query.size(-2), key.size(-2)
-        causal_mask = registered_causal_mask[
-            :, :, key_length - query_length : key_length, :key_length
-        ]
-        mask_value = torch.finfo(attn_weights.dtype).min
-        mask_value = torch.tensor(mask_value, dtype=attn_weights.dtype).to(
-            attn_weights.device
-        )
-        attn_weights = torch.where(causal_mask, attn_weights, mask_value)
-        if attention_mask is not None:
-            attn_weights = attn_weights + attention_mask
-        attn_weights = nn.functional.softmax(attn_weights, dim=-1)
-        if attn_weights.dtype != torch.float32:
-            raise RuntimeError(
-                "Error with upcasting, attn_weights does not have dtype torch.float32"
-            )
-        attn_weights = attn_weights.type(value.dtype)
-        attn_weights = self.attn_dropout(attn_weights)
-        if head_mask is not None:
-            attn_weights = attn_weights * head_mask
-        attn_output = torch.matmul(attn_weights, value)
-        return attn_output, attn_weights
-    def _split_heads(self, tensor, num_heads, attn_head_size):
-        new_shape = tensor.size()[:-1] + (num_heads, attn_head_size)
-        tensor = tensor.view(new_shape)
-        return tensor
-    def _merge_heads(self, tensor, num_heads, attn_head_size):
-        tensor = tensor.contiguous()
-        new_shape = tensor.size()[:-2] + (num_heads * attn_head_size,)
-        return tensor.view(new_shape)
-    def forward(
-        self,
-        hidden_states: Optional[Tuple[torch.FloatTensor]],
-        rotary_pos_emb: Optional[List[torch.Tensor]] = None,
-        registered_causal_mask: Optional[torch.Tensor] = None,
-        layer_past: Optional[Tuple[torch.Tensor]] = None,
-        attention_mask: Optional[torch.FloatTensor] = None,
-        head_mask: Optional[torch.FloatTensor] = None,
-        encoder_hidden_states: Optional[torch.Tensor] = None,
-        encoder_attention_mask: Optional[torch.FloatTensor] = None,
-        output_attentions: Optional[bool] = False,
-        use_cache: Optional[bool] = False,
-    ):
-        mixed_x_layer = self.c_attn(hidden_states)
-        query, key, value = mixed_x_layer.split(self.split_size, dim=2)
-        query = self._split_heads(query, self.num_heads, self.head_dim)
-        key = self._split_heads(key, self.num_heads, self.head_dim)
-        value = self._split_heads(value, self.num_heads, self.head_dim)
-        if rotary_pos_emb is not None:
-            cur_len = query.shape[1]
-            rotary_pos_emb = [i[:, -cur_len:, :, :] for i in rotary_pos_emb]
-            rotary_pos_emb = (rotary_pos_emb,) * 2
-            q_pos_emb, k_pos_emb = rotary_pos_emb
-            # Slice the pos emb for current inference
-            query = apply_rotary_pos_emb(query, q_pos_emb)
-            key = apply_rotary_pos_emb(key, k_pos_emb)
-        if layer_past is not None:
-            past_key, past_value = layer_past[0], layer_past[1]
-            key = torch.cat((past_key, key), dim=1)
-            value = torch.cat((past_value, value), dim=1)
-        if use_cache:
-            present = (key, value)
-        else:
-            present = None
-        if self.use_logn_attn and not self.training:
-            if self.logn_tensor.device != query.device or self.logn_tensor.dtype != query.dtype:
-                self.logn_tensor = self.logn_tensor.to(query.device).type_as(query)
-            seq_start = key.size(1) - query.size(1)
-            seq_end = key.size(1)
-            logn_tensor = self.logn_tensor[:, seq_start:seq_end, :, :]
-            query = query * logn_tensor.expand_as(query)
-        query = query.permute(0, 2, 1, 3)
-        key = key.permute(0, 2, 1, 3)
-        value = value.permute(0, 2, 1, 3)
-        attn_output, attn_weight = self._attn(
-            query, key, value, registered_causal_mask, attention_mask, head_mask
-        )
-        context_layer = self._merge_heads(
-            attn_output, self.num_heads, self.head_dim
-        )
-        attn_output = self.c_proj(context_layer)
-        outputs = (attn_output, present)
-        if output_attentions:
-            outputs += (attn_weight,)
-        return outputs
-class QWenMLP(nn.Module):
-    def __init__(self, config):
-        super().__init__()
-        self.w1 = nn.Linear(
-            config.hidden_size, config.intermediate_size // 2, bias=not config.no_bias
-        )
-        self.w2 = nn.Linear(
-            config.hidden_size, config.intermediate_size // 2, bias=not config.no_bias
-        )
-        ff_dim_in = config.intermediate_size // 2
-        self.c_proj = nn.Linear(ff_dim_in, config.hidden_size, bias=not config.no_bias)
-    def forward(self, hidden_states):
-        a1 = self.w1(hidden_states)
-        a2 = self.w2(hidden_states)
-        intermediate_parallel = a1 * F.silu(a2)
-        output = self.c_proj(intermediate_parallel)
-        return output
-class QWenBlock(nn.Module):
-    def __init__(self, config):
-        super().__init__()
-        hidden_size = config.hidden_size
-        self.bf16 = config.bf16
-        self.ln_1 = RMSNorm(
-            hidden_size,
-            eps=config.layer_norm_epsilon,
-        )
-        self.attn = QWenAttention(config)
-        self.ln_2 = RMSNorm(
-            hidden_size,
-            eps=config.layer_norm_epsilon,
-        )
-        self.mlp = QWenMLP(config)
-    def forward(
-        self,
-        hidden_states: Optional[Tuple[torch.FloatTensor]],
-        rotary_pos_emb: Optional[List[torch.Tensor]] = None,
-        registered_causal_mask: Optional[torch.Tensor] = None,
-        layer_past: Optional[Tuple[torch.Tensor]] = None,
-        attention_mask: Optional[torch.FloatTensor] = None,
-        head_mask: Optional[torch.FloatTensor] = None,
-        encoder_hidden_states: Optional[torch.Tensor] = None,
-        encoder_attention_mask: Optional[torch.FloatTensor] = None,
-        use_cache: Optional[bool] = False,
-        output_attentions: Optional[bool] = False,
-    ):
-        layernorm_output = self.ln_1(hidden_states)
-        attn_outputs = self.attn(
-            layernorm_output,
-            rotary_pos_emb,
-            registered_causal_mask=registered_causal_mask,
-            layer_past=layer_past,
-            attention_mask=attention_mask,
-            head_mask=head_mask,
-            use_cache=use_cache,
-            output_attentions=output_attentions,
-        )
-        attn_output = attn_outputs[0]
-        outputs = attn_outputs[1:]
-        residual = hidden_states
-        layernorm_input = attn_output + residual
-        layernorm_output = self.ln_2(layernorm_input)
-        residual = layernorm_input
-        mlp_output = self.mlp(layernorm_output)
-        hidden_states = residual + mlp_output
-        if use_cache:
-            outputs = (hidden_states,) + outputs
-        else:
-            outputs = (hidden_states,) + outputs[1:]
-        return outputs
-class QWenPreTrainedModel(PreTrainedModel):
-    config_class = QWenConfig
-    base_model_prefix = "transformer"
-    is_parallelizable = False
-    supports_gradient_checkpointing = True
-    _no_split_modules = ["QWenBlock"]
-    def __init__(self, *inputs, **kwargs):
-        super().__init__(*inputs, **kwargs)
-    def _init_weights(self, module):
-        """Initialize the weights."""
-        if isinstance(module, nn.Linear):
-            module.weight.data.normal_(mean=0.0, std=self.config.initializer_range)
-            if module.bias is not None:
-                module.bias.data.zero_()
-        elif isinstance(module, nn.Embedding):
-            module.weight.data.normal_(mean=0.0, std=self.config.initializer_range)
-            if module.padding_idx is not None:
-                module.weight.data[module.padding_idx].zero_()
-        elif isinstance(module, RMSNorm):
-            module.weight.data.fill_(1.0)
-        for name, p in module.named_parameters():
-            if name == "c_proj.weight":
-                p.data.normal_(
-                    mean=0.0,
-                    std=(
-                        self.config.initializer_range
-                        / math.sqrt(2 * self.config.num_hidden_layers)
-                    ),
-                )
-    def _set_gradient_checkpointing(self, module, value=False):
-        if isinstance(module, QWenModel):
-            module.gradient_checkpointing = value
-class QWenModel(QWenPreTrainedModel):
-    _keys_to_ignore_on_load_missing = ["attn.masked_bias"]
-    def __init__(self, config):
-        super().__init__(config)
-        self.vocab_size = config.vocab_size
-        self.num_hidden_layers = config.num_hidden_layers
-        self.embed_dim = config.hidden_size
-        self.gradient_checkpointing = False
-        self.use_dynamic_ntk = config.use_dynamic_ntk
-        self.seq_length = config.seq_length
-        self.wte = nn.Embedding(self.vocab_size, self.embed_dim)
-        self.drop = nn.Dropout(config.emb_dropout_prob)
-        if config.rotary_pct == 1.0:
-            self.rotary_ndims = None
-        else:
-            assert config.rotary_pct < 1
-            self.rotary_ndims = int(
-                config.kv_channels * config.rotary_pct
-            )
-        dim = (
-            self.rotary_ndims
-            if self.rotary_ndims is not None
-            else config.kv_channels
-        )
-        self.rotary_emb = RotaryEmbedding(dim, base=config.rotary_emb_base)
-        self.use_flash_attn = config.use_flash_attn
-        self.is_fp32 = not (config.bf16 or config.fp16)
-        self.registered_causal_mask = None
-        # if (
-        #     self.use_flash_attn
-        #     and flash_attn_unpadded_func is not None
-        #     and not self.is_fp32
-        # ):
-        #     self.registered_causal_mask = None
-        # else:
-        #     max_positions = config.max_position_embeddings
-        #     self.register_buffer(
-        #         "registered_causal_mask",
-        #         torch.tril(
-        #             torch.ones((max_positions, max_positions), dtype=torch.bool)
-        #         ).view(1, 1, max_positions, max_positions),
-        #         persistent=False,
-        #     )
-        self.h = nn.ModuleList(
-            [
-                QWenBlock(
-                    config
-                )
-                for i in range(config.num_hidden_layers)
-            ]
-        )
-        self.ln_f = RMSNorm(
-            self.embed_dim,
-            eps=config.layer_norm_epsilon,
-        )
-        self.visual = VisionTransformer(**config.visual)
-        self.post_init()
-    def get_input_embeddings(self):
-        return self.wte
-    def set_input_embeddings(self, new_embeddings):
-        self.wte = new_embeddings
-    # Copied from transformers.models.bart.modeling_bart.BartDecoder._prepare_decoder_attention_mask
-    def _prepare_decoder_attention_mask(self, attention_mask, input_shape, inputs_embeds, past_key_values_length):
-        # create causal mask
-        # [bsz, seq_len] -> [bsz, 1, tgt_seq_len, src_seq_len]
-        combined_attention_mask = None
-        if input_shape[-1] > 1:
-            combined_attention_mask = _make_causal_mask(
-                input_shape,
-                inputs_embeds.dtype,
-                device=inputs_embeds.device,
-                past_key_values_length=past_key_values_length,
-            )
-        if attention_mask is not None:
-            # [bsz, seq_len] -> [bsz, 1, tgt_seq_len, src_seq_len]
-            expanded_attn_mask = _expand_mask(attention_mask, inputs_embeds.dtype, tgt_len=input_shape[-1]).to(
-                inputs_embeds.device
-            )
-            combined_attention_mask = (
-                expanded_attn_mask if combined_attention_mask is None else expanded_attn_mask + combined_attention_mask
-            )
-        return combined_attention_mask
-    def forward(
-        self,
-        input_ids: Optional[torch.LongTensor] = None,
-        past_key_values: Optional[Tuple[Tuple[torch.Tensor]]] = None,
-        attention_mask: Optional[torch.FloatTensor] = None,
-        token_type_ids: Optional[torch.LongTensor] = None,
-        position_ids: Optional[torch.LongTensor] = None,
-        head_mask: Optional[torch.FloatTensor] = None,
-        inputs_embeds: Optional[torch.FloatTensor] = None,
-        encoder_hidden_states: Optional[torch.Tensor] = None,
-        encoder_attention_mask: Optional[torch.FloatTensor] = None,
-        use_cache: Optional[bool] = None,
-        output_attentions: Optional[bool] = None,
-        output_hidden_states: Optional[bool] = None,
-        return_dict: Optional[bool] = None,
-        masks_ids: Optional[torch.LongTensor] = None,
-    ):
-        if past_key_values is None and torch.any(input_ids == self.config.visual['image_start_id']):
-            bos_pos = torch.where(input_ids == self.config.visual['image_start_id'])
-            eos_pos = torch.where(input_ids == self.config.visual['image_start_id'] + 1)
-            assert (bos_pos[0] == eos_pos[0]).all()
-            img_pos = torch.stack((bos_pos[0], bos_pos[1], eos_pos[1]), dim=1)
-            images = []
-            masks = []
-            for i, a, b in img_pos:
-                image = input_ids[i][a + 1 : b - 1].tolist()
-                image = image[ : image.index(self.config.visual['image_start_id'] + 2)]
-                images.append(bytes(image).decode('utf-8'))
-                if masks_ids is not None:
-                    mask = int_list_to_str(masks_ids[i][masks_ids[i] != -1].tolist())
-                    masks.append(mask)
-                else:
-                    masks.append('')
-            images = self.visual.encode(images,masks)
-            assert images.shape[0] == len(images)
-            fake_images = None
-        elif self.training:
-            fake_images=torch.zeros(1,3,224,224).to(
-                dtype=self.visual.conv1.weight.dtype, device=self.visual.conv1.weight.device)
-            images = self.visual(fake_images)
-        else:
-            fake_images = None
-            images = None
-        output_attentions = (
-            output_attentions
-            if output_attentions is not None
-            else self.config.output_attentions
-        )
-        output_hidden_states = (
-            output_hidden_states
-            if output_hidden_states is not None
-            else self.config.output_hidden_states
-        )
-        use_cache = use_cache if use_cache is not None else self.config.use_cache
-        return_dict = (
-            return_dict if return_dict is not None else self.config.use_return_dict
-        )
-        if input_ids is not None and inputs_embeds is not None:
-            raise ValueError(
-                "You cannot specify both input_ids and inputs_embeds at the same time"
-            )
-        elif input_ids is not None:
-            input_shape = input_ids.size()
-            input_ids = input_ids.view(-1, input_shape[-1])
-            batch_size = input_ids.shape[0]
-        elif inputs_embeds is not None:
-            input_shape = inputs_embeds.size()[:-1]
-            batch_size = inputs_embeds.shape[0]
-        else:
-            raise ValueError("You have to specify either input_ids or inputs_embeds")
-        device = input_ids.device if input_ids is not None else inputs_embeds.device
-        if token_type_ids is not None:
-            token_type_ids = token_type_ids.view(-1, input_shape[-1])
-        if position_ids is not None:
-            position_ids = position_ids.view(-1, input_shape[-1])
-        if past_key_values is None:
-            past_length = 0
-            past_key_values = tuple([None] * len(self.h))
-        else:
-            past_length = past_key_values[0][0].size(-2)
-        if position_ids is None:
-            position_ids = torch.arange(
-                past_length,
-                input_shape[-1] + past_length,
-                dtype=torch.long,
-                device=device,
-            )
-            position_ids = position_ids.unsqueeze(0).view(-1, input_shape[-1])
-        encoder_attention_mask = None
-        head_mask = self.get_head_mask(head_mask, self.config.num_hidden_layers)
-        if inputs_embeds is None:
-            inputs_embeds = self.wte(input_ids)
-        if batch_size <= 0:
-            raise ValueError("batch_size has to be defined and > 0")
-        attention_mask = self._prepare_decoder_attention_mask(
-            attention_mask, input_shape, inputs_embeds, past_length
-        )
-        hidden_states = inputs_embeds
-        kv_seq_len = hidden_states.size()[1]
-        if past_key_values[0] is not None:
-            # past key values[0][0] shape: bs * seq_len * head_num * dim
-            kv_seq_len += past_key_values[0][0].shape[1]
-        if (
-            self.use_dynamic_ntk
-            and kv_seq_len == hidden_states.size()[1]
-            and not self.training
-        ):
-            context_value = math.log(kv_seq_len / self.seq_length, 2) + 1
-            ntk_alpha = 2 ** math.ceil(context_value) - 1
-            ntk_alpha = max(ntk_alpha, 1)
-        else:
-            ntk_alpha = self.rotary_emb._ntk_alpha_cached
-        rotary_pos_emb = self.rotary_emb(kv_seq_len, ntk_alpha=ntk_alpha)
-        for idx in range(len(rotary_pos_emb)):
-            rotary_pos_emb[idx] = rotary_pos_emb[idx].to(hidden_states.device)
-        hidden_states = self.drop(hidden_states).clone()
-        if fake_images is not None:
-            hidden_states = hidden_states + images.mean()*0
-        elif images is not None:
-            for idx, (i, a, b) in enumerate(img_pos):
-                hidden_states[i][a + 1 : b] = images[idx]
-        output_shape = input_shape + (hidden_states.size(-1),)
-        if self.gradient_checkpointing and self.training:
-            if use_cache:
-                logger.warning_once(
-                    "`use_cache=True` is incompatible with gradient checkpointing. Setting `use_cache=False`..."
-                )
-                use_cache = False
-        presents = () if use_cache else None
-        all_self_attentions = () if output_attentions else None
-        all_hidden_states = () if output_hidden_states else None
-        for i, (block, layer_past) in enumerate(zip(self.h, past_key_values)):
-            if output_hidden_states:
-                all_hidden_states = all_hidden_states + (hidden_states,)
-            if self.gradient_checkpointing and self.training:
-                def create_custom_forward(module):
-                    def custom_forward(*inputs):
-                        # None for past_key_value
-                        return module(*inputs, use_cache, output_attentions)
-                    return custom_forward
-                outputs = torch.utils.checkpoint.checkpoint(
-                    create_custom_forward(block),
-                    hidden_states,
-                    rotary_pos_emb,
-                    self.registered_causal_mask,
-                    None,
-                    attention_mask,
-                    head_mask[i],
-                    encoder_hidden_states,
-                    encoder_attention_mask,
-                )
-            else:
-                outputs = block(
-                    hidden_states,
-                    layer_past=layer_past,
-                    rotary_pos_emb=rotary_pos_emb,
-                    registered_causal_mask=self.registered_causal_mask,
-                    attention_mask=attention_mask,
-                    head_mask=head_mask[i],
-                    encoder_hidden_states=encoder_hidden_states,
-                    encoder_attention_mask=encoder_attention_mask,
-                    use_cache=use_cache,
-                    output_attentions=output_attentions,
-                )
-            hidden_states = outputs[0]
-            if use_cache is True:
-                presents = presents + (outputs[1],)
-            if output_attentions:
-                all_self_attentions = all_self_attentions + (outputs[2 if use_cache else 1],)
-        hidden_states = self.ln_f(hidden_states)
-        hidden_states = hidden_states.view(output_shape)
-        # Add last hidden state
-        if output_hidden_states:
-            all_hidden_states = all_hidden_states + (hidden_states,)
-        if not return_dict:
-            return tuple(
-                v for v in [hidden_states, presents, all_hidden_states] if v is not None
-            )
-        return BaseModelOutputWithPast(
-            last_hidden_state=hidden_states,
-            past_key_values=presents,
-            hidden_states=all_hidden_states,
-            attentions=all_self_attentions,
-        )
-class QWenLMHeadModel(QWenPreTrainedModel):
-    _keys_to_ignore_on_load_missing = [r"h\.\d+\.attn\.rotary_emb\.inv_freq"]
-    _keys_to_ignore_on_load_unexpected = [r"h\.\d+\.attn\.masked_bias"]
-    def __init__(self, config):
-        super().__init__(config)
-        assert (
-            config.bf16 + config.fp16 + config.fp32 <= 1
-        ), "Only one of \"bf16\", \"fp16\", \"fp32\" can be true"
-        autoset_precision = config.bf16 + config.fp16 + config.fp32 == 0
-        if autoset_precision:
-            if SUPPORT_BF16:
-                logger.warn(
-                    "The model is automatically converting to bf16 for faster inference. "
-                    "If you want to disable the automatic precision, please manually add bf16/fp16/fp32=True to \"AutoModelForCausalLM.from_pretrained\"."
-                )
-                config.bf16 = True
-            elif SUPPORT_FP16:
-                logger.warn(
-                    "The model is automatically converting to fp16 for faster inference. "
-                    "If you want to disable the automatic precision, please manually add bf16/fp16/fp32=True to \"AutoModelForCausalLM.from_pretrained\"."
-                )
-                config.fp16 = True
-            else:
-                config.fp32 = True
-        if config.bf16 and SUPPORT_CUDA and not SUPPORT_BF16:
-            logger.warn("Your device does NOT seem to support bf16, you can switch to fp16 or fp32 by by passing fp16/fp32=True in \"AutoModelForCausalLM.from_pretrained\".")
-        if config.fp16 and SUPPORT_CUDA and not SUPPORT_FP16:
-            logger.warn("Your device does NOT support faster inference with fp16, please switch to fp32 which is likely to be faster")
-        if config.fp32:
-            if SUPPORT_BF16:
-                logger.warn("Your device support faster inference by passing bf16=True in \"AutoModelForCausalLM.from_pretrained\".")
-            elif SUPPORT_FP16:
-                logger.warn("Your device support faster inference by passing fp16=True in \"AutoModelForCausalLM.from_pretrained\".")
-        self.transformer = QWenModel(config)
-        self.lm_head = nn.Linear(config.hidden_size, config.vocab_size, bias=False)
-        if config.bf16:
-            self.transformer.bfloat16()
-            self.lm_head.bfloat16()
-        if config.fp16:
-            self.transformer.half()
-            self.lm_head.half()
-        self.post_init()
-    def get_output_embeddings(self):
-        return self.lm_head
-    def set_output_embeddings(self, new_embeddings):
-        self.lm_head = new_embeddings
-    def prepare_inputs_for_generation(
-        self, input_ids, past_key_values=None, inputs_embeds=None, **kwargs
-    ):
-        token_type_ids = kwargs.get("token_type_ids", None)
-        if past_key_values:
-            input_ids = input_ids[:, -1].unsqueeze(-1)
-            if token_type_ids is not None:
-                token_type_ids = token_type_ids[:, -1].unsqueeze(-1)
-        attention_mask = kwargs.get("attention_mask", None)
-        position_ids = kwargs.get("position_ids", None)
-        if attention_mask is not None and position_ids is None:
-            position_ids = attention_mask.long().cumsum(-1) - 1
-            position_ids.masked_fill_(attention_mask == 0, 1)
-            if past_key_values:
-                position_ids = position_ids[:, -1].unsqueeze(-1)
-        else:
-            position_ids = None
-        if inputs_embeds is not None and past_key_values is None:
-            model_inputs = {"inputs_embeds": inputs_embeds}
-        else:
-            model_inputs = {"input_ids": input_ids}
-        model_inputs.update(
-            {
-                "past_key_values": past_key_values,
-                "use_cache": kwargs.get("use_cache"),
-                "position_ids": position_ids,
-                "attention_mask": attention_mask,
-                "token_type_ids": token_type_ids,
-            }
-        )
-        return model_inputs
-    def forward(
-        self,
-        input_ids: Optional[torch.LongTensor] = None,
-        past_key_values: Optional[Tuple[Tuple[torch.Tensor]]] = None,
-        attention_mask: Optional[torch.FloatTensor] = None,
-        token_type_ids: Optional[torch.LongTensor] = None,
-        position_ids: Optional[torch.LongTensor] = None,
-        head_mask: Optional[torch.FloatTensor] = None,
-        inputs_embeds: Optional[torch.FloatTensor] = None,
-        encoder_hidden_states: Optional[torch.Tensor] = None,
-        encoder_attention_mask: Optional[torch.FloatTensor] = None,
-        labels: Optional[torch.LongTensor] = None,
-        use_cache: Optional[bool] = None,
-        output_attentions: Optional[bool] = None,
-        output_hidden_states: Optional[bool] = None,
-        return_dict: Optional[bool] = None,
-        masks_ids: Optional[torch.LongTensor] = None,
-    ) -> Union[Tuple, CausalLMOutputWithPast]:
-        return_dict = (
-            return_dict if return_dict is not None else self.config.use_return_dict
-        )
-        transformer_outputs = self.transformer(
-            input_ids,
-            past_key_values=past_key_values,
-            attention_mask=attention_mask,
-            token_type_ids=token_type_ids,
-            position_ids=position_ids,
-            head_mask=head_mask,
-            inputs_embeds=inputs_embeds,
-            encoder_hidden_states=encoder_hidden_states,
-            encoder_attention_mask=encoder_attention_mask,
-            use_cache=use_cache,
-            output_attentions=output_attentions,
-            output_hidden_states=output_hidden_states,
-            return_dict=return_dict,
-            masks_ids=masks_ids,
-        )
-        hidden_states = transformer_outputs[0]
-        lm_logits = self.lm_head(hidden_states)
-        loss = None
-        if labels is not None:
-            labels = labels.to(lm_logits.device)
-            shift_logits = lm_logits[..., :-1, :].contiguous()
-            shift_labels = labels[..., 1:].contiguous()
-            loss_fct = CrossEntropyLoss()
-            loss = loss_fct(
-                shift_logits.view(-1, shift_logits.size(-1)), shift_labels.view(-1)
-            )
-        if not return_dict:
-            output = (lm_logits,) + transformer_outputs[1:]
-            return ((loss,) + output) if loss is not None else output
-        return CausalLMOutputWithPast(
-            loss=loss,
-            logits=lm_logits,
-            past_key_values=transformer_outputs.past_key_values,
-            hidden_states=transformer_outputs.hidden_states,
-            attentions=transformer_outputs.attentions,
-        )
-    @staticmethod
-    def _reorder_cache(
-        past_key_values: Tuple[Tuple[torch.Tensor]], beam_idx: torch.Tensor
-    ) -> Tuple[Tuple[torch.Tensor]]:
-        return tuple(
-            tuple(
-                past_state.index_select(0, beam_idx.to(past_state.device))
-                for past_state in layer_past
-            )
-            for layer_past in past_key_values
-        )
-    def chat(
-        self,
-        tokenizer: PreTrainedTokenizer,
-        query: str,
-        history: Optional[HistoryType],
-        system: str = "You are a helpful assistant.",
-        append_history: bool = True,
-        stream: Optional[bool] = _SENTINEL,
-        stop_words_ids: Optional[List[List[int]]] = None,
-        generation_config: Optional[GenerationConfig] = None,
-        **kwargs,
-    ) -> Tuple[str, HistoryType]:
-        generation_config = generation_config if generation_config is not None else self.generation_config
-        assert stream is _SENTINEL, _ERROR_STREAM_IN_CHAT
-        assert generation_config.chat_format == 'chatml', _ERROR_BAD_CHAT_FORMAT
-        if history is None:
-            history = []
-        if stop_words_ids is None:
-            stop_words_ids = []
-        max_window_size = kwargs.get('max_window_size', None)
-        if max_window_size is None:
-            max_window_size = generation_config.max_window_size
-        raw_text, context_tokens = make_context(
-            tokenizer,
-            query,
-            history=history,
-            system=system,
-            max_window_size=max_window_size,
-            chat_format=generation_config.chat_format,
-        )
-        stop_words_ids.extend(get_stop_words_ids(
-            generation_config.chat_format, tokenizer
-        ))
-        input_ids = torch.tensor([context_tokens]).to(self.device)
-        outputs = self.generate(
-                    input_ids,
-                    stop_words_ids=stop_words_ids,
-                    return_dict_in_generate=False,
-                    generation_config=generation_config,
-                    **kwargs,
-                )
-        response = decode_tokens(
-            outputs[0],
-            tokenizer,
-            raw_text_len=len(raw_text),
-            context_length=len(context_tokens),
-            chat_format=generation_config.chat_format,
-            verbose=False,
-            errors='replace'
-        )
-        if append_history:
-            history.append((query, response))
-        return response, history
-    def chat_stream(
-            self,
-            tokenizer: PreTrainedTokenizer,
-            query: str,
-            history: Optional[HistoryType],
-            system: str = "You are a helpful assistant.",
-            stop_words_ids: Optional[List[List[int]]] = None,
-            logits_processor: Optional[LogitsProcessorList] = None,
-            generation_config: Optional[GenerationConfig] = None,
-            **kwargs,
-    ) -> Generator[str, Any, None]:
-        generation_config = generation_config if generation_config is not None else self.generation_config
-        assert generation_config.chat_format == 'chatml', _ERROR_BAD_CHAT_FORMAT
-        if history is None:
-            history = []
-        if stop_words_ids is None:
-            stop_words_ids = []
-        max_window_size = kwargs.get('max_window_size', None)
-        if max_window_size is None:
-            max_window_size = generation_config.max_window_size
-        raw_text, context_tokens = make_context(
-            tokenizer,
-            query,
-            history=history,
-            system=system,
-            max_window_size=max_window_size,
-            chat_format=generation_config.chat_format,
-        )
-        stop_words_ids.extend(get_stop_words_ids(
-            generation_config.chat_format, tokenizer
-        ))
-        if stop_words_ids is not None:
-            stop_words_logits_processor = StopWordsLogitsProcessor(
-                stop_words_ids=stop_words_ids,
-                eos_token_id=generation_config.eos_token_id,
-            )
-            if logits_processor is None:
-                logits_processor = LogitsProcessorList([stop_words_logits_processor])
-            else:
-                logits_processor.append(stop_words_logits_processor)
-        input_ids = torch.tensor([context_tokens]).to(self.device)
-        from transformers_stream_generator.main import NewGenerationMixin, StreamGenerationConfig
-        self.__class__.generate_stream = NewGenerationMixin.generate
-        self.__class__.sample_stream = NewGenerationMixin.sample_stream
-        stream_config = StreamGenerationConfig(**generation_config.to_dict(), do_stream=True)
-        def stream_generator():
-            outputs = []
-            for token in self.generate_stream(
-                    input_ids,
-                    return_dict_in_generate=False,
-                    generation_config=stream_config,
-                    logits_processor=logits_processor,
-                    seed=-1,
-                    **kwargs):
-                outputs.append(token.item())
-                yield tokenizer.decode(outputs, skip_special_tokens=True, errors='ignore', keep_image_special=True)
-        return stream_generator()
-    def generate(
-        self,
-        inputs: Optional[torch.Tensor] = None,
-        generation_config: Optional[GenerationConfig] = None,
-        logits_processor: Optional[LogitsProcessorList] = None,
-        stopping_criteria: Optional[StoppingCriteriaList] = None,
-        prefix_allowed_tokens_fn: Optional[
-            Callable[[int, torch.Tensor], List[int]]
-        ] = None,
-        synced_gpus: Optional[bool] = None,
-        assistant_model: Optional["PreTrainedModel"] = None,
-        streamer: Optional["BaseStreamer"] = None,
-        **kwargs,
-    ) -> Union[GenerateOutput, torch.LongTensor]:
-        generation_config = generation_config if generation_config is not None else self.generation_config
-        # Process stop_words_ids.
-        stop_words_ids = kwargs.pop("stop_words_ids", None)
-        if stop_words_ids is None and generation_config is not None:
-            stop_words_ids = getattr(generation_config, "stop_words_ids", None)
-        if stop_words_ids is None:
-            stop_words_ids = getattr(generation_config, "stop_words_ids", None)
-        if stop_words_ids is not None:
-            stop_words_logits_processor = StopWordsLogitsProcessor(
-                stop_words_ids=stop_words_ids,
-                eos_token_id=generation_config.eos_token_id,
-            )
-            if logits_processor is None:
-                logits_processor = LogitsProcessorList([stop_words_logits_processor])
-            else:
-                logits_processor.append(stop_words_logits_processor)
-        return super().generate(
-            inputs,
-            generation_config=generation_config,
-            logits_processor=logits_processor,
-            stopping_criteria=stopping_criteria,
-            prefix_allowed_tokens_fn=prefix_allowed_tokens_fn,
-            synced_gpus=synced_gpus,
-            assistant_model=assistant_model,
-            streamer=streamer,
-            **kwargs,
-        )
-class RotaryEmbedding(torch.nn.Module):
-    def __init__(self, dim, base=10000):
-        super().__init__()
-        self.dim = dim
-        self.base = base
-        self.inv_freq = 1.0 / (base ** (torch.arange(0, dim, 2).float() / dim))
-        if importlib.util.find_spec("einops") is None:
-            raise RuntimeError("einops is required for Rotary Embedding")
-        self._rotary_pos_emb_cache = None
-        self._seq_len_cached = 0
-        self._ntk_alpha_cached = 1.0
-    def update_rotary_pos_emb_cache(self, max_seq_len, offset=0, ntk_alpha=1.0):
-        seqlen = max_seq_len + offset
-        if seqlen > self._seq_len_cached or ntk_alpha != self._ntk_alpha_cached:
-            base = self.base * ntk_alpha ** (self.dim / (self.dim - 2))
-            self.inv_freq = 1.0 / (
-                base
-                ** (
-                    torch.arange(0, self.dim, 2, device=self.inv_freq.device).float()
-                    / self.dim
-                )
-            )
-            self._seq_len_cached = max(2 * seqlen, 16)
-            self._ntk_alpha_cached = ntk_alpha
-            seq = torch.arange(self._seq_len_cached, device=self.inv_freq.device)
-            freqs = torch.outer(seq.type_as(self.inv_freq), self.inv_freq)
-            emb = torch.cat((freqs, freqs), dim=-1)
-            from einops import rearrange
-            emb = rearrange(emb, "n d -> 1 n 1 d")
-            cos, sin = emb.cos(), emb.sin()
-            self._rotary_pos_emb_cache = [cos, sin]
-    def forward(self, max_seq_len, offset=0, ntk_alpha=1.0):
-        self.update_rotary_pos_emb_cache(max_seq_len, offset, ntk_alpha)
-        cos, sin = self._rotary_pos_emb_cache
-        return [cos[:, offset : offset + max_seq_len], sin[:, offset : offset + max_seq_len]]
-def _rotate_half(x):
-    from einops import rearrange
-    x = rearrange(x, "... (j d) -> ... j d", j=2)
-    x1, x2 = x.unbind(dim=-2)
-    return torch.cat((-x2, x1), dim=-1)
-def apply_rotary_pos_emb(t, freqs):
-    cos, sin = freqs
-    if apply_rotary_emb_func is not None and t.is_cuda:
-        t_ = t.float()
-        cos = cos.squeeze(0).squeeze(1)[:, : cos.shape[-1] // 2]
-        sin = sin.squeeze(0).squeeze(1)[:, : sin.shape[-1] // 2]
-        output = apply_rotary_emb_func(t_, cos, sin).type_as(t)
-        return output
-    else:
-        rot_dim = freqs[0].shape[-1]
-        cos, sin = freqs
-        t_, t_pass_ = t[..., :rot_dim], t[..., rot_dim:]
-        t_ = t_.float()
-        t_pass_ = t_pass_.float()
-        t_ = (t_ * cos) + (_rotate_half(t_) * sin)
-        return torch.cat((t_, t_pass_), dim=-1).type_as(t)
-class RMSNorm(torch.nn.Module):
-    def __init__(self, dim: int, eps: float = 1e-6):
-        super().__init__()
-        self.eps = eps
-        self.weight = nn.Parameter(torch.ones(dim))
-    def _norm(self, x):
-        return x * torch.rsqrt(x.pow(2).mean(-1, keepdim=True) + self.eps)
-    def forward(self, x):
-        if rms_norm is not None and x.is_cuda:
-            return rms_norm(x, self.weight, self.eps)
-        else:
-            output = self._norm(x.float()).type_as(x)
-            return output * self.weight

zzzmmz/SegAgent-Model/qwen.tiktoken DELETED Viewed

The diff for this file is too large to render. See raw diff

zzzmmz/SegAgent-Model/qwen_generation_utils.py DELETED Viewed

@@ -1,420 +0,0 @@
-# Copyright (c) Alibaba Cloud.
-#
-# This source code is licensed under the license found in the
-# LICENSE file in the root directory of this source tree.
-"""Generation support."""
-from typing import Tuple, List, Union, Iterable
-import numpy as np
-import torch
-import torch.nn.functional as F
-from transformers import PreTrainedTokenizer
-from transformers import logging
-from transformers.generation import LogitsProcessor
-logger = logging.get_logger(__name__)
-# Types.
-HistoryType = List[Tuple[str, str]]
-TokensType = List[int]
-BatchTokensType = List[List[int]]
-def pad_batch(batch: BatchTokensType, pad_id: int, seq_length: int) -> BatchTokensType:
-    for tokens in batch:
-        context_length = len(tokens)
-        if context_length < seq_length:
-            tokens.extend([pad_id] * (seq_length - context_length))
-    return batch
-def get_ltor_masks_and_position_ids(
-    data,
-    eod_token,
-    reset_position_ids,
-    reset_attention_mask,
-    eod_mask_loss,
-):
-    """Build masks and position id for left to right model."""
-    # Extract batch size and sequence length.
-    micro_batch_size, seq_length = data.size()
-    # Attention mask (lower triangular).
-    if reset_attention_mask:
-        att_mask_batch = micro_batch_size
-    else:
-        att_mask_batch = 1
-    attention_mask = torch.tril(
-        torch.ones((att_mask_batch, seq_length, seq_length), device=data.device)
-    ).view(att_mask_batch, 1, seq_length, seq_length)
-    # Loss mask.
-    loss_mask = torch.ones(data.size(), dtype=torch.float, device=data.device)
-    if eod_mask_loss:
-        loss_mask[data == eod_token] = 0.0
-    # Position ids.
-    position_ids = torch.arange(seq_length, dtype=torch.long, device=data.device)
-    position_ids = position_ids.unsqueeze(0).expand_as(data)
-    # We need to clone as the ids will be modifed based on batch index.
-    if reset_position_ids:
-        position_ids = position_ids.clone()
-    if reset_position_ids or reset_attention_mask:
-        # Loop through the batches:
-        for b in range(micro_batch_size):
-            # Find indecies where EOD token is.
-            eod_index = position_ids[b, data[b] == eod_token]
-            # Detach indecies from positions if going to modify positions.
-            if reset_position_ids:
-                eod_index = eod_index.clone()
-            # Loop through EOD indecies:
-            prev_index = 0
-            for j in range(eod_index.size()[0]):
-                i = eod_index[j]
-                # Mask attention loss.
-                if reset_attention_mask:
-                    attention_mask[b, 0, (i + 1) :, : (i + 1)] = 0
-                # Reset positions.
-                if reset_position_ids:
-                    position_ids[b, (i + 1) :] -= i + 1 - prev_index
-                    prev_index = i + 1
-    # Convert attention mask to binary:
-    attention_mask = attention_mask < 0.5
-    return attention_mask, loss_mask, position_ids
-def get_batch(context_tokens: torch.LongTensor, eod_id: int):
-    """Generate batch from context tokens."""
-    # Move to GPU.
-    tokens = context_tokens.contiguous().to(context_tokens.device)
-    # Get the attention mask and postition ids.
-    attention_mask, _, position_ids = get_ltor_masks_and_position_ids(
-        tokens,
-        eod_id,
-        reset_position_ids=False,
-        reset_attention_mask=False,
-        eod_mask_loss=False,
-    )
-    return tokens, attention_mask, position_ids
-def get_stop_words_ids(chat_format, tokenizer):
-    if chat_format == "raw":
-        stop_words_ids = [tokenizer.encode("Human:"), [tokenizer.eod_id]]
-    elif chat_format == "chatml":
-        stop_words_ids = [[tokenizer.im_end_id], [tokenizer.im_start_id]]
-    else:
-        raise NotImplementedError(f"Unknown chat format {chat_format!r}")
-    return stop_words_ids
-def make_context(
-    tokenizer: PreTrainedTokenizer,
-    query: str,
-    history: List[Tuple[str, str]] = None,
-    system: str = "",
-    max_window_size: int = 6144,
-    chat_format: str = "chatml",
-):
-    if history is None:
-        history = []
-    if chat_format == "chatml":
-        im_start, im_end = "<|im_start|>", "<|im_end|>"
-        im_start_tokens = [tokenizer.im_start_id]
-        im_end_tokens = [tokenizer.im_end_id]
-        nl_tokens = tokenizer.encode("\n")
-        def _tokenize_str(role, content):
-            return f"{role}\n{content}", tokenizer.encode(
-                role, allowed_special=set(tokenizer.IMAGE_ST)
-            ) + nl_tokens + tokenizer.encode(content, allowed_special=set(tokenizer.IMAGE_ST))
-        system_text, system_tokens_part = _tokenize_str("system", system)
-        system_tokens = im_start_tokens + system_tokens_part + im_end_tokens
-        raw_text = ""
-        context_tokens = []
-        for turn_query, turn_response in reversed(history):
-            query_text, query_tokens_part = _tokenize_str("user", turn_query)
-            query_tokens = im_start_tokens + query_tokens_part + im_end_tokens
-            if turn_response is not None:
-                response_text, response_tokens_part = _tokenize_str(
-                    "assistant", turn_response
-                )
-                response_tokens = im_start_tokens + response_tokens_part + im_end_tokens
-                next_context_tokens = nl_tokens + query_tokens + nl_tokens + response_tokens
-                prev_chat = (
-                    f"\n{im_start}{query_text}{im_end}\n{im_start}{response_text}{im_end}"
-                )
-            else:
-                next_context_tokens = nl_tokens + query_tokens + nl_tokens
-                prev_chat = f"\n{im_start}{query_text}{im_end}\n"
-            current_context_size = (
-                len(system_tokens) + len(next_context_tokens) + len(context_tokens)
-            )
-            if current_context_size < max_window_size:
-                context_tokens = next_context_tokens + context_tokens
-                raw_text = prev_chat + raw_text
-            else:
-                break
-        context_tokens = system_tokens + context_tokens
-        raw_text = f"{im_start}{system_text}{im_end}" + raw_text
-        context_tokens += (
-            nl_tokens
-            + im_start_tokens
-            + _tokenize_str("user", query)[1]
-            + im_end_tokens
-            + nl_tokens
-            + im_start_tokens
-            + tokenizer.encode("assistant")
-            + nl_tokens
-        )
-        raw_text += f"\n{im_start}user\n{query}{im_end}\n{im_start}assistant\n"
-    elif chat_format == "raw":
-        raw_text = query
-        context_tokens = tokenizer.encode(raw_text)
-    else:
-        raise NotImplementedError(f"Unknown chat format {chat_format!r}")
-    return raw_text, context_tokens
-def _decode_default(
-    tokens: List[int],
-    *,
-    stop_words: List[str],
-    eod_words: List[str],
-    tokenizer: PreTrainedTokenizer,
-    raw_text_len: int,
-    verbose: bool = False,
-    return_end_reason: bool = False,
-    errors: str='replace',
-):
-    trim_decode_tokens = tokenizer.decode(tokens, errors=errors)[raw_text_len:]
-    if verbose:
-        print("\nRaw Generate: ", trim_decode_tokens)
-    end_reason = f"Gen length {len(tokens)}"
-    for stop_word in stop_words:
-        trim_decode_tokens = trim_decode_tokens.replace(stop_word, "").strip()
-    for eod_word in eod_words:
-        if eod_word in trim_decode_tokens:
-            end_reason = f"Gen {eod_word!r}"
-        trim_decode_tokens = trim_decode_tokens.split(eod_word)[0]
-    trim_decode_tokens = trim_decode_tokens.strip()
-    if verbose:
-        print("\nEnd Reason:", end_reason)
-        print("\nGenerate: ", trim_decode_tokens)
-    if return_end_reason:
-        return trim_decode_tokens, end_reason
-    else:
-        return trim_decode_tokens
-def _decode_chatml(
-    tokens: List[int],
-    *,
-    stop_words: List[str],
-    eod_token_ids: List[int],
-    tokenizer: PreTrainedTokenizer,
-    raw_text_len: int,
-    context_length: int,
-    verbose: bool = False,
-    return_end_reason: bool = False,
-    errors: str='replace'
-):
-    end_reason = f"Gen length {len(tokens)}"
-    eod_token_idx = context_length
-    for eod_token_idx in range(context_length, len(tokens)):
-        if tokens[eod_token_idx] in eod_token_ids:
-            end_reason = f"Gen {tokenizer.decode([tokens[eod_token_idx]])!r}"
-            break
-    trim_decode_tokens = tokenizer.decode(tokens[:eod_token_idx], errors=errors)[raw_text_len:]
-    if verbose:
-        print("\nRaw Generate w/o EOD:", tokenizer.decode(tokens, errors=errors)[raw_text_len:])
-        print("\nRaw Generate:", trim_decode_tokens)
-        print("\nEnd Reason:", end_reason)
-    for stop_word in stop_words:
-        trim_decode_tokens = trim_decode_tokens.replace(stop_word, "").strip()
-    trim_decode_tokens = trim_decode_tokens.strip()
-    if verbose:
-        print("\nGenerate:", trim_decode_tokens)
-    if return_end_reason:
-        return trim_decode_tokens, end_reason
-    else:
-        return trim_decode_tokens
-def decode_tokens(
-    tokens: Union[torch.LongTensor, TokensType],
-    tokenizer: PreTrainedTokenizer,
-    raw_text_len: int,
-    context_length: int,
-    chat_format: str,
-    verbose: bool = False,
-    return_end_reason: bool = False,
-    errors: str="replace",
-) -> str:
-    if torch.is_tensor(tokens):
-        tokens = tokens.cpu().numpy().tolist()
-    if chat_format == "chatml":
-        return _decode_chatml(
-            tokens,
-            stop_words=[],
-            eod_token_ids=[tokenizer.im_start_id, tokenizer.im_end_id],
-            tokenizer=tokenizer,
-            raw_text_len=raw_text_len,
-            context_length=context_length,
-            verbose=verbose,
-            return_end_reason=return_end_reason,
-            errors=errors,
-        )
-    elif chat_format == "raw":
-        return _decode_default(
-            tokens,
-            stop_words=["<|endoftext|>"],
-            eod_words=["<|endoftext|>"],
-            tokenizer=tokenizer,
-            raw_text_len=raw_text_len,
-            verbose=verbose,
-            return_end_reason=return_end_reason,
-            errors=errors,
-        )
-    else:
-        raise NotImplementedError(f"Unknown chat format {chat_format!r}")
-class StopWordsLogitsProcessor(LogitsProcessor):
-    """
-    :class:`transformers.LogitsProcessor` that enforces that when specified sequences appear, stop geration.
-    Args:
-        stop_words_ids (:obj:`List[List[int]]`):
-            List of list of token ids of stop ids. In order to get the tokens of the words
-            that should not appear in the generated text, use :obj:`tokenizer(bad_word,
-            add_prefix_space=True).input_ids`.
-        eos_token_id (:obj:`int`):
-            The id of the `end-of-sequence` token.
-    """
-    def __init__(self, stop_words_ids: Iterable[Iterable[int]], eos_token_id: int):
-        if not isinstance(stop_words_ids, List) or len(stop_words_ids) == 0:
-            raise ValueError(
-                f"`stop_words_ids` has to be a non-emtpy list, but is {stop_words_ids}."
-            )
-        if any(not isinstance(bad_word_ids, list) for bad_word_ids in stop_words_ids):
-            raise ValueError(
-                f"`stop_words_ids` has to be a list of lists, but is {stop_words_ids}."
-            )
-        if any(
-            any(
-                (not isinstance(token_id, (int, np.integer)) or token_id < 0)
-                for token_id in stop_word_ids
-            )
-            for stop_word_ids in stop_words_ids
-        ):
-            raise ValueError(
-                f"Each list in `stop_words_ids` has to be a list of positive integers, but is {stop_words_ids}."
-            )
-        self.stop_words_ids = list(
-            filter(
-                lambda bad_token_seq: bad_token_seq != [eos_token_id], stop_words_ids
-            )
-        )
-        self.eos_token_id = eos_token_id
-        for stop_token_seq in self.stop_words_ids:
-            assert (
-                len(stop_token_seq) > 0
-            ), "Stop words token sequences {} cannot have an empty list".format(
-                stop_words_ids
-            )
-    def __call__(
-        self, input_ids: torch.LongTensor, scores: torch.FloatTensor
-    ) -> torch.FloatTensor:
-        stopped_samples = self._calc_stopped_samples(input_ids)
-        for i, should_stop in enumerate(stopped_samples):
-            if should_stop:
-                scores[i, self.eos_token_id] = float(2**15)
-        return scores
-    def _tokens_match(self, prev_tokens: torch.LongTensor, tokens: List[int]) -> bool:
-        if len(tokens) == 0:
-            # if bad word tokens is just one token always ban it
-            return True
-        elif len(tokens) > len(prev_tokens):
-            # if bad word tokens are longer then prev input_ids they can't be equal
-            return False
-        elif prev_tokens[-len(tokens) :].tolist() == tokens:
-            # if tokens match
-            return True
-        else:
-            return False
-    def _calc_stopped_samples(self, prev_input_ids: Iterable[int]) -> Iterable[int]:
-        stopped_samples = []
-        for prev_input_ids_slice in prev_input_ids:
-            match = False
-            for stop_token_seq in self.stop_words_ids:
-                if self._tokens_match(prev_input_ids_slice, stop_token_seq):
-                    # if tokens do not match continue
-                    match = True
-                    break
-            stopped_samples.append(match)
-        return stopped_samples
-def top_k_logits(logits, top_k=0, top_p=0.0, filter_value=-float("Inf")):
-    """This function has been mostly taken from huggingface conversational
-    ai code at
-        https://medium.com/huggingface/how-to-build-a-state-of-the-art-
-             conversational-ai-with-transfer-learning-2d818ac26313"""
-    if top_k > 0:
-        # Remove all tokens with a probability less than the
-        # last token of the top-k
-        indices_to_remove = logits < torch.topk(logits, top_k)[0][..., -1, None]
-        logits[indices_to_remove] = filter_value
-    if top_p > 0.0:
-        # Cconvert to 1D
-        sorted_logits, sorted_indices = torch.sort(logits, descending=True, dim=-1)
-        cumulative_probs = torch.cumsum(F.softmax(sorted_logits, dim=-1), dim=-1)
-        # Remove tokens with cumulative probability above the threshold
-        sorted_indices_to_remove = cumulative_probs > top_p
-        # Shift the indices to the right to keep also the first token
-        # above the threshold
-        sorted_indices_to_remove[..., 1:] = sorted_indices_to_remove[..., :-1].clone()
-        sorted_indices_to_remove[..., 0] = 0
-        for i in range(sorted_indices.size(0)):
-            indices_to_remove = sorted_indices[i][sorted_indices_to_remove[i]]
-            logits[i][indices_to_remove] = filter_value
-    return logits
-def switch(val1, val2, boolean):
-    boolean = boolean.type_as(val1)
-    return (1 - boolean) * val1 + boolean * val2

zzzmmz/SegAgent-Model/special_tokens_map.json DELETED Viewed

@@ -1,3 +0,0 @@
-{
-  "pad_token": "<|endoftext|>"
-}

zzzmmz/SegAgent-Model/tokenization_qwen.py DELETED Viewed

@@ -1,598 +0,0 @@
-# Copyright (c) Alibaba Cloud.
-#
-# This source code is licensed under the license found in the
-# LICENSE file in the root directory of this source tree.
-"""Tokenization classes for QWen."""
-import base64
-import logging
-import os
-import requests
-import unicodedata
-from typing import Collection, Dict, List, Set, Tuple, Union, Any, Callable, Optional
-import tiktoken
-import numpy as np
-from PIL import Image
-from PIL import ImageFont
-from PIL import ImageDraw
-from transformers import PreTrainedTokenizer, AddedToken
-from transformers.utils import try_to_load_from_cache
-import matplotlib.colors as mcolors
-from matplotlib.font_manager import FontProperties
-logger = logging.getLogger(__name__)
-VOCAB_FILES_NAMES = {"vocab_file": "qwen.tiktoken", "ttf": "SimSun.ttf"}
-FONT_PATH = try_to_load_from_cache("Qwen/Qwen-VL-Chat", "SimSun.ttf")
-if FONT_PATH is None:
-    if not os.path.exists("SimSun.ttf"):
-        ttf = requests.get("https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/SimSun.ttf")
-        open("SimSun.ttf", "wb").write(ttf.content)
-    FONT_PATH = "SimSun.ttf"
-PAT_STR = r"""(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\r\n\p{L}\p{N}]?\p{L}+|\p{N}| ?[^\s\p{L}\p{N}]+[\r\n]*|\s*[\r\n]+|\s+(?!\S)|\s+"""
-ENDOFTEXT = "<|endoftext|>"
-IMSTART = "<|im_start|>"
-IMEND = "<|im_end|>"
-# as the default behavior is changed to allow special tokens in
-# regular texts, the surface forms of special tokens need to be
-# as different as possible to minimize the impact
-EXTRAS = tuple((f"<|extra_{i}|>" for i in range(205)))
-SPECIAL_TOKENS = (
-    ENDOFTEXT,
-    IMSTART,
-    IMEND,
-) + EXTRAS
-IMG_TOKEN_SPAN = 256
-def _load_tiktoken_bpe(tiktoken_bpe_file: str) -> Dict[bytes, int]:
-    with open(tiktoken_bpe_file, "rb") as f:
-        contents = f.read()
-    return {
-        base64.b64decode(token): int(rank)
-        for token, rank in (line.split() for line in contents.splitlines() if line)
-    }
-def _list_find(
-    input_list: List[Any],
-    candidates: Tuple[Any],
-    start: int = 0,
-):
-    for i in range(start, len(input_list)):
-        if input_list[i] in candidates:
-            return i
-    return -1
-def _replace_closed_tag(
-    input_tokens: List[Any],
-    start_tags: Union[Any, Tuple[Any]],
-    end_tags: Union[Any, Tuple[Any]],
-    inclusive_replace_func: Callable,
-    exclusive_replace_func: Callable = lambda x: x,
-):
-    if isinstance(start_tags, (str, int)):
-        start_tags = (start_tags,)
-    if isinstance(end_tags, (str, int)):
-        end_tags = (end_tags,)
-    assert len(start_tags) == len(end_tags)
-    output_tokens = []
-    end = 0
-    while True:
-        start = _list_find(input_tokens, start_tags, end)
-        if start == -1:
-            break
-        output_tokens.extend(exclusive_replace_func(input_tokens[end : start]))
-        tag_idx = start_tags.index(input_tokens[start])
-        end = _list_find(input_tokens, (end_tags[tag_idx],), start)
-        if end == -1:
-            raise ValueError("Unclosed image token")
-        output_tokens.extend(inclusive_replace_func(input_tokens[start : end + 1]))
-        end += 1
-    output_tokens.extend(exclusive_replace_func(input_tokens[end : ]))
-    return output_tokens
-class QWenTokenizer(PreTrainedTokenizer):
-    """QWen tokenizer."""
-    vocab_files_names = VOCAB_FILES_NAMES
-    def __init__(
-        self,
-        vocab_file,
-        errors="replace",
-        image_start_tag='<img>',
-        image_end_tag='</img>',
-        image_pad_tag='<imgpad>',
-        ref_start_tag='<ref>',
-        ref_end_tag='</ref>',
-        box_start_tag='<box>',
-        box_end_tag='</box>',
-        quad_start_tag='<quad>',
-        quad_end_tag='</quad>',
-        **kwargs,
-    ):
-        super().__init__(**kwargs)
-        self.image_start_tag = image_start_tag
-        self.image_end_tag = image_end_tag
-        self.image_pad_tag = image_pad_tag
-        self.ref_start_tag = ref_start_tag
-        self.ref_end_tag = ref_end_tag
-        self.box_start_tag = box_start_tag
-        self.box_end_tag = box_end_tag
-        self.quad_start_tag = quad_start_tag
-        self.quad_end_tag = quad_end_tag
-        self.IMAGE_ST = (
-            ref_start_tag, ref_end_tag,
-            box_start_tag, box_end_tag,
-            quad_start_tag, quad_end_tag,
-            image_start_tag, image_end_tag,
-            image_pad_tag
-        )
-        self.errors = errors  # how to handle errors in decoding
-        self.mergeable_ranks = _load_tiktoken_bpe(vocab_file)  # type: dict[bytes, int]
-        self.special_tokens = {
-            token: index
-            for index, token in enumerate(
-                SPECIAL_TOKENS + self.IMAGE_ST, start=len(self.mergeable_ranks)
-            )
-        }
-        self.img_start_id = self.special_tokens[self.image_start_tag]
-        self.img_end_id = self.special_tokens[self.image_end_tag]
-        self.img_pad_id = self.special_tokens[self.image_pad_tag]
-        self.ref_start_id = self.special_tokens[self.ref_start_tag]
-        self.ref_end_id = self.special_tokens[self.ref_end_tag]
-        self.box_start_id = self.special_tokens[self.box_start_tag]
-        self.box_end_id = self.special_tokens[self.box_end_tag]
-        self.quad_start_id = self.special_tokens[self.quad_start_tag]
-        self.quad_end_id = self.special_tokens[self.quad_end_tag]
-        self.image_special_tokens = set([
-            self.ref_start_id, self.ref_end_id, self.box_start_id, self.box_end_id,
-            self.quad_start_id, self.quad_end_id,
-        ])
-        enc = tiktoken.Encoding(
-            "Qwen",
-            pat_str=PAT_STR,
-            mergeable_ranks=self.mergeable_ranks,
-            special_tokens=self.special_tokens,
-        )
-        assert (
-            len(self.mergeable_ranks) + len(self.special_tokens) == enc.n_vocab
-        ), f"{len(self.mergeable_ranks) + len(self.special_tokens)} != {enc.n_vocab} in encoding"
-        self.decoder = {
-            v: k for k, v in self.mergeable_ranks.items()
-        }  # type: dict[int, bytes|str]
-        self.decoder.update({v: k for k, v in self.special_tokens.items()})
-        self.tokenizer = enc  # type: tiktoken.Encoding
-        self.eod_id = self.tokenizer.eot_token
-        self.im_start_id = self.special_tokens[IMSTART]
-        self.im_end_id = self.special_tokens[IMEND]
-    def __getstate__(self):
-        # for pickle lovers
-        state = self.__dict__.copy()
-        del state['tokenizer']
-        return state
-    def __setstate__(self, state):
-        # tokenizer is not python native; don't pass it; rebuild it
-        self.__dict__.update(state)
-        enc = tiktoken.Encoding(
-            "Qwen",
-            pat_str=PAT_STR,
-            mergeable_ranks=self.mergeable_ranks,
-            special_tokens=self.special_tokens,
-        )
-        self.tokenizer = enc
-    def __len__(self) -> int:
-        return self.tokenizer.n_vocab
-    def get_vocab(self) -> Dict[bytes, int]:
-        return self.mergeable_ranks
-    def convert_tokens_to_ids(
-        self, tokens: Union[bytes, str, List[Union[bytes, str]]]
-    ) -> List[int]:
-        ids = []
-        if isinstance(tokens, (str, bytes)):
-            if tokens in self.special_tokens:
-                return self.special_tokens[tokens]
-            else:
-                return self.mergeable_ranks.get(tokens)
-        for token in tokens:
-            if token in self.special_tokens:
-                ids.append(self.special_tokens[token])
-            else:
-                ids.append(self.mergeable_ranks.get(token))
-        return ids
-    def _add_tokens(self, new_tokens: Union[List[str], List[AddedToken]], special_tokens: bool = False) -> int:
-        if not special_tokens and new_tokens:
-            raise ValueError('Adding regular tokens is not supported')
-        for token in new_tokens:
-            surface_form = token.content if isinstance(token, AddedToken) else token
-            if surface_form not in SPECIAL_TOKENS + self.IMAGE_ST:
-                raise ValueError('Adding unknown special tokens is not supported')
-        return 0
-    def save_vocabulary(self, save_directory: str, **kwargs) -> Tuple[str]:
-        """
-        Save only the vocabulary of the tokenizer (vocabulary).
-        Returns:
-            `Tuple(str)`: Paths to the files saved.
-        """
-        file_path = os.path.join(save_directory, "qwen.tiktoken")
-        with open(file_path, "w", encoding="utf8") as w:
-            for k, v in self.mergeable_ranks.items():
-                line = base64.b64encode(k).decode("utf8") + " " + str(v) + "\n"
-                w.write(line)
-        return (file_path,)
-    def tokenize(
-        self,
-        text: str,
-        allowed_special: Union[Set, str] = "all",
-        disallowed_special: Union[Collection, str] = (),
-        **kwargs,
-    ) -> List[Union[bytes, str]]:
-        """
-        Converts a string in a sequence of tokens.
-        Args:
-            text (`str`):
-                The sequence to be encoded.
-            allowed_special (`Literal["all"]` or `set`):
-                The surface forms of the tokens to be encoded as special tokens in regular texts.
-                Default to "all".
-            disallowed_special (`Literal["all"]` or `Collection`):
-                The surface forms of the tokens that should not be in regular texts and trigger errors.
-                Default to an empty tuple.
-            kwargs (additional keyword arguments, *optional*):
-                Will be passed to the underlying model specific encode method.
-        Returns:
-            `List[bytes|str]`: The list of tokens.
-        """
-        tokens = []
-        text = unicodedata.normalize("NFC", text)
-        # this implementation takes a detour: text -> token id -> token surface forms
-        for t in self.tokenizer.encode(
-            text, allowed_special=allowed_special, disallowed_special=disallowed_special
-        ):
-            tokens.append(self.decoder[t])
-        def _encode_imgurl(img_tokens):
-            assert img_tokens[0] == self.image_start_tag and img_tokens[-1] == self.image_end_tag
-            img_tokens = img_tokens[1:-1]
-            img_url = b''.join(img_tokens)
-            out_img_tokens = list(map(self.decoder.get, img_url))
-            if len(out_img_tokens) > IMG_TOKEN_SPAN:
-                raise ValueError("The content in {}..{} is too long".format(
-                    self.image_start_tag, self.image_end_tag))
-            out_img_tokens.extend([self.image_pad_tag] * (IMG_TOKEN_SPAN - len(out_img_tokens)))
-            out_img_tokens = [self.image_start_tag] + out_img_tokens + [self.image_end_tag]
-            return out_img_tokens
-        return _replace_closed_tag(tokens, self.image_start_tag, self.image_end_tag, _encode_imgurl)
-    def convert_tokens_to_string(self, tokens: List[Union[bytes, str]]) -> str:
-        """
-        Converts a sequence of tokens in a single string.
-        """
-        text = ""
-        temp = b""
-        for t in tokens:
-            if isinstance(t, str):
-                if temp:
-                    text += temp.decode("utf-8", errors=self.errors)
-                    temp = b""
-                text += t
-            elif isinstance(t, bytes):
-                temp += t
-            else:
-                raise TypeError("token should only be of type types or str")
-        if temp:
-            text += temp.decode("utf-8", errors=self.errors)
-        return text
-    @property
-    def vocab_size(self):
-        return self.tokenizer.n_vocab
-    def _convert_id_to_token(self, index: int) -> Union[bytes, str]:
-        """Converts an id to a token, special tokens included"""
-        if index in self.decoder:
-            return self.decoder[index]
-        raise ValueError("unknown ids")
-    def _convert_token_to_id(self, token: Union[bytes, str]) -> int:
-        """Converts a token to an id using the vocab, special tokens included"""
-        if token in self.special_tokens:
-            return self.special_tokens[token]
-        if token in self.mergeable_ranks:
-            return self.mergeable_ranks[token]
-        raise ValueError("unknown token")
-    def _tokenize(self, text: str, **kwargs):
-        """
-        Converts a string in a sequence of tokens (string), using the tokenizer. Split in words for word-based
-        vocabulary or sub-words for sub-word-based vocabularies (BPE/SentencePieces/WordPieces).
-        Do NOT take care of added tokens.
-        """
-        raise NotImplementedError
-    def _decode(
-        self,
-        token_ids: Union[int, List[int]],
-        skip_special_tokens: bool = False,
-        errors: str = None,
-        **kwargs,
-    ) -> str:
-        if isinstance(token_ids, int):
-            token_ids = [token_ids]
-        def _decode_imgurl(img_token_ids):
-            assert img_token_ids[0] == self.img_start_id and img_token_ids[-1] == self.img_end_id
-            img_token_ids = img_token_ids[1:-1]
-            img_token_ids = img_token_ids[ : img_token_ids.index(self.img_pad_id)]
-            img_url = bytes(img_token_ids).decode('utf-8')
-            return [self.img_start_id] + self.tokenizer.encode(img_url) + [self.img_end_id]
-        token_ids = _replace_closed_tag(token_ids, self.img_start_id, self.img_end_id, _decode_imgurl)
-        if skip_special_tokens:
-            if kwargs.get('keep_image_special', False):
-                token_ids = [i for i in token_ids if i < self.eod_id
-                    or i in self.image_special_tokens]
-            else:
-                token_ids = [i for i in token_ids if i < self.eod_id]
-        return self.tokenizer.decode(token_ids, errors=errors or self.errors)
-    def to_list_format(self, text: str):
-        text = unicodedata.normalize("NFC", text)
-        token_ids = self.tokenizer.encode(
-            text, allowed_special=set(self.IMAGE_ST + (ENDOFTEXT,)))
-        def _encode_vl_info(tokens):
-            if len(tokens) == 0:
-                return []
-            if tokens[0] == self.img_start_id and tokens[-1] == self.img_end_id:
-                key = 'image'
-            elif tokens[0] == self.ref_start_id and tokens[-1] == self.ref_end_id:
-                key = 'ref'
-            elif tokens[0] == self.box_start_id and tokens[-1] == self.box_end_id:
-                key = 'box'
-            elif tokens[0] == self.quad_start_id and tokens[-1] == self.quad_end_id:
-                key = 'quad'
-            else:
-                _tobytes = lambda x: x.encode('utf-8') if isinstance(x, str) else x
-                return [{'text': b''.join(map(_tobytes, map(self.decoder.get, tokens))).decode('utf-8')}]
-            _tobytes = lambda x: x.encode('utf-8') if isinstance(x, str) else x
-            val = b''.join(map(_tobytes, map(self.decoder.get, tokens[1:-1]))).decode('utf-8')
-            return [{key: val}]
-        return _replace_closed_tag(
-            token_ids,
-            (self.img_start_id, self.ref_start_id, self.box_start_id, self.quad_start_id),
-            (self.img_end_id, self.ref_end_id, self.box_end_id, self.quad_end_id),
-            _encode_vl_info,
-            _encode_vl_info,
-        )
-    def from_list_format(self, list_format: List[Dict]):
-        text = ''
-        num_images = 0
-        for ele in list_format:
-            if 'image' in ele:
-                num_images += 1
-                text += f'Picture {num_images}: '
-                text += self.image_start_tag + ele['image'] + self.image_end_tag
-                text += '\n'
-            elif 'text' in ele:
-                text += ele['text']
-            elif 'box' in ele:
-                if 'ref' in ele:
-                    text += self.ref_start_tag + ele['ref'] + self.ref_end_tag
-                for box in ele['box']:
-                    text += self.box_start_tag + '(%d,%d),(%d,%d)' % (box[0], box[1], box[2], box[3]) + self.box_end_tag
-            else:
-                raise ValueError("Unsupport element: " + str(ele))
-        return text
-    def _fetch_latest_picture(self, response, history):
-        if history is None:
-            history = []
-        _history = history + [(response, None)]
-        for q, r in _history[::-1]:
-            for ele in self.to_list_format(q)[::-1]:
-                if 'image' in ele:
-                    return ele['image']
-        return None
-    def _fetch_all_box_with_ref(self, text):
-        list_format = self.to_list_format(text)
-        output = []
-        for i, ele in enumerate(list_format):
-            if 'box' in ele:
-                bbox = tuple(map(int, ele['box'].replace('(', '').replace(')', '').split(',')))
-                assert len(bbox) == 4
-                output.append({'box': bbox})
-                if i > 0 and 'ref' in list_format[i-1]:
-                    output[-1]['ref'] = list_format[i-1]['ref'].strip()
-        return output
-    def draw_bbox_on_latest_picture(
-        self,
-        response,
-        history=None,
-    ) -> Optional[Image.Image]:
-        image = self._fetch_latest_picture(response, history)
-        if image is None:
-            return None
-        if image.startswith("http://") or image.startswith("https://"):
-            image = Image.open(requests.get(image, stream=True).raw).convert("RGB")
-            h, w = image.height, image.width
-        else:
-            image = np.asarray(Image.open(image).convert("RGB"))
-            h, w = image.shape[0], image.shape[1]
-        visualizer = Visualizer(image)
-        boxes = self._fetch_all_box_with_ref(response)
-        if not boxes:
-            return None
-        color = random.choice([_ for _ in mcolors.TABLEAU_COLORS.keys()]) # init color
-        for box in boxes:
-            if 'ref' in box: # random new color for new refexps
-                color = random.choice([_ for _ in mcolors.TABLEAU_COLORS.keys()])
-            x1, y1, x2, y2 = box['box']
-            x1, y1, x2, y2 = (int(x1 / 1000 * w), int(y1 / 1000 * h), int(x2 / 1000 * w), int(y2 / 1000 * h))
-            visualizer.draw_box((x1, y1, x2, y2), alpha=1, edge_color=color)
-            if 'ref' in box:
-                visualizer.draw_text(box['ref'], (x1, y1), color=color, horizontal_alignment="left")
-        return visualizer.output
-import colorsys
-import logging
-import math
-import numpy as np
-import matplotlib as mpl
-import matplotlib.colors as mplc
-import matplotlib.figure as mplfigure
-import torch
-from matplotlib.backends.backend_agg import FigureCanvasAgg
-from PIL import Image
-import random
-logger = logging.getLogger(__name__)
-class VisImage:
-    def __init__(self, img, scale=1.0):
-        self.img = img
-        self.scale = scale
-        self.width, self.height = img.shape[1], img.shape[0]
-        self._setup_figure(img)
-    def _setup_figure(self, img):
-        fig = mplfigure.Figure(frameon=False)
-        self.dpi = fig.get_dpi()
-        # add a small 1e-2 to avoid precision lost due to matplotlib's truncation
-        # (https://github.com/matplotlib/matplotlib/issues/15363)
-        fig.set_size_inches(
-            (self.width * self.scale + 1e-2) / self.dpi,
-            (self.height * self.scale + 1e-2) / self.dpi,
-        )
-        self.canvas = FigureCanvasAgg(fig)
-        # self.canvas = mpl.backends.backend_cairo.FigureCanvasCairo(fig)
-        ax = fig.add_axes([0.0, 0.0, 1.0, 1.0])
-        ax.axis("off")
-        self.fig = fig
-        self.ax = ax
-        self.reset_image(img)
-    def reset_image(self, img):
-        img = img.astype("uint8")
-        self.ax.imshow(img, extent=(0, self.width, self.height, 0), interpolation="nearest")
-    def save(self, filepath):
-        self.fig.savefig(filepath)
-    def get_image(self):
-        canvas = self.canvas
-        s, (width, height) = canvas.print_to_buffer()
-        buffer = np.frombuffer(s, dtype="uint8")
-        img_rgba = buffer.reshape(height, width, 4)
-        rgb, alpha = np.split(img_rgba, [3], axis=2)
-        return rgb.astype("uint8")
-class Visualizer:
-    def __init__(self, img_rgb, metadata=None, scale=1.0):
-        self.img = np.asarray(img_rgb).clip(0, 255).astype(np.uint8)
-        self.font_path = FONT_PATH
-        self.output = VisImage(self.img, scale=scale)
-        self.cpu_device = torch.device("cpu")
-        # too small texts are useless, therefore clamp to 14
-        self._default_font_size = max(
-            np.sqrt(self.output.height * self.output.width) // 30, 15 // scale
-        )
-    def draw_text(
-        self,
-        text,
-        position,
-        *,
-        font_size=None,
-        color="g",
-        horizontal_alignment="center",
-        rotation=0,
-    ):
-        if not font_size:
-            font_size = self._default_font_size
-        # since the text background is dark, we don't want the text to be dark
-        color = np.maximum(list(mplc.to_rgb(color)), 0.2)
-        color[np.argmax(color)] = max(0.8, np.max(color))
-        x, y = position
-        self.output.ax.text(
-            x,
-            y,
-            text,
-            size=font_size * self.output.scale,
-            fontproperties=FontProperties(fname=self.font_path),
-            bbox={"facecolor": "black", "alpha": 0.8, "pad": 0.7, "edgecolor": "none"},
-            verticalalignment="top",
-            horizontalalignment=horizontal_alignment,
-            color=color,
-            zorder=10,
-            rotation=rotation,
-        )
-        return self.output
-    def draw_box(self, box_coord, alpha=0.5, edge_color="g", line_style="-"):
-        x0, y0, x1, y1 = box_coord
-        width = x1 - x0
-        height = y1 - y0
-        linewidth = max(self._default_font_size / 4, 1)
-        self.output.ax.add_patch(
-            mpl.patches.Rectangle(
-                (x0, y0),
-                width,
-                height,
-                fill=False,
-                edgecolor=edge_color,
-                linewidth=linewidth * self.output.scale,
-                alpha=alpha,
-                linestyle=line_style,
-            )
-        )
-        return self.output
-    def get_output(self):
-        return self.output

zzzmmz/SegAgent-Model/tokenizer_config.json DELETED Viewed

@@ -1,14 +0,0 @@
-{
-  "added_tokens_decoder": {},
-  "auto_map": {
-    "AutoTokenizer": [
-      "tokenization_qwen.QWenTokenizer",
-      null
-    ]
-  },
-  "clean_up_tokenization_spaces": true,
-  "model_max_length": 2048,
-  "pad_token": "<|endoftext|>",
-  "padding_side": "right",
-  "tokenizer_class": "QWenTokenizer"
-}

zzzmmz/SegAgent-Model/trainer_state.json DELETED Viewed

The diff for this file is too large to render. See raw diff

zzzmmz/SegAgent-Model/visual.py DELETED Viewed

@@ -1,472 +0,0 @@
-# Copyright (c) Alibaba Cloud.
-#
-# This source code is licensed under the license found in the
-# LICENSE file in the root directory of this source tree.
-from collections import OrderedDict
-import math
-import requests
-from io import BytesIO
-from functools import partial
-from PIL import Image,ImageOps
-from typing import Callable, Optional, Sequence, Tuple, List
-import numpy as np
-import torch
-from torch import nn
-from torch.nn import functional as F
-from torch.nn.init import normal_
-from torchvision import transforms
-from torchvision.transforms import InterpolationMode
-from pycocotools import mask as maskUtils
-import os
-import torchshow
-def overlay_mask(image, mask, color=(0, 255, 0), alpha=0.5):
-    overlay = image.copy()
-    binary_mask = mask
-    color_mask = np.zeros_like(image)
-    color_mask[binary_mask > 0] = color
-    for c in range(0, 3):
-        overlay[:, :, c] = np.where(binary_mask > 0,
-                                    overlay[:, :, c] * (1 - alpha) + color_mask[:, :, c] * alpha,
-                                    overlay[:, :, c])
-    return overlay
-def get_abs_pos(abs_pos, tgt_size):
-    # abs_pos: L, C
-    # tgt_size: M
-    # return: M, C
-    src_size = int(math.sqrt(abs_pos.size(0)))
-    tgt_size = int(math.sqrt(tgt_size))
-    dtype = abs_pos.dtype
-    if src_size != tgt_size:
-        return F.interpolate(
-            abs_pos.float().reshape(1, src_size, src_size, -1).permute(0, 3, 1, 2),
-            size=(tgt_size, tgt_size),
-            mode="bicubic",
-            align_corners=False,
-        ).permute(0, 2, 3, 1).flatten(0, 2).to(dtype=dtype)
-    else:
-        return abs_pos
-# https://github.com/facebookresearch/mae/blob/efb2a8062c206524e35e47d04501ed4f544c0ae8/util/pos_embed.py#L20
-def get_2d_sincos_pos_embed(embed_dim, grid_size, cls_token=False):
-    """
-    grid_size: int of the grid height and width
-    return:
-    pos_embed: [grid_size*grid_size, embed_dim] or [1+grid_size*grid_size, embed_dim] (w/ or w/o cls_token)
-    """
-    grid_h = np.arange(grid_size, dtype=np.float32)
-    grid_w = np.arange(grid_size, dtype=np.float32)
-    grid = np.meshgrid(grid_w, grid_h)  # here w goes first
-    grid = np.stack(grid, axis=0)
-    grid = grid.reshape([2, 1, grid_size, grid_size])
-    pos_embed = get_2d_sincos_pos_embed_from_grid(embed_dim, grid)
-    if cls_token:
-        pos_embed = np.concatenate([np.zeros([1, embed_dim]), pos_embed], axis=0)
-    return pos_embed
-def get_2d_sincos_pos_embed_from_grid(embed_dim, grid):
-    assert embed_dim % 2 == 0
-    # use half of dimensions to encode grid_h
-    emb_h = get_1d_sincos_pos_embed_from_grid(embed_dim // 2, grid[0])  # (H*W, D/2)
-    emb_w = get_1d_sincos_pos_embed_from_grid(embed_dim // 2, grid[1])  # (H*W, D/2)
-    emb = np.concatenate([emb_h, emb_w], axis=1) # (H*W, D)
-    return emb
-def get_1d_sincos_pos_embed_from_grid(embed_dim, pos):
-    """
-    embed_dim: output dimension for each position
-    pos: a list of positions to be encoded: size (M,)
-    out: (M, D)
-    """
-    assert embed_dim % 2 == 0
-    omega = np.arange(embed_dim // 2, dtype=np.float32)
-    omega /= embed_dim / 2.
-    omega = 1. / 10000**omega  # (D/2,)
-    pos = pos.reshape(-1)  # (M,)
-    out = np.einsum('m,d->md', pos, omega)  # (M, D/2), outer product
-    emb_sin = np.sin(out) # (M, D/2)
-    emb_cos = np.cos(out) # (M, D/2)
-    emb = np.concatenate([emb_sin, emb_cos], axis=1)  # (M, D)
-    return emb
-class Resampler(nn.Module):
-    """
-    A 2D perceiver-resampler network with one cross attention layers by
-        (grid_size**2) learnable queries and 2d sincos pos_emb
-    Outputs:
-        A tensor with the shape of (grid_size**2, embed_dim)
-    """
-    def __init__(
-            self,
-            grid_size,
-            embed_dim,
-            num_heads,
-            kv_dim=None,
-            norm_layer=nn.LayerNorm
-    ):
-        super().__init__()
-        self.num_queries = grid_size ** 2
-        self.embed_dim = embed_dim
-        self.num_heads = num_heads
-        self.pos_embed = nn.Parameter(
-            torch.from_numpy(get_2d_sincos_pos_embed(embed_dim, grid_size)).float()
-        ).requires_grad_(False)
-        self.query = nn.Parameter(torch.zeros(self.num_queries, embed_dim))
-        normal_(self.query, std=.02)
-        if kv_dim is not None and kv_dim != embed_dim:
-            self.kv_proj = nn.Linear(kv_dim, embed_dim, bias=False)
-        else:
-            self.kv_proj = nn.Identity()
-        self.attn = nn.MultiheadAttention(embed_dim, num_heads)
-        self.ln_q = norm_layer(embed_dim)
-        self.ln_kv = norm_layer(embed_dim)
-        # self.apply(self._init_weights)
-    def _init_weights(self, m):
-        if isinstance(m, nn.Linear):
-            normal_(m.weight, std=.02)
-            if isinstance(m, nn.Linear) and m.bias is not None:
-                nn.init.constant_(m.bias, 0)
-        elif isinstance(m, nn.LayerNorm):
-            nn.init.constant_(m.bias, 0)
-            nn.init.constant_(m.weight, 1.0)
-    def forward(self, x, attn_mask=None):
-        pos_embed = get_abs_pos(self.pos_embed, x.size(1))
-        x = self.kv_proj(x)
-        x = self.ln_kv(x).permute(1, 0, 2)
-        N = x.shape[1]
-        q = self.ln_q(self.query)
-        out = self.attn(
-            self._repeat(q, N) + self.pos_embed.unsqueeze(1),
-            x + pos_embed.unsqueeze(1),
-            x,
-            attn_mask=attn_mask)[0]
-        return out.permute(1, 0, 2)
-    def _repeat(self, query, N: int):
-        return query.unsqueeze(1).repeat(1, N, 1)
-class VisualAttention(nn.Module):
-    """self-attention layer class.
-    Self-attention layer takes input with size [s, b, h]
-    and returns output of the same size.
-    """
-    def __init__(self, embed_dim, num_heads,
-                 bias=True, kdim=None, vdim=None):
-        super(VisualAttention, self).__init__()
-        self.embed_dim = embed_dim
-        self.kdim = kdim if kdim is not None else embed_dim
-        self.vdim = vdim if vdim is not None else embed_dim
-        self._qkv_same_embed_dim = self.kdim == embed_dim and self.vdim == embed_dim
-        self.num_heads = num_heads
-        # Per attention head and per partition values.
-        assert embed_dim % num_heads == 0
-        self.hidden_size_per_attention_head = embed_dim // num_heads
-        self.num_attention_heads_per_partition = num_heads
-        self.hidden_size_per_partition = embed_dim
-        # Strided linear layer.
-        assert self._qkv_same_embed_dim, 'Only Support SelfAttention Currently'
-        self.in_proj = nn.Linear(embed_dim, 3 * embed_dim)
-        self.out_proj = nn.Linear(embed_dim, embed_dim)
-        self.norm_factor = math.sqrt(self.hidden_size_per_attention_head)
-    def forward(self, query, key, value, attn_mask = None):
-        # query/key/value: [sq, b, h]
-        sq, b, _ = query.size()
-        assert torch.allclose(query, key), 'Only Support Self-Attention Currently'
-        sk = sq
-        mixed_x_layer = self.in_proj(query)
-        # [sq, b, (np * 3 * hn)] --> [sq, b, np, 3 * hn]
-        new_tensor_shape = mixed_x_layer.size()[:-1] + \
-            (self.num_attention_heads_per_partition,
-             3 * self.hidden_size_per_attention_head)
-        mixed_x_layer = mixed_x_layer.view(*new_tensor_shape)
-        # [sq, b, np, 3 * hn] --> 3 [sq, b, np, hn]
-        query_layer, key_layer, value_layer = mixed_x_layer.split(
-            self.hidden_size_per_attention_head, dim=-1)
-        # [sq, b, np, hn] -> [sq, b * np, hn]
-        query_layer = query_layer.view(sq,
-            b * self.num_attention_heads_per_partition,
-            self.hidden_size_per_attention_head).transpose(0, 1)
-        # [sk, b, np, hn] -> [sk, b * np, hn]
-        key_layer = key_layer.view(sk,
-            b * self.num_attention_heads_per_partition,
-            self.hidden_size_per_attention_head).transpose(0, 1)
-        q_scaled = query_layer / self.norm_factor
-        if attn_mask is not None:
-            attention_probs = torch.baddbmm(attn_mask, q_scaled, key_layer.transpose(-2, -1))
-        else:
-            attention_probs = torch.bmm(q_scaled, key_layer.transpose(-2, -1))
-        attention_probs = attention_probs.softmax(dim=-1)
-        value_layer = value_layer.view(sk,
-            b * self.num_attention_heads_per_partition,
-            self.hidden_size_per_attention_head).transpose(0, 1)
-        # matmul: [b * np, sq, hn]
-        context_layer = torch.bmm(attention_probs, value_layer)
-        # change view [b, np, sq, hn]
-        context_layer = context_layer.view(b,
-            self.num_attention_heads_per_partition,
-            sq, self.hidden_size_per_attention_head)
-        # [b, np, sq, hn] --> [sq, b, np, hn]
-        context_layer = context_layer.permute(2, 0, 1, 3).contiguous()
-        # [sq, b, np, hn] --> [sq, b, hp]
-        new_context_layer_shape = context_layer.size()[:-2] + \
-            (self.hidden_size_per_partition,)
-        context_layer = context_layer.view(*new_context_layer_shape)
-        output = self.out_proj(context_layer)
-        return output
-class VisualAttentionBlock(nn.Module):
-    def __init__(
-            self,
-            d_model: int,
-            n_head: int,
-            mlp_ratio: float = 4.0,
-            act_layer: Callable = nn.GELU,
-            norm_layer: Callable = nn.LayerNorm,
-            is_cross_attention: bool = False,
-    ):
-        super().__init__()
-        self.ln_1 = norm_layer(d_model)
-        if is_cross_attention:
-            self.ln_1_kv = norm_layer(d_model)
-        self.ln_2 = norm_layer(d_model)
-        mlp_width = int(d_model * mlp_ratio)
-        self.attn = VisualAttention(d_model, n_head)
-        self.mlp = nn.Sequential(OrderedDict([
-            ("c_fc", nn.Linear(d_model, mlp_width)),
-            ("gelu", act_layer()),
-            ("c_proj", nn.Linear(mlp_width, d_model))
-        ]))
-    def attention(
-            self,
-            q_x: torch.Tensor,
-            k_x: Optional[torch.Tensor] = None,
-            v_x: Optional[torch.Tensor] = None,
-            attn_mask: Optional[torch.Tensor] = None,
-    ):
-        k_x = k_x if k_x is not None else q_x
-        v_x = v_x if v_x is not None else q_x
-        attn_mask = attn_mask.to(q_x.dtype) if attn_mask is not None else None
-        return self.attn(q_x, k_x, v_x, attn_mask=attn_mask)
-    def forward(
-            self,
-            q_x: torch.Tensor,
-            k_x: Optional[torch.Tensor] = None,
-            v_x: Optional[torch.Tensor] = None,
-            attn_mask: Optional[torch.Tensor] = None,
-    ):
-        k_x = self.ln_1_kv(k_x) if hasattr(self, "ln_1_kv") and k_x is not None else None
-        v_x = self.ln_1_kv(v_x) if hasattr(self, "ln_1_kv") and v_x is not None else None
-        x = q_x + self.attention(q_x=self.ln_1(q_x), k_x=k_x, v_x=v_x, attn_mask=attn_mask)
-        x = x + self.mlp(self.ln_2(x))
-        return x
-class TransformerBlock(nn.Module):
-    def __init__(
-            self,
-            width: int,
-            layers: int,
-            heads: int,
-            mlp_ratio: float = 4.0,
-            act_layer: Callable = nn.GELU,
-            norm_layer: Callable = nn.LayerNorm,
-    ):
-        super().__init__()
-        self.width = width
-        self.layers = layers
-        self.resblocks = nn.ModuleList([
-            VisualAttentionBlock(
-                width, heads, mlp_ratio, act_layer=act_layer, norm_layer=norm_layer)
-            for _ in range(layers)
-        ])
-    def get_cast_dtype(self) -> torch.dtype:
-        return self.resblocks[0].mlp.c_fc.weight.dtype
-    def get_cast_device(self) -> torch.device:
-        return self.resblocks[0].mlp.c_fc.weight.device
-    def forward(self, x: torch.Tensor, attn_mask: Optional[torch.Tensor] = None):
-        for r in self.resblocks:
-            x = r(x, attn_mask=attn_mask)
-        return x
-class VisionTransformer(nn.Module):
-    def __init__(
-            self,
-            image_size: int,
-            patch_size: int,
-            width: int,
-            layers: int,
-            heads: int,
-            mlp_ratio: float,
-            n_queries: int = 256,
-            output_dim: int = 512,
-            **kwargs
-    ):
-        super().__init__()
-        image_height, image_width = self.image_size = (image_size, image_size)
-        patch_height, patch_width = self.patch_size = (patch_size, patch_size)
-        self.grid_size = (image_height // patch_height, image_width // patch_width)
-        self.output_dim = output_dim
-        mean = (0.48145466, 0.4578275, 0.40821073)
-        std = (0.26862954, 0.26130258, 0.27577711)
-        self.image_transform = transforms.Compose([
-            transforms.Resize(
-                (image_size, image_size),
-                interpolation=InterpolationMode.BICUBIC
-            ),
-            transforms.ToTensor(),
-            transforms.Normalize(mean=mean, std=std),
-        ])
-        self.conv1 = nn.Conv2d(in_channels=3, out_channels=width, kernel_size=patch_size, stride=patch_size, bias=False)
-        # class embeddings and positional embeddings
-        scale = width ** -0.5
-        self.positional_embedding = nn.Parameter(scale * torch.randn(256, width))
-        norm_layer = partial(nn.LayerNorm, eps=1e-6)
-        act_layer = nn.GELU
-        self.ln_pre = norm_layer(width)
-        self.transformer = TransformerBlock(
-            width,
-            layers,
-            heads,
-            mlp_ratio,
-            act_layer=act_layer,
-            norm_layer=norm_layer,
-        )
-        self.attn_pool = Resampler(
-            grid_size=int(math.sqrt(n_queries)),
-            embed_dim=output_dim,
-            num_heads=output_dim // 128,
-            kv_dim=width,
-            norm_layer=norm_layer,
-        )
-        self.ln_post = norm_layer(output_dim)
-        self.proj = nn.Parameter((output_dim** -0.5) * torch.randn(output_dim, output_dim))
-        self.data_path = os.environ.get("DATA", "./")
-        self.visualize = os.environ.get("VISUALIZE", False)
-    def forward(self, x: torch.Tensor):
-        x = x.to(
-            dtype=self.transformer.get_cast_dtype(),
-            device=self.transformer.get_cast_device(),
-        )
-        # to patches
-        x = self.conv1(x)  # shape = [*, width, grid, grid]
-        x = x.reshape(x.shape[0], x.shape[1], -1)  # shape = [*, width, grid ** 2]
-        x = x.permute(0, 2, 1)  # shape = [*, grid ** 2, width]
-        x = x + get_abs_pos(self.positional_embedding, x.size(1))
-        x = self.ln_pre(x)
-        x = x.permute(1, 0, 2)  # NLD -> LND
-        x = self.transformer(x)
-        x = x.permute(1, 0, 2)  # LND -> NLD
-        x = self.attn_pool(x)
-        x = self.ln_post(x)
-        x = x @ self.proj
-        return x
-    def encode(self, image_paths: List[str], masks:List[str] ):
-        images = []
-        for image_path, mask in zip(image_paths, masks):
-            if image_path.startswith("http://") or image_path.startswith("https://"):
-                image = Image.open(requests.get(image_path, stream=True).raw)
-            else:
-                image = Image.open(image_path)
-                image = ImageOps.exif_transpose(image)
-            image = image.convert("RGB")
-            if len(mask)>0:
-                print(len(mask))
-                seg = {'counts': mask, 'size': [image.size[1], image.size[0]]}
-            else:
-                seg = None
-            if seg is not None:
-                mask = self.annToMask(seg, image.size[1], image.size[0])
-                image = overlay_mask(np.array(image), mask)
-                image = Image.fromarray(image)
-            visulize = self.visualize
-            if visulize:
-                self.data_path =  "visualize"
-                os.makedirs(self.data_path, exist_ok=True)
-                torchshow.save(image, os.path.join(self.data_path, image_path.split("/")[-1]))
-            images.append(self.image_transform(image))
-        images = torch.stack(images, dim=0)
-        #torchshow.save(images)
-        return self(images)
-    def annToMask(self, mask_ann, h, w):
-        if mask_ann is None:
-            return np.zeros((h, w), dtype=np.uint8)
-        if isinstance(mask_ann, list):
-            rles = maskUtils.frPyObjects(mask_ann, h, w)
-            rle = maskUtils.merge(rles)
-        elif isinstance(mask_ann['counts'], list):
-            # uncompressed RLE
-            rle = maskUtils.frPyObjects(mask_ann, h, w)
-        else:
-            # rle
-            rle = mask_ann
-        mask = maskUtils.decode(rle)
-        return mask

zzzmmz/SegAgent-Model/zero_to_fp32.py DELETED Viewed

@@ -1,587 +0,0 @@
-#!/usr/bin/env python
-# Copyright (c) Microsoft Corporation.
-# SPDX-License-Identifier: Apache-2.0
-# DeepSpeed Team
-# This script extracts fp32 consolidated weights from a zero 1, 2 and 3 DeepSpeed checkpoints. It gets
-# copied into the top level checkpoint dir, so the user can easily do the conversion at any point in
-# the future. Once extracted, the weights don't require DeepSpeed and can be used in any
-# application.
-#
-# example: python zero_to_fp32.py . pytorch_model.bin
-import argparse
-import torch
-import glob
-import math
-import os
-import re
-from collections import OrderedDict
-from dataclasses import dataclass
-# while this script doesn't use deepspeed to recover data, since the checkpoints are pickled with
-# DeepSpeed data structures it has to be available in the current python environment.
-from deepspeed.utils import logger
-from deepspeed.checkpoint.constants import (DS_VERSION, OPTIMIZER_STATE_DICT, SINGLE_PARTITION_OF_FP32_GROUPS,
-                                            FP32_FLAT_GROUPS, ZERO_STAGE, PARTITION_COUNT, PARAM_SHAPES, BUFFER_NAMES,
-                                            FROZEN_PARAM_SHAPES, FROZEN_PARAM_FRAGMENTS)
-@dataclass
-class zero_model_state:
-    buffers: dict()
-    param_shapes: dict()
-    shared_params: list
-    ds_version: int
-    frozen_param_shapes: dict()
-    frozen_param_fragments: dict()
-debug = 0
-# load to cpu
-device = torch.device('cpu')
-def atoi(text):
-    return int(text) if text.isdigit() else text
-def natural_keys(text):
-    '''
-    alist.sort(key=natural_keys) sorts in human order
-    http://nedbatchelder.com/blog/200712/human_sorting.html
-    (See Toothy's implementation in the comments)
-    '''
-    return [atoi(c) for c in re.split(r'(\d+)', text)]
-def get_model_state_file(checkpoint_dir, zero_stage):
-    if not os.path.isdir(checkpoint_dir):
-        raise FileNotFoundError(f"Directory '{checkpoint_dir}' doesn't exist")
-    # there should be only one file
-    if zero_stage <= 2:
-        file = os.path.join(checkpoint_dir, "mp_rank_00_model_states.pt")
-    elif zero_stage == 3:
-        file = os.path.join(checkpoint_dir, "zero_pp_rank_0_mp_rank_00_model_states.pt")
-    if not os.path.exists(file):
-        raise FileNotFoundError(f"can't find model states file at '{file}'")
-    return file
-def get_checkpoint_files(checkpoint_dir, glob_pattern):
-    # XXX: need to test that this simple glob rule works for multi-node setup too
-    ckpt_files = sorted(glob.glob(os.path.join(checkpoint_dir, glob_pattern)), key=natural_keys)
-    if len(ckpt_files) == 0:
-        raise FileNotFoundError(f"can't find {glob_pattern} files in directory '{checkpoint_dir}'")
-    return ckpt_files
-def get_optim_files(checkpoint_dir):
-    return get_checkpoint_files(checkpoint_dir, "*_optim_states.pt")
-def get_model_state_files(checkpoint_dir):
-    return get_checkpoint_files(checkpoint_dir, "*_model_states.pt")
-def parse_model_states(files):
-    zero_model_states = []
-    for file in files:
-        state_dict = torch.load(file, map_location=device)
-        if BUFFER_NAMES not in state_dict:
-            raise ValueError(f"{file} is not a model state checkpoint")
-        buffer_names = state_dict[BUFFER_NAMES]
-        if debug:
-            print("Found buffers:", buffer_names)
-        # recover just the buffers while restoring them to fp32 if they were saved in fp16
-        buffers = {k: v.float() for k, v in state_dict["module"].items() if k in buffer_names}
-        param_shapes = state_dict[PARAM_SHAPES]
-        # collect parameters that are included in param_shapes
-        param_names = []
-        for s in param_shapes:
-            for name in s.keys():
-                param_names.append(name)
-        # update with frozen parameters
-        frozen_param_shapes = state_dict.get(FROZEN_PARAM_SHAPES, None)
-        if frozen_param_shapes is not None:
-            if debug:
-                print(f"Found frozen_param_shapes: {frozen_param_shapes}")
-            param_names += list(frozen_param_shapes.keys())
-        # handle shared params
-        shared_params = [[k, v] for k, v in state_dict["shared_params"].items()]
-        ds_version = state_dict.get(DS_VERSION, None)
-        frozen_param_fragments = state_dict.get(FROZEN_PARAM_FRAGMENTS, None)
-        z_model_state = zero_model_state(buffers=buffers,
-                                         param_shapes=param_shapes,
-                                         shared_params=shared_params,
-                                         ds_version=ds_version,
-                                         frozen_param_shapes=frozen_param_shapes,
-                                         frozen_param_fragments=frozen_param_fragments)
-        zero_model_states.append(z_model_state)
-    return zero_model_states
-def parse_optim_states(files, ds_checkpoint_dir):
-    total_files = len(files)
-    state_dicts = []
-    for f in files:
-        state_dict = torch.load(f, map_location=device)
-        # immediately discard the potentially huge 2 optimizer states as we only care for fp32 master weights
-        # and also handle the case where it was already removed by another helper script
-        state_dict["optimizer_state_dict"].pop("optimizer_state_dict", None)
-        state_dicts.append(state_dict)
-    if not ZERO_STAGE in state_dicts[0][OPTIMIZER_STATE_DICT]:
-        raise ValueError(f"{files[0]} is not a zero checkpoint")
-    zero_stage = state_dicts[0][OPTIMIZER_STATE_DICT][ZERO_STAGE]
-    world_size = state_dicts[0][OPTIMIZER_STATE_DICT][PARTITION_COUNT]
-    # For ZeRO-2 each param group can have different partition_count as data parallelism for expert
-    # parameters can be different from data parallelism for non-expert parameters. So we can just
-    # use the max of the partition_count to get the dp world_size.
-    if type(world_size) is list:
-        world_size = max(world_size)
-    if world_size != total_files:
-        raise ValueError(
-            f"Expected {world_size} of '*_optim_states.pt' under '{ds_checkpoint_dir}' but found {total_files} files. "
-            "Possibly due to an overwrite of an old checkpoint, or a checkpoint didn't get saved by one or more processes."
-        )
-    # the groups are named differently in each stage
-    if zero_stage <= 2:
-        fp32_groups_key = SINGLE_PARTITION_OF_FP32_GROUPS
-    elif zero_stage == 3:
-        fp32_groups_key = FP32_FLAT_GROUPS
-    else:
-        raise ValueError(f"unknown zero stage {zero_stage}")
-    if zero_stage <= 2:
-        fp32_flat_groups = [state_dicts[i][OPTIMIZER_STATE_DICT][fp32_groups_key] for i in range(len(state_dicts))]
-    elif zero_stage == 3:
-        # if there is more than one param group, there will be multiple flattened tensors - one
-        # flattened tensor per group - for simplicity merge them into a single tensor
-        #
-        # XXX: could make the script more memory efficient for when there are multiple groups - it
-        # will require matching the sub-lists of param_shapes for each param group flattened tensor
-        fp32_flat_groups = [
-            torch.cat(state_dicts[i][OPTIMIZER_STATE_DICT][fp32_groups_key], 0) for i in range(len(state_dicts))
-        ]
-    return zero_stage, world_size, fp32_flat_groups
-def _get_fp32_state_dict_from_zero_checkpoint(ds_checkpoint_dir):
-    """
-    Returns fp32 state_dict reconstructed from ds checkpoint
-    Args:
-        - ``ds_checkpoint_dir``: path to the deepspeed checkpoint folder (where the optimizer files are)
-    """
-    print(f"Processing zero checkpoint '{ds_checkpoint_dir}'")
-    optim_files = get_optim_files(ds_checkpoint_dir)
-    zero_stage, world_size, fp32_flat_groups = parse_optim_states(optim_files, ds_checkpoint_dir)
-    print(f"Detected checkpoint of type zero stage {zero_stage}, world_size: {world_size}")
-    model_files = get_model_state_files(ds_checkpoint_dir)
-    zero_model_states = parse_model_states(model_files)
-    print(f'Parsing checkpoint created by deepspeed=={zero_model_states[0].ds_version}')
-    if zero_stage <= 2:
-        return _get_fp32_state_dict_from_zero2_checkpoint(world_size, fp32_flat_groups, zero_model_states)
-    elif zero_stage == 3:
-        return _get_fp32_state_dict_from_zero3_checkpoint(world_size, fp32_flat_groups, zero_model_states)
-def _zero2_merge_frozen_params(state_dict, zero_model_states):
-    if zero_model_states[0].frozen_param_shapes is None or len(zero_model_states[0].frozen_param_shapes) == 0:
-        return
-    frozen_param_shapes = zero_model_states[0].frozen_param_shapes
-    frozen_param_fragments = zero_model_states[0].frozen_param_fragments
-    if debug:
-        num_elem = sum(s.numel() for s in frozen_param_shapes.values())
-        print(f'rank 0: {FROZEN_PARAM_SHAPES}.numel = {num_elem}')
-        wanted_params = len(frozen_param_shapes)
-        wanted_numel = sum(s.numel() for s in frozen_param_shapes.values())
-        avail_numel = sum([p.numel() for p in frozen_param_fragments.values()])
-        print(f'Frozen params: Have {avail_numel} numels to process.')
-        print(f'Frozen params: Need {wanted_numel} numels in {wanted_params} params')
-    total_params = 0
-    total_numel = 0
-    for name, shape in frozen_param_shapes.items():
-        total_params += 1
-        unpartitioned_numel = shape.numel()
-        total_numel += unpartitioned_numel
-        state_dict[name] = frozen_param_fragments[name]
-        if debug:
-            print(f"{name} full shape: {shape} unpartitioned numel {unpartitioned_numel} ")
-    print(f"Reconstructed Frozen fp32 state dict with {total_params} params {total_numel} elements")
-def _zero2_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states):
-    param_shapes = zero_model_states[0].param_shapes
-    # Reconstruction protocol:
-    #
-    # XXX: document this
-    if debug:
-        for i in range(world_size):
-            for j in range(len(fp32_flat_groups[0])):
-                print(f"{FP32_FLAT_GROUPS}[{i}][{j}].shape={fp32_flat_groups[i][j].shape}")
-    # XXX: memory usage doubles here (zero2)
-    num_param_groups = len(fp32_flat_groups[0])
-    merged_single_partition_of_fp32_groups = []
-    for i in range(num_param_groups):
-        merged_partitions = [sd[i] for sd in fp32_flat_groups]
-        full_single_fp32_vector = torch.cat(merged_partitions, 0)
-        merged_single_partition_of_fp32_groups.append(full_single_fp32_vector)
-    avail_numel = sum(
-        [full_single_fp32_vector.numel() for full_single_fp32_vector in merged_single_partition_of_fp32_groups])
-    if debug:
-        wanted_params = sum([len(shapes) for shapes in param_shapes])
-        wanted_numel = sum([sum(shape.numel() for shape in shapes.values()) for shapes in param_shapes])
-        # not asserting if there is a mismatch due to possible padding
-        print(f"Have {avail_numel} numels to process.")
-        print(f"Need {wanted_numel} numels in {wanted_params} params.")
-    # params
-    # XXX: for huge models that can't fit into the host's RAM we will have to recode this to support
-    # out-of-core computing solution
-    total_numel = 0
-    total_params = 0
-    for shapes, full_single_fp32_vector in zip(param_shapes, merged_single_partition_of_fp32_groups):
-        offset = 0
-        avail_numel = full_single_fp32_vector.numel()
-        for name, shape in shapes.items():
-            unpartitioned_numel = shape.numel()
-            total_numel += unpartitioned_numel
-            total_params += 1
-            if debug:
-                print(f"{name} full shape: {shape} unpartitioned numel {unpartitioned_numel} ")
-            state_dict[name] = full_single_fp32_vector.narrow(0, offset, unpartitioned_numel).view(shape)
-            offset += unpartitioned_numel
-        # Z2 started to align to 2*world_size to improve nccl performance. Therefore both offset and
-        # avail_numel can differ by anywhere between 0..2*world_size. Due to two unrelated complex
-        # paddings performed in the code it's almost impossible to predict the exact numbers w/o the
-        # live optimizer object, so we are checking that the numbers are within the right range
-        align_to = 2 * world_size
-        def zero2_align(x):
-            return align_to * math.ceil(x / align_to)
-        if debug:
-            print(f"original offset={offset}, avail_numel={avail_numel}")
-        offset = zero2_align(offset)
-        avail_numel = zero2_align(avail_numel)
-        if debug:
-            print(f"aligned  offset={offset}, avail_numel={avail_numel}")
-        # Sanity check
-        if offset != avail_numel:
-            raise ValueError(f"consumed {offset} numels out of {avail_numel} - something is wrong")
-    print(f"Reconstructed fp32 state dict with {total_params} params {total_numel} elements")
-def _get_fp32_state_dict_from_zero2_checkpoint(world_size, fp32_flat_groups, zero_model_states):
-    state_dict = OrderedDict()
-    # buffers
-    buffers = zero_model_states[0].buffers
-    state_dict.update(buffers)
-    if debug:
-        print(f"added {len(buffers)} buffers")
-    _zero2_merge_frozen_params(state_dict, zero_model_states)
-    _zero2_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states)
-    # recover shared parameters
-    for pair in zero_model_states[0].shared_params:
-        if pair[1] in state_dict:
-            state_dict[pair[0]] = state_dict[pair[1]]
-    return state_dict
-def zero3_partitioned_param_info(unpartitioned_numel, world_size):
-    remainder = unpartitioned_numel % world_size
-    padding_numel = (world_size - remainder) if remainder else 0
-    partitioned_numel = math.ceil(unpartitioned_numel / world_size)
-    return partitioned_numel, padding_numel
-def _zero3_merge_frozen_params(state_dict, world_size, zero_model_states):
-    if zero_model_states[0].frozen_param_shapes is None or len(zero_model_states[0].frozen_param_shapes) == 0:
-        return
-    if debug:
-        for i in range(world_size):
-            num_elem = sum(s.numel() for s in zero_model_states[i].frozen_param_fragments.values())
-            print(f'rank {i}: {FROZEN_PARAM_SHAPES}.numel = {num_elem}')
-        frozen_param_shapes = zero_model_states[0].frozen_param_shapes
-        wanted_params = len(frozen_param_shapes)
-        wanted_numel = sum(s.numel() for s in frozen_param_shapes.values())
-        avail_numel = sum([p.numel() for p in zero_model_states[0].frozen_param_fragments.values()]) * world_size
-        print(f'Frozen params: Have {avail_numel} numels to process.')
-        print(f'Frozen params: Need {wanted_numel} numels in {wanted_params} params')
-    total_params = 0
-    total_numel = 0
-    for name, shape in zero_model_states[0].frozen_param_shapes.items():
-        total_params += 1
-        unpartitioned_numel = shape.numel()
-        total_numel += unpartitioned_numel
-        param_frags = tuple(model_state.frozen_param_fragments[name] for model_state in zero_model_states)
-        state_dict[name] = torch.cat(param_frags, 0).narrow(0, 0, unpartitioned_numel).view(shape)
-        partitioned_numel, partitioned_padding_numel = zero3_partitioned_param_info(unpartitioned_numel, world_size)
-        if debug:
-            print(
-                f"Frozen params: {total_params} {name} full shape: {shape} partition0 numel={partitioned_numel} partitioned_padding_numel={partitioned_padding_numel}"
-            )
-    print(f"Reconstructed Frozen fp32 state dict with {total_params} params {total_numel} elements")
-def _zero3_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states):
-    param_shapes = zero_model_states[0].param_shapes
-    avail_numel = fp32_flat_groups[0].numel() * world_size
-    # Reconstruction protocol: For zero3 we need to zip the partitions together at boundary of each
-    # param, re-consolidating each param, while dealing with padding if any
-    # merge list of dicts, preserving order
-    param_shapes = {k: v for d in param_shapes for k, v in d.items()}
-    if debug:
-        for i in range(world_size):
-            print(f"{FP32_FLAT_GROUPS}[{i}].shape={fp32_flat_groups[i].shape}")
-        wanted_params = len(param_shapes)
-        wanted_numel = sum(shape.numel() for shape in param_shapes.values())
-        # not asserting if there is a mismatch due to possible padding
-        avail_numel = fp32_flat_groups[0].numel() * world_size
-        print(f"Trainable params: Have {avail_numel} numels to process.")
-        print(f"Trainable params: Need {wanted_numel} numels in {wanted_params} params.")
-    # params
-    # XXX: for huge models that can't fit into the host's RAM we will have to recode this to support
-    # out-of-core computing solution
-    offset = 0
-    total_numel = 0
-    total_params = 0
-    for name, shape in param_shapes.items():
-        unpartitioned_numel = shape.numel()
-        total_numel += unpartitioned_numel
-        total_params += 1
-        partitioned_numel, partitioned_padding_numel = zero3_partitioned_param_info(unpartitioned_numel, world_size)
-        if debug:
-            print(
-                f"Trainable params: {total_params} {name} full shape: {shape} partition0 numel={partitioned_numel} partitioned_padding_numel={partitioned_padding_numel}"
-            )
-        # XXX: memory usage doubles here
-        state_dict[name] = torch.cat(
-            tuple(fp32_flat_groups[i].narrow(0, offset, partitioned_numel) for i in range(world_size)),
-            0).narrow(0, 0, unpartitioned_numel).view(shape)
-        offset += partitioned_numel
-    offset *= world_size
-    # Sanity check
-    if offset != avail_numel:
-        raise ValueError(f"consumed {offset} numels out of {avail_numel} - something is wrong")
-    print(f"Reconstructed Trainable fp32 state dict with {total_params} params {total_numel} elements")
-def _get_fp32_state_dict_from_zero3_checkpoint(world_size, fp32_flat_groups, zero_model_states):
-    state_dict = OrderedDict()
-    # buffers
-    buffers = zero_model_states[0].buffers
-    state_dict.update(buffers)
-    if debug:
-        print(f"added {len(buffers)} buffers")
-    _zero3_merge_frozen_params(state_dict, world_size, zero_model_states)
-    _zero3_merge_trainable_params(state_dict, world_size, fp32_flat_groups, zero_model_states)
-    # recover shared parameters
-    for pair in zero_model_states[0].shared_params:
-        if pair[1] in state_dict:
-            state_dict[pair[0]] = state_dict[pair[1]]
-    return state_dict
-def get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir, tag=None):
-    """
-    Convert ZeRO 2 or 3 checkpoint into a single fp32 consolidated state_dict that can be loaded with
-    ``load_state_dict()`` and used for training without DeepSpeed or shared with others, for example
-    via a model hub.
-    Args:
-        - ``checkpoint_dir``: path to the desired checkpoint folder
-        - ``tag``: checkpoint tag used as a unique identifier for checkpoint. If not provided will attempt to load tag in 'latest' file. e.g., ``global_step14``
-    Returns:
-        - pytorch ``state_dict``
-    Note: this approach may not work if your application doesn't have sufficient free CPU memory and
-    you may need to use the offline approach using the ``zero_to_fp32.py`` script that is saved with
-    the checkpoint.
-    A typical usage might be ::
-        from deepspeed.utils.zero_to_fp32 import get_fp32_state_dict_from_zero_checkpoint
-        # do the training and checkpoint saving
-        state_dict = get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir) # already on cpu
-        model = model.cpu() # move to cpu
-        model.load_state_dict(state_dict)
-        # submit to model hub or save the model to share with others
-    In this example the ``model`` will no longer be usable in the deepspeed context of the same
-    application. i.e. you will need to re-initialize the deepspeed engine, since
-    ``model.load_state_dict(state_dict)`` will remove all the deepspeed magic from it.
-    If you want it all done for you, use ``load_state_dict_from_zero_checkpoint`` instead.
-    """
-    if tag is None:
-        latest_path = os.path.join(checkpoint_dir, 'latest')
-        if os.path.isfile(latest_path):
-            with open(latest_path, 'r') as fd:
-                tag = fd.read().strip()
-        else:
-            raise ValueError(f"Unable to find 'latest' file at {latest_path}")
-    ds_checkpoint_dir = os.path.join(checkpoint_dir, tag)
-    if not os.path.isdir(ds_checkpoint_dir):
-        raise FileNotFoundError(f"Directory '{ds_checkpoint_dir}' doesn't exist")
-    return _get_fp32_state_dict_from_zero_checkpoint(ds_checkpoint_dir)
-def convert_zero_checkpoint_to_fp32_state_dict(checkpoint_dir, output_file, tag=None):
-    """
-    Convert ZeRO 2 or 3 checkpoint into a single fp32 consolidated ``state_dict`` file that can be
-    loaded with ``torch.load(file)`` + ``load_state_dict()`` and used for training without DeepSpeed.
-    Args:
-        - ``checkpoint_dir``: path to the desired checkpoint folder. (one that contains the tag-folder, like ``global_step14``)
-        - ``output_file``: path to the pytorch fp32 state_dict output file (e.g. path/pytorch_model.bin)
-        - ``tag``: checkpoint tag used as a unique identifier for checkpoint. If not provided will attempt to load tag in the file named ``latest`` in the checkpoint folder, e.g., ``global_step14``
-    """
-    state_dict = get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir, tag)
-    print(f"Saving fp32 state dict to {output_file}")
-    torch.save(state_dict, output_file)
-def load_state_dict_from_zero_checkpoint(model, checkpoint_dir, tag=None):
-    """
-    1. Put the provided model to cpu
-    2. Convert ZeRO 2 or 3 checkpoint into a single fp32 consolidated ``state_dict``
-    3. Load it into the provided model
-    Args:
-        - ``model``: the model object to update
-        - ``checkpoint_dir``: path to the desired checkpoint folder. (one that contains the tag-folder, like ``global_step14``)
-        - ``tag``: checkpoint tag used as a unique identifier for checkpoint. If not provided will attempt to load tag in the file named ``latest`` in the checkpoint folder, e.g., ``global_step14``
-    Returns:
-        - ``model`: modified model
-    Make sure you have plenty of CPU memory available before you call this function. If you don't
-    have enough use the ``zero_to_fp32.py`` utility to do the conversion. You will find it
-    conveniently placed for you in the checkpoint folder.
-    A typical usage might be ::
-        from deepspeed.utils.zero_to_fp32 import load_state_dict_from_zero_checkpoint
-        model = load_state_dict_from_zero_checkpoint(trainer.model, checkpoint_dir)
-        # submit to model hub or save the model to share with others
-    Note, that once this was run, the ``model`` will no longer be usable in the deepspeed context
-    of the same application. i.e. you will need to re-initialize the deepspeed engine, since
-    ``model.load_state_dict(state_dict)`` will remove all the deepspeed magic from it.
-    """
-    logger.info(f"Extracting fp32 weights")
-    state_dict = get_fp32_state_dict_from_zero_checkpoint(checkpoint_dir, tag)
-    logger.info(f"Overwriting model with fp32 weights")
-    model = model.cpu()
-    model.load_state_dict(state_dict, strict=False)
-    return model
-if __name__ == "__main__":
-    parser = argparse.ArgumentParser()
-    parser.add_argument("checkpoint_dir",
-                        type=str,
-                        help="path to the desired checkpoint folder, e.g., path/checkpoint-12")
-    parser.add_argument(
-        "output_file",
-        type=str,
-        help="path to the pytorch fp32 state_dict output file (e.g. path/checkpoint-12/pytorch_model.bin)")
-    parser.add_argument("-t",
-                        "--tag",
-                        type=str,
-                        default=None,
-                        help="checkpoint tag used as a unique identifier for checkpoint. e.g., global_step1")
-    parser.add_argument("-d", "--debug", action='store_true', help="enable debug")
-    args = parser.parse_args()
-    debug = args.debug
-    convert_zero_checkpoint_to_fp32_state_dict(args.checkpoint_dir, args.output_file, tag=args.tag)