InternImage 是由上海人工智能实验室、清华大学及其他机构的研究人员研发的前沿视觉基础模型。与基于 Transformer 的模型不同,InternImage 采用 DCNv3 作为核心算子。该方案为模型提供了目标检测、分割等下游任务所需的动态且高效感受野,同时支持自适应空间聚合。
| Hugging Face 名称 | 模型名称 | 预训练 | 分辨率 | 参数量 |
|---|---|---|---|---|
| internimage_l_22k_384 | InternImage-L | IN-22K | 384x384 | 223M |
| internimage_xl_22k_384 | InternImage-XL | IN-22K | 384x384 | 335M |
| internimage_h_jointto22k_384 | InternImage-H | Joint 427M -> IN-22K | 384x384 | 1.08B |
| internimage_g_jointto22k_384 | InternImage-G | Joint 427M -> IN-22K | 384x384 | 3B |
| Hugging Face 名称 | 模型名称 | 预训练 | 分辨率 | acc@1 | 参数量 | FLOPs |
|---|---|---|---|---|---|---|
| internimage_t_1k_224 | InternImage-T | IN-1K | 224x224 | 83.5 | 30M | 5G |
| internimage_s_1k_224 | InternImage-S | IN-1K | 224x224 | 84.2 | 50M | 8G |
| internimage_b_1k_224 | InternImage-B | IN-1K | 224x224 | 84.9 | 97M | 16G |
| internimage_l_22kto1k_384 | InternImage-L | IN-22K | 384x384 | 87.7 | 223M | 108G |
| internimage_xl_22kto1k_384 | InternImage-XL | IN-22K | 384x384 | 88.0 | 335M | 163G |
| internimage_h_22kto1k_640 | InternImage-H | Joint 427M -> IN-22K | 640x640 | 89.6 | 1.08B | 1478G |
| internimage_g_22kto1k_512 | InternImage-G | Joint 427M -> IN-22K | 512x512 | 90.1 | 3B | 2700G |
如果未安装 DCNv3 的 CUDA 版本,InternImage 将自动回退至 PyTorch 实现。但是,CUDA 实现可以显著降低 GPU 显存占用,并提升推理效率。
安装教程:
打开终端并运行:
git clone https://github.com/OpenGVLab/InternImage.git
cd InternImage/classification/ops_dcnv3确保有可用于编译的 GPU,然后运行:
sh make.sh这将编译 DCNv3 的 CUDA 版本。安装完成后,InternImage 将自动利用经过优化的 CUDA 实现,以获得更好的性能。
以下是 InternImage 在 Transformers 框架中的两个使用示例:
import torch
from PIL import Image
from transformers import AutoModel, CLIPImageProcessor
# Replace 'model_name' with the appropriate model identifier
model_name = "OpenGVLab/internimage_t_1k_224" # example model
# Prepare the image
image_path = 'img.png'
image_processor = CLIPImageProcessor.from_pretrained(model_name)
image = Image.open(image_path)
image = image_processor(images=image, return_tensors='pt').pixel_values
print('image shape:', image.shape)
# Load the model as a backbone
model = AutoModel.from_pretrained(model_name, trust_remote_code=True)
# 'hidden_states' contains the outputs from the 4 stages of the InternImage backbone
hidden_states = model(image).hidden_statesimport torch
from PIL import Image
from transformers import AutoModelForImageClassification, CLIPImageProcessor
# Replace 'model_name' with the appropriate model identifier
model_name = "OpenGVLab/internimage_t_1k_224" # example model
# Prepare the image
image_path = 'img.png'
image_processor = CLIPImageProcessor.from_pretrained(model_name)
image = Image.open(image_path)
image = image_processor(images=image, return_tensors='pt').pixel_values
print('image shape:', image.shape)
# Load the model as an image classifier
model = AutoModelForImageClassification.from_pretrained(model_name, trust_remote_code=True)
logits = model(image).logits
label = torch.argmax(logits, dim=1)
print("Predicted label:", label.item())若本研究对您的工作有所帮助,敬请引用以下 BibTeX 条目。
@inproceedings{wang2023internimage,
title={Internimage: Exploring large-scale vision foundation models with deformable convolutions},
author={Wang, Wenhai and Dai, Jifeng and Chen, Zhe and Huang, Zhenhang and Li, Zhiqi and Zhu, Xizhou and Hu, Xiaowei and Lu, Tong and Lu, Lewei and Li, Hongsheng and others},
booktitle={Proceedings of the IEEE/CVF conference on computer vision and pattern recognition},
pages={14408--14419},
year={2023}
}