Exploring Deformable Convolution-Based Large-Scale Visual Foundation Models
Core components of this framework include a deformable convolution-based genarel visual backbone InternImage, the M3I-Pretraining self-supervised/semi-supervised training algorithm, the Uni-Perceiver universal decoder family, and the BEVFormer autonomous driving perception encoder suite.
Key Highlights
A 3-billion-parameter InternImage-G as th ...
Posted on Sat, 05 Sep 2026 16:16:53 +0000 by l3asturd