Exploring Deformable Convolution-Based Large-Scale Visual Foundation Models

Core components of this framework include a deformable convolution-based genarel visual backbone InternImage, the M3I-Pretraining self-supervised/semi-supervised training algorithm, the Uni-Perceiver universal decoder family, and the BEVFormer autonomous driving perception encoder suite. Key Highlights A 3-billion-parameter InternImage-G as th ...

Posted on Sat, 05 Sep 2026 16:16:53 +0000 by l3asturd