GhostNet introduces a novel approach to neural network lightweighting by rethinking how feature maps are generated—replacing expensive convolutions with structured, low-cost operations. Its evolution from V1 to V2 integrates spatial attention in a hardware-friendly manner, enabling stronger representational power without compromising inference speed on edge devices.
GhostNet V1: Efficient Feature Generation via Ghost Modules
The core innovation of GhostNet V1 lies in the Ghost Module, a plug-and-play replacement for standard convolutional layers. Instead of computing all output channels directly via costly 1×1 or depthwise convolutions, it splits the process into two stages:
- A primary convolution generates m intrinsic feature maps (where m < n, and n is the target number of channels).
- A set of inexpensive linear transformations—implemented as depthwise convolutions—generates s−1 "ghost" maps per intrinsic map, yielding n = m × s total channels.
This design exploits redundancy in intermediate CNN features: many channels are highly correlated, so generating them via cheap operations preserves expressiveness while drastically cutting FLOPs and parameters. The theoretical speedup ratio approximates s, assuming kernel sizes and channel counts follow typical mobile-CNN configurations.
The module’s PyTorch implementation reflects this decomposition:
class GhostModule(nn.Module):
def __init__(self, in_c, out_c, s=2, k=1, stride=1, act=True):
super().__init__()
self.intrinsic = out_c // s
self.ghost = self.intrinsic * (s - 1)
self.primary = nn.Sequential(
nn.Conv2d(in_c, self.intrinsic, k, stride, k//2, bias=False),
nn.BatchNorm2d(self.intrinsic),
nn.ReLU(inplace=True) if act else nn.Identity()
)
self.cheap = nn.Sequential(
nn.Conv2d(self.intrinsic, self.ghost, 3, 1, 1, groups=self.intrinsic, bias=False),
nn.BatchNorm2d(self.ghost),
nn.ReLU(inplace=True) if act else nn.Identity()
)
def forward(self, x):
x1 = self.primary(x)
x2 = self.cheap(x1)
return torch.cat([x1, x2], dim=1)
Ghost Bottleneck: Residual Building Block
The Ghost bottleneck adapts the MobileNetV2-style inverted residual structure, substituting both pointwise expansions and projections with Ghost Modules. It consists of:
- An expansion Ghost Module (increasing channel count),
- An optional depthwise convolution (for downsampling when
stride=2), - An optional Squeeze-and-Excite (SE) block,
- A projection Ghost Module (restoring channel count, no activation after).
Residual connections mirror those in ResNet—identity mapping for stride=1, and depthwise + pointwise projection for stride=2. This maintains spatial alignment while minimizing parameter overhead.
class GhostBottleneck(nn.Module):
def __init__(self, in_c, mid_c, out_c, k=3, stride=1, se_ratio=0.):
super().__init__()
self.use_se = se_ratio > 0.
self.stride = stride
self.expand = GhostModule(in_c, mid_c, act=True)
self.dwconv = nn.Conv2d(mid_c, mid_c, k, stride, k//2, groups=mid_c, bias=False) if stride > 1 else nn.Identity()
self.bn = nn.BatchNorm2d(mid_c) if stride > 1 else nn.Identity()
self.se = SqueezeExcite(mid_c, se_ratio) if self.use_se else nn.Identity()
self.project = GhostModule(mid_c, out_c, act=False)
self.shortcut = nn.Sequential(
nn.Conv2d(in_c, out_c, 1, stride=stride, bias=False) if stride > 1 else nn.Identity(),
nn.BatchNorm2d(out_c) if stride > 1 else nn.Identity(),
nn.Conv2d(in_c, out_c, k, stride=stride, padding=k//2, groups=in_c, bias=False) if stride > 1 else nn.Identity(),
nn.BatchNorm2d(out_c) if stride > 1 else nn.Identity()
) if (in_c != out_c or stride > 1) else nn.Identity()
def forward(self, x):
residual = self.shortcut(x)
x = self.expand(x)
x = self.dwconv(x)
x = self.bn(x)
x = self.se(x)
x = self.project(x)
return x + residual
GhostNet V2: Enhancing Spatial Awareness with Decoupled FC Attention
While GhostNet V1 excels in parameter efficiency, its local receptive fields limit long-range dependency modeling. GhostNet V2 addresses this with the Decoupled Fully Connected (DFC) attention mechanism—a lightweight alternative to self-attention that retains global context while remaining deployable on mobile accelerators.
Instead of computing dense pairwise interactions across all H×W tokens (O(H²W²)), DFC decomposes attention into two sequential 1D passes:
- A horizontal pass: applies a learnable weight matrix across width dimension → O(HW·KW),
- A vertical pass: applies another across height dimension → O(HW·KH).
In practice, these are realized using two depthwise convolutions: one with kernel 1×KH, the other with KW×1. This avoids expensive reshaping and transpose ops, making DFC fully compatible with ONNX and TFLite runtimes.
Crucially, DFC is applied only to the intermediate expansion layer inside the bottleneck—not the input or output—balancing expressiveness and cost. An attention gate modulates the Ghost Module’s output via element-wise multiplication after upsampling the attention map to match spatial dimensions.
class GhostModuleV2(nn.Module):
def __init__(self, in_c, out_c, s=2, k=1, dw_k=3, stride=1, act=True, attn=False):
super().__init__()
self.attn = attn
self.intrinsic = out_c // s
self.ghost = self.intrinsic * (s - 1)
self.primary = nn.Sequential(
nn.Conv2d(in_c, self.intrinsic, k, stride, k//2, bias=False),
nn.BatchNorm2d(self.intrinsic),
nn.ReLU(inplace=True) if act else nn.Identity()
)
self.cheap = nn.Sequential(
nn.Conv2d(self.intrinsic, self.ghost, dw_k, 1, dw_k//2, groups=self.intrinsic, bias=False),
nn.BatchNorm2d(self.ghost),
nn.ReLU(inplace=True) if act else nn.Identity()
)
if self.attn:
self.attn_path = nn.Sequential(
nn.AvgPool2d(2, 2),
nn.Conv2d(in_c, out_c, k, stride, k//2, bias=False),
nn.BatchNorm2d(out_c),
nn.Conv2d(out_c, out_c, (1, 5), padding=(0, 2), groups=out_c, bias=False),
nn.BatchNorm2d(out_c),
nn.Conv2d(out_c, out_c, (5, 1), padding=(2, 0), groups=out_c, bias=False),
nn.BatchNorm2d(out_c)
)
self.gate = nn.Sigmoid()
def forward(self, x):
x1 = self.primary(x)
x2 = self.cheap(x1)
out = torch.cat([x1, x2], dim=1)[:, :x.size(1)*2, :, :]
if self.attn:
attn_map = self.attn_path(x)
attn_map = F.interpolate(attn_map, size=out.shape[-2:], mode='nearest')
out = out * self.gate(attn_map)
return out
Architectural Integration
In GhostNet V2, attention-augmented Ghost Modules are selectively deployed—typically starting from deeper layers where contextual reasoning matters most. Earlier bottlenecks retain the original Ghost Module to preserve latency budgets. This staged enhancement strategy achivees measurable accuracy gains (e.g., +0.8% Top-1 on ImageNet at similar FLOPs) without increasing peak memory footprint or degrading throughput on ARM or NPU backends.