Non-linear transformations are fundamental to neural network architectures, enabling models to capture complex data distributions. An effective activation mapping typically satisfies three criteria: it must be non-linear yet predominantly differentiable to support backpropagation; computationally lightweight to minimize training overhead; and its gradient magnitude should remain bounded to prevent optimization instability.
Logistic Sigmoid
The sigmoid function compresses arbitrary real numbers into the $(0, 1)$ interval. Its mathematical definition is:
σ(x) = 1 / (1 + e<sup>-x</sup>)
The derivative exhibits a useful closed-form property: σ'(x) = σ(x) * (1 - σ(x)). Near the origin, the curve behaves approximately linearly, while larrge positive or negative inputs saturate toward 1 or 0, respectively. This saturation behavior leads to several practical drawbacks:
- Gradients diminish rapidly in saturated regions, slowing convergence in deep layers.
- Outputs are strictly positive, causing shifted mean activations in subsequent layers (bias shift) and complicating optimization dynamics.
- Exponential computations introduce higher computational cost compared to linear operations.
import torch
import matplotlib.pyplot as plt
import numpy as np
def plot_sigmoid(domain):
fig, ax = plt.subplots(figsize=(8, 4))
input_tensor = torch.from_numpy(domain).float()
activations = torch.sigmoid(input_tensor)
ax.plot(domain, activations.numpy(), linewidth=2.5)
ax.set(title="Logistic Sigmoid Curve", xlabel="Input Domain", ylabel="Output Range")
ax.grid(alpha=0.3)
plt.tight_layout()
plt.show()
values = np.linspace(-5.0, 5.0, 400)
plot_sigmoid(values)
Hyperbolic Tangent (Tanh)
The tanh function shifts and scales the sigmoid curve to span $(-1, 1)$. This zero-centered property mitigates the bias shift issue, improving gradient flow during backpropagation. It relates directly to the logistic function via tanh(x) = 2 * σ(2x) - 1.
import torch
import matplotlib.pyplot as plt
import numpy as np
def compare_centered_activations(x_vals):
x_t = torch.tensor(x_vals, dtype=torch.float32)
sig_out = torch.sigmoid(x_t).numpy()
tanh_out = torch.tanh(x_t).numpy()
fig, ax = plt.subplots()
ax.plot(x_vals, sig_out, label='Sigmoid', linestyle='--')
ax.plot(x_vals, tanh_out, label='Tanh', color='tab:orange')
ax.axhline(0, color='black', linewidth=0.5)
ax.set(xlabel="x", ylabel="f(x)", title="Activation Centering Comparison")
ax.legend(loc='upper left')
ax.grid(True, linestyle=':')
plt.show()
compare_centered_activations(np.linspace(-6, 6, 500))
Rectified Linear Unit (ReLU)
Defined as f(x) = max(0, x), ReLU has become the default choice for hidden layers. Its advantages include computational efficiency, biological plausibility regarding neural sparsity, and gradient preservation for positive inputs, which alleviates vanishing gradients. However, two limitations persist: outputs remain non-zero-centered, and neurons receiving consistently negative pre-activations receive zero gradients permanently, leading to the "dying ReLU" phenomenon.
To address these shortcomings, several extensions were developed:
- Leaky ReLU: Assigns a small constant slope
α(e.g., 0.01) to negative inputs:f(x) = max(αx, x). - Parametric ReLU (PReLU): Treats the negative slope as a learnable tensor, allowing the model to adapt the gradient flow during training.
- Exponential Linear Unit (ELU): Uses
f(x) = xforx > 0andα(e<sup>x</sup> - 1)forx ≤ 0. This pushes mean activations closer to zero, accelerating convergence. - SoftPlus: A smooth, everywhere-differentiable approximation:
ln(1 + e<sup>x</sup>). It shares ReLU's one-sided saturation but lacks sparsity.
import torch
import matplotlib.pyplot as plt
import numpy as np
def visualize_relu_variants():
domain = np.linspace(-4.0, 4.0, 300)
x = torch.from_numpy(domain).float()
fig, ax = plt.subplots()
ax.plot(domain, torch.relu(x).numpy(), label='Standard ReLU')
ax.plot(domain, torch.nn.functional.leaky_relu(x, negative_slope=0.1).numpy(),
label='Leaky ReLU', linestyle='-.')
ax.plot(domain, torch.nn.functional.elu(x).numpy(), label='ELU', linewidth=2)
ax.plot(domain, torch.nn.functional.softplus(x).numpy(), label='Softplus', linestyle='--')
ax.set(title="Rectified Activation Variants", xlabel="Input", ylabel="Activation")
ax.legend()
ax.grid(alpha=0.4)
plt.show()
visualize_relu_variants()
Maxout
Instead of relying on a fixed mathematical formula, Maxout learns the activation shape. Given an input vector x and k affine transformations, it computes z<sub>i</sub> = W<sub>i</sub>x + b<sub>i</sub> and outputs max(z<sub>1</sub>, ..., z<sub>k</sub>). This approach can approximate any convex piecewise-linear function but increases parameter count linearly with k.
Mish
Mish combines self-regularization with smoothness: Mish(x) = x * tanh(Softplus(x)). Its continuous gradient and non-monotonic behavior allow it to preserve information across deeper layers while maintaining bounded outputs for negative inputs.
import torch
import matplotlib.pyplot as plt
import numpy as np
def render_mish():
x_arr = np.linspace(-6.0, 6.0, 400)
x_t = torch.from_numpy(x_arr).float()
y = x_t * torch.tanh(torch.nn.functional.softplus(x_t)).numpy()
plt.figure()
plt.plot(x_arr, y, color='purple', linewidth=2.5)
plt.title("Mish Activation Profile")
plt.xlabel("Input Space")
plt.ylabel("Transformed Value")
plt.grid(linestyle=':')
plt.show()
render_mish()
Swish / SiLU
Swish introduces a gated structure: f(x) = x * σ(βx). The sigmoid component acts as a continuous gate that modulates how much of the original input x passes through. The hyperparameter β controls the transition: β=0 yields a linear half-identity, β=1 produces the standard Swish curve, and large β values approximate ReLU.
import torch
import matplotlib.pyplot as plt
import numpy as np
def render_swish_family():
x_vals = np.linspace(-5.0, 5.0, 350)
x_t = torch.tensor(x_vals, dtype=torch.float32)
beta_val = 1.0
gated_output = x_t * torch.sigmoid(beta_val * x_t).numpy()
plt.plot(x_vals, gated_output, label='Swish (beta=1)')
plt.axhline(0, color='gray', linewidth=0.5)
plt.title("Gated Activation (SiLU) Curve")
plt.xlabel("Input Domain")
plt.ylabel("Activation Output")
plt.legend()
plt.grid(alpha=0.3)
plt.show()
render_swish_family()
Gaussian Error Linear Unit (GELU)
GELU probabilistically weights inputs using the cumulative distribution function (CDF) of a standard normal distribution: GELU(x) = x * Φ(x). In practice, its approximated via a computationally efficient cubic polynomial: 0.5x * (1 + tanh(√(2/π) * (x + 0.044715x<sup>3</sup>))). The function exhibits smooth, non-monotonic characteristics that improve training stability in transformer blocks.
import torch
import matplotlib.pyplot as plt
import numpy as np
def plot_gelu_approx():
inputs = np.linspace(-4.0, 4.0, 300)
x_t = torch.from_numpy(inputs).float()
# Using PyTorch's built-in GELU for accuracy
gelu_out = torch.nn.functional.gelu(x_t).numpy()
fig, ax = plt.subplots(figsize=(7, 3.5))
ax.plot(inputs, gelu_out, linewidth=2.5)
ax.set(xlabel="Input", ylabel="GELU(x)", title="Probabilistic Gating Curve")
ax.grid(linestyle='--', alpha=0.4)
plt.tight_layout()
plt.show()
plot_gelu_approx()
Gated Linear Architectures (GLU & Variants)
GLU partitions the input into two branches, applying a sigmoid gate to one: GLU(a, b) = a ⊗ σ(b). This mechanism enables selective information routing, which is highly effective for sequence modeling. Recent large language models have adopted refined variants that replace the linear branch with modern non-linearities:
- ReGLU:
max(0, xW + b) ⊗ (xV + c). Favored for its training stability in certain embedding architectures. - SwiGLU:
Swish(xW + b) ⊗ (xV + c). Widely deployed in LLaMA and Baichuan families. - GEGLU:
GELU(xW + b) ⊗ (xV + c). Adopted in GLM-130B for its smooth gradient properties.