In video-based computer vision systems—such as those used for traffic monitoring—it's critical that the camera remains fixed relative to the scene. Even minor shifts or vibrations can misalign predefined detection zones (e.g., lanes or direction markers), leading to false positives or missed detections. To maintain system reliability, an automatic mechanism is needed to detect when the camera has moved and suspend analysis until it’s realigned.
This article presents a lightweight solution using feature point matching between consecutive video frames. By comparing the spatial displacement of matched keypoints, the system classifies camera status into four categories: no shift, minor jitter, significant shift, or complete misalignment. The entire implementation requires fewer than 200 lines of Python code.
Feature Detection and Matching
Images contain distinctive regions known as feature points, which remain identifiable despite changes in scale, rotation, or lighting. Common algorithms include SIFT, SURF, ORB, and FAST. OpenCV provides efficient implementations of these methods.
- Corners: Points at intersections of edges (e.g., building corners).
- Blobs: Distinctive textured or intensity-based regions.
Algorithms like SIFT offer invariance to rotation and scale, making them ideal for comparing frames captured under slightly varying conditions.
Detecting Camera Movement
If a camera is stationary, corresponding feature points in two successive frames should occupy nearly identical pixel coordinates (modulo minor motion from scene dynamics). A significant average displacement across multiple matched points indicates physical camera movement.
Key considerations:
- Use a minimum number of matches (e.g., 10) to avoid noise-driven decisions.
- Exclude dynamic overlays (e.g., timestamps or logos) that may falsely appear stable.
- Compute the mean Euclidean distance between matched keyopint positions.
Based on this average displacement, the system assigns one of four states:
- Complete misalignment: Too few matches (
< MIN_MATCH_COUNT). - Significant shift: Average displacement ≥ 10 pixels.
- Minor jitter: Displacement between 4 and 10 pixels.
- No shift: Displacement < 4 pixels.
Implementation
The core logic uses SIFT for feature extraction and FLANN-based matching for efficiency. Frames are downsampled to 30% of original size to speed up processing without sacrificing robustness.
import numpy as np
import cv2
MIN_MATCH_THRESHOLD = 10
IDEAL_MAX_DISPLACEMENT = 4.0
ACCEPTABLE_MAX_DISPLACEMENT = 10.0
def extract_and_match(frame_a, frame_b):
gray_a = cv2.cvtColor(frame_a, cv2.COLOR_BGR2GRAY)
gray_b = cv2.cvtColor(frame_b, cv2.COLOR_BGR2GRAY)
# Downscale for performance
h, w = gray_a.shape[:2]
gray_a = cv2.resize(gray_a, (int(w * 0.3), int(h * 0.3)))
gray_b = cv2.resize(gray_b, (int(w * 0.3), int(h * 0.3)))
detector = cv2.SIFT_create()
kp_a, desc_a = detector.detectAndCompute(gray_a, None)
kp_b, desc_b = detector.detectAndCompute(gray_b, None)
if desc_a is None or desc_b is None:
return 0 # No features
index_params = dict(algorithm=0, trees=5)
search_params = dict(checks=50)
matcher = cv2.FlannBasedMatcher(index_params, search_params)
raw_matches = matcher.knnMatch(desc_a, desc_b, k=2)
good_matches = []
for m, n in raw_matches:
if m.distance < 0.7 * n.distance:
good_matches.append(m)
if len(good_matches) < MIN_MATCH_THRESHOLD:
return 0 # Scene change or total misalignment
total_dist = 0.0
for match in good_matches:
pt_a = kp_a[match.queryIdx].pt
pt_b = kp_b[match.trainIdx].pt
dx = pt_a[0] - pt_b[0]
dy = pt_a[1] - pt_b[1]
total_dist += np.sqrt(dx*dx + dy*dy)
avg_displacement = total_dist / len(good_matches)
if avg_displacement < IDEAL_MAX_DISPLACEMENT:
return 3 # Stable
elif avg_displacement < ACCEPTABLE_MAX_DISPLACEMENT:
return 2 # Minor jitter
else:
return 1 # Major shift
Video Processing Loop
The main loop reads video frames, compares every Nth frame against a reference (updated only when stability is confirmed), and overlays the current status.
import cv2
from PIL import Image, ImageDraw, ImageFont
import numpy as np
STATUS_LABELS = ["Complete Misalignment", "Major Shift", "Minor Jitter", "Stable"]
def annotate_frame(frame, status_text):
rgb = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)
pil_img = Image.fromarray(rgb)
draw = ImageDraw.Draw(pil_img)
try:
font = ImageFont.truetype("arial.ttf", 28)
except:
font = ImageFont.load_default()
draw.text((40, 40), status_text, fill=(255, 255, 0), font=font)
return cv2.cvtColor(np.array(pil_img), cv2.COLOR_RGB2BGR)
cap = cv2.VideoCapture('input_video.mp4')
reference_frame = None
frame_counter = 0
while cap.isOpened():
ret, current = cap.read()
if not ret:
break
if reference_frame is None:
reference_frame = current.copy()
continue
frame_counter += 1
if frame_counter % 25 == 0: # Analyze every 25 frames
state = extract_and_match(reference_frame, current)
label = STATUS_LABELS[state]
if state >= 2: # Update reference only if stable or minor jitter
reference_frame = current.copy()
display = current.copy()
if display.shape[1] > 800:
display = cv2.resize(display, (800, int(800 * display.shape[0] / display.shape[1])))
annotated = annotate_frame(display, label)
cv2.imshow('Camera Stability Monitor', annotated)
if cv2.waitKey(1) == ord('q'):
break
cap.release()
cv2.destroyAllWindows()