← Back to Detection Studio

Technical Methodology & Formulas

A deep dive into the algorithmic pillars powering our multi-modal deepfake detection system.

1. Core Architecture (Spatial Ensemble)

The foundation of our detection pipeline is a dual-model CNN ensemble comprising Xception and EfficientNet-B4. These models complement each other perfectly:

  • Xception: Uses depthwise separable convolutions to identify microscopic texture anomalies and high-frequency noise.
  • EfficientNet: Uses compound scaling to capture macro-spatial deformations like asymmetrical facial warping and incorrect lighting gradients.

2. Focal Loss for Class Imbalance

Deepfake datasets (like FaceForensics++) suffer from heavy class imbalance (mostly fake videos, fewer real videos). Training with standard Binary Cross-Entropy causes models to collapse and overfit. To solve this, we optimize the network using Focal Loss, forcing it to focus on hard, realistic examples.

FL(p_t) = -α_t(1 - p_t)^γ log(p_t)

Where γ = 2.0 (focusing parameter) and α = 0.14 (weight for the positive class).

3. Top-k Weighted Temporal Aggregation

A naive approach to video classification averages the scores of all frames. This heavily dilutes transient manipulation glitches (false negatives). Conversely, taking the absolute maximum frame score triggers false positives on motion blur. We solve this mathematically by blending the top-k (k=3) frame scores with the overall mean.

S_final = 0.5 * (Avg of Top-3 Scores) + 0.5 * (Avg of All Scores)

4. Biological Signals (rPPG Heartbeat SNR)

Synthetic faces generated by GANs do not replicate the volumetric blood flow of a living human. Every time a human heart beats, facial capillaries absorb slightly more green light. Our system extracts the Remote Photoplethysmography (rPPG) signal from the forehead over 90 frames, applies a Butterworth Bandpass Filter (0.7 - 2.5 Hz), and computes the Fast Fourier Transform (FFT).

SNR = E_peak / (E_total - E_peak)

A low Signal-to-Noise Ratio (SNR < 1.5) triggers a heavy deepfake penalty.

5. Frequency Domain Analysis (FFT)

Generative AI upsampling layers weave periodic checkerboard artifacts into images. Invisible in RGB, these are glaringly obvious in the frequency domain. We compute the 2D FFT and calculate the ratio of high-frequency energy to low-frequency energy.

M(u, v) = 20 log(|F_shift(u, v)| + ε)

High-Freq Ratio = E_high / E_low

6. Full Stack Tech Stack

Backend: FastAPI, PyTorch, MTCNN, SciPy, MoviePy
Frontend: Next.js (React), CSS Modules
Deployment: Hugging Face Spaces (GPU Inference), Vercel