πŸ”¬ UnivFD β€’ CVPR 2023 β€’ CLIP ViT-L/14

Universal AI Image Detector

State-of-the-art universal detector for synthetic AI imagery. Powered by frozen CLIP ViT-L/14 feature representations, 2D FFT frequency lattice analysis, PRNU sensor noise forensics, and C2PA provenance triage.

🧠 UnivFD Neural Probe
πŸ“Š 2D FFT Spectral Lattice
πŸ” PRNU Sensor Fingerprint
πŸ›‘οΈ C2PA & EXIF Provenance
πŸ”¬ Deep Forensic Scan (5-Point Spatial Ensemble) Ensemble

Evaluates center and 4 corner spatial crops. Detects localized inpainting, facial deepfakes, and mixed edits.

⚑ Fast Scan (Global Center Crop) < 2s

High-speed single-pass inference (< 2s). Optimal for quick triage and whole-image generations.

⚠️
πŸ”¬

Analyzing Image with UnivFD...

Forensic Inspection Completed

AI Probability

AI Probability
Authentic Probability
Confidence Level
Analysis Time

UnivFD (Universal Fake Detector, CVPR 2023) maps images into CLIP ViT-L/14 high-dimensional latent space. Generative models (GANs, Diffusion, Flow Matching) inhabit distinct sub-manifolds from authentic camera captures.

Spatial Patch Breakdown

Evaluation across 5 strategic image regions to detect localized editing or face replacement:

Center
Top-Left
Top-Right
Bottom-Left
Bottom-Right

2D Fast Fourier Transform (FFT) decomposes pixel luminance into spatial frequency components. Generative upsampling (e.g. transposed convolution) creates unnatural periodic grid spikes along cardinal axes.

Radial frequency energy exhibits high-frequency distribution with an axial spike ratio of .

Interpreting the Map: High-frequency lattice dots and十字-axis rays reveal artificial convolutional upsampling (e.g. Stable Diffusion UNet layers).

Physical camera sensors exhibit Poisson-Gaussian shot noise and Photo-Response Non-Uniformity (PRNU). AI-generated imagery exhibits synthetic smoothing, missing grain, or uniform mathematical noise.

High-pass Laplacian residual variance measured at .

Interpreting the Map: Physical optical camera sensors exhibit fine, stochastic grain (PRNU). Pure synthetic images exhibit smooth patches or unnatural uniform noise.

Cryptographic C2PA manifests, EXIF camera hardware parameters, and PNG generation parameters (prompts, samplers, model hashes).

C2PA / Provenance Status

How UnivFD Universal AI Detection Works

Most AI image detectors are trained on a single generator and fail when tested on new diffusion architectures. UnivFD solves this fundamental generalization problem.

1

1. Why CLIP ViT-L/14 Generalizes

UnivFD leverages OpenAI CLIP ViT-L/14 backbone, pre-trained on 400M diverse image-text pairs. Frozen features preserve rich semantic and structural invariants, while a trained linear probe separates authentic photos from synthetic distributions.

2

2. 2D Spectral Frequency Signatures

Neural generators produce subtle checkerboard frequency artifacts due to deconvolutional upsampling. 2D FFT spectral analysis unmasks these invisible mathematical fingerprints regardless of post-processing filters.

3

3. Multi-Crop Inpainting & Face Swap Detection

By segmenting the image into spatial crops, UnivFD pinpoints whether an image is 100% synthetic or an authentic photo with localized AI-generated modifications (e.g. generative fill or deepfake face swaps).

Frequently Asked Questions

UnivFD detects images created with Midjourney (v4, v5, v6), Stable Diffusion (1.5, 2.1, XL, SD3), Flux, DALL-E 2/3, Adobe Firefly, Imagen, StyleGAN, ProGAN, BigGAN, and generative inpainting tools.
Yes. CLIP representations are remarkably resilient to JPEG compression, resizing, color adjustments, and social media transcoding.
A score above 75% indicates strong synthetic fingerprints. A score below 25% signifies authentic optical sensor characteristics. Intermediate scores (40-60%) indicate heavily compressed, stylized, or hybrid composite images.
No. Images are processed temporarily for neural inference and deleted automatically according to our privacy retention policy. They are never used to train models.