Skip to the case study
Morris Lin
Selected work

Computer Vision / ResearchJan — Apr 2026

AI Facial Recognition Research

Comparing five pretrained face-verification models on accuracy, false acceptance, and false rejection — 50,000 verification runs, one evaluation framework, on the LFW benchmark.

Try the live demoUpload two photos and run the same detect, embed and compare pipeline in your browser, with an adjustable match threshold.Open demo

The question

As a Research Assistant at UGA's Multispectral Imagery Lab (advised by Dr. Thirimachos Bourlai), I set out to answer a practical question: how do five pretrained face-verification models actually compare on accuracy, false acceptance, and false rejection when you run them through the same evaluation framework?

The goal was never to build a new face-recognition algorithm. It was to treat five well-known models — VGG-Face, ArcFace, FaceNet, FaceNet512, and OpenFace — as black boxes, run them through an identical pipeline on the same 50,000 image pairs, and find out which one actually holds up.

Dataset and approach

Everything runs on Labeled Faces in the Wild (LFW) — 13,233 unconstrained face photos of 5,749 people, collected from the web with natural variation in pose, lighting, expression, and occlusion. From the full set, I selected every identity with at least 3 images, giving 901 people and 7,606 images to work with.

From that pool I generated 15,235 image pairs: 6,225 genuine (same person) and 9,010 impostor (different people), with a duplicate check on every pair so the same comparison never got counted twice. The full pipeline: preprocessing (grayscale, rotation correction, cropping) → face detection (OpenCV) → feature extraction (each model's pretrained embedding) → cosine-distance comparison → accept/reject against a threshold.

Four-stage preprocessing pipeline applied to LFW sample images: original, grayscale, histogram-equalized, and cropped
Fig 1. Preprocessing pipeline applied to LFW sample images — original, grayscale, histogram-equalized, cropped.

Getting the pipeline right first

Before scaling up, I ran the pipeline on a small custom set — my own photos against public figures — to sanity-check basic behavior. One early result mattered more than the rest: with enforce_detection turned off, an image where no face could be detected still got compared, treating the entire frame as "the face." That reliably produced a false verdict no matter who was actually in the photo.

That finding changed the entire evaluation design — enforce_detection was set to True for every subsequent test, and any pair where a face couldn't be found was skipped rather than forced through. A 100-pair pilot on LFW (50 genuine, 50 impostor) confirmed the pipeline worked before committing to the full run: 96% accuracy, 0% FRR, 8% FAR.

Large-scale evaluation: 15,235 pairs

Running ArcFace across the full 15,235-pair set (901 identities, 7,606 images) gave 94.87% accuracy, a 10.55% false-rejection rate, and a 1.39% false-acceptance rate. In plain terms: about 1 in 10 genuine matches got rejected, mostly from pose changes, occlusion (glasses, hats), and lighting differences between photos of the same person — while the system stayed conservative about accepting impostors.

Looking at the actual error cases makes the failure modes concrete rather than abstract statistics.

Grid of same-person image pairs that ArcFace incorrectly rejected as different people, with distance scores
Fig 2a. False rejections — same person, incorrectly split by pose, lighting, or angle.

The other kind of error

False acceptances tell a different story — mostly pairs where facial features were partially obscured, resolution was low, or two different people simply shared enough structural similarity under similar lighting to fool the embedding.

Grid of different-person image pairs that ArcFace incorrectly accepted as matches, with distance scores
Fig 2b. False acceptances — different people, incorrectly matched.

Model comparison — 5 models, 10 runs each, 50,000 pairs

The core experiment: VGG-Face, ArcFace, FaceNet, FaceNet512, and OpenFace, each run 10 independent times over 1,000 randomly-sampled pairs per run — 50,000 verification evaluations total, at each model's own default threshold.

ModelAvg AccuracyAvg FARAvg FRRAvg Time
ArcFace93.8%1.6%10.7%~610s
VGG-Face93.0%2.1%11.4%~470s
FaceNet~56%0.3%42.5%~780s
FaceNet512~56%0.3%42.5%~780s
OpenFace51.1%~0%~97%~290s
Line graphs showing accuracy, FAR, and FRR across 10 independent test runs for each of five models
Fig 3. Accuracy, FAR, and FRR across all 10 runs per model — ArcFace and VGG-Face are consistently stable; OpenFace is consistently bad.

Why the gap is so large

ArcFace came out on top with the best balance of accuracy, FAR, and FRR, and low run-to-run variance — meaning it generalizes well across random subsets of the data, not just one lucky sample. VGG-Face was a close, faster second. FaceNet and FaceNet512 were the surprise: both rejected over 40% of genuine pairs at their default thresholds, which is impractical however low their FAR looked on paper. OpenFace effectively rejected almost everyone (97% FRR) — functioning as close to a random rejector as a "working" system can get.

This was the most counterintuitive finding of the whole project: FaceNet is widely cited as high-performing, but its triplet-loss training pushes different identities aggressively far apart — great for strict security contexts, badly mismatched to LFW's pose and lighting variation.

Grouped bar chart of mean accuracy, FAR, and FRR with standard deviation error bars for all five models
Fig 4. Mean performance ± standard deviation across all five models. ArcFace: high accuracy, low error, low variance.

Threshold tuning

Default thresholds are tuned on whatever dataset each model was originally trained on, not on LFW. So I swept threshold values for ArcFace, FaceNet, and VGG-Face and tracked how accuracy, FAR, and FRR moved together, to find where each model actually balances on this dataset.

ModelOptimal ThresholdAccuracyFARFRR
ArcFace0.7094.8%1.6%8.8%
FaceNet0.6092.3%5.0%~12%
VGG-Face0.7093.7%5.2%~9%
Heatmaps of accuracy, false acceptance rate, and false rejection rate across threshold values for three models
Fig 5. Threshold sweep heatmaps — accuracy peaks around 0.6–0.7 for every model tested.

What tuning actually bought

After tuning, ArcFace at 0.70 improved from 93.8% to 94.8% accuracy while cutting its FRR from 10.7% down to 8.8% — more genuine users correctly let through, with almost no added security risk. FaceNet's accuracy jumped the most (its default threshold was badly miscalibrated for LFW), but that came at the cost of a 5.0% FAR, worse than ArcFace's. VGG-Face improved similarly but landed at a less favorable 5.2% FAR.

The takeaway generalizes past this one dataset: a model's out-of-the-box threshold is tuned for whatever benchmark its authors used, not necessarily yours — any real deployment needs its own calibration step.

Bar chart comparing accuracy, FAR, and FRR before and after threshold optimization for three models
Fig 6. Before vs. after threshold optimization — FRR drops substantially for all three models, FAR trades up slightly.

ROC curve — how good is ArcFace, really?

To evaluate ArcFace independent of any single threshold choice, I plotted a full ROC curve — true positive rate against false positive rate across every tested threshold. The result: AUC = 0.9786, meaning ArcFace correctly ranks a genuine pair above an impostor pair 97.86% of the time, across the entire operating range. That's a strong result for an uncontrolled, real-world benchmark like LFW.

The selected operating point (threshold 0.70) sits at 91.2% TPR and 1.6% FPR — near the curve's knee, where you can't meaningfully trade more security for more usability or vice versa without giving something up.

ROC curve for ArcFace on the LFW dataset showing true positive rate versus false positive rate, AUC 0.9786
Fig 7. ROC curve for ArcFace on LFW (AUC = 0.9786). The blue curve hugging the top-left corner is what a strong verifier looks like.

Where errors actually come from

Pose variation was the single biggest driver of false rejections — a frontal photo compared against a profile or three-quarter view of the same person, which OpenCV's detector handles poorly. Accessories (glasses, hats, masks) caused both kinds of error: rejections when present in only one photo, occasional false acceptances when they dominated the embedding. Poor lighting (backlight, heavy shadow, low exposure) pushed genuine pairs' embeddings apart enough to read as different people.

Two engineering lessons stood out. First, enforce_detection is not a minor setting — forcing detection on an undetectable face silently corrupts results rather than failing loudly. Second, aggregate accuracy alone hides the real story: FAR and FRR are different failure modes (security risk vs. usability risk) and have to be read together, which is exactly what the ROC and threshold-sweep analysis were for.

Limitations

No liveness detection — a printed photo or screen replay could plausibly spoof this system, and that was never tested. Everything ran on static images; there's no evaluation on video, near-infrared, or thermal imagery. All five models were used strictly as black boxes with zero fine-tuning, which likely caps performance below what a domain-adapted version could reach. And this is a single-dataset result — LFW has known demographic imbalance and a heavy skew toward celebrity photos, so these numbers may not transfer cleanly to a different population or capture setup.

Future work

Fine-tuning ArcFace on a domain-specific dataset to push FRR down further; adding a liveness-detection stage to close the spoofing gap; validating against additional benchmarks (VGGFace2, IJB-C) to check whether these findings generalize; a more rigorous demographic fairness analysis with actual statistical testing; and trying alternative detector backends (RetinaFace, MTCNN) that may handle difficult poses better than the OpenCV Haar cascade used here.

Closing

Across 50,000 verification runs on real, unconstrained photos, ArcFace with a tuned threshold of 0.70 was the clear, reproducible winner for this dataset and task. But the more durable finding is procedural: a face-verification system's reported accuracy means very little without knowing its FAR/FRR trade-off, whether its threshold was calibrated for your actual data, and what its failure modes look like when you go looking for them.

Full source code (pair generation, model comparison, threshold tuning, all scripts) is available at github.com/Linsanity12/face-recognition-uga.