Face RestoreOn-device · No upload

How Face Restore works

Face Restore runs the classic open-source face-restoration pipeline — detect, align, restore, match tone, upscale, blend — inside the browser, on the visitor's own GPU or CPU. This page records exactly which models it runs, where their bytes come from, how the browser build was checked against a reference implementation, and what it cannot do.

The five stages

  1. Input. The photo is decoded with the browser's image decoder into a canvas. Photos whose long side exceeds 1024 px are reduced before the neural upscaler (2048 px for the other paths) so a 4× result stays within 4096 px and GPU memory.
  2. Face & landmark detection — YuNet. The photo is letterboxed into the model's fixed 640×640 input (BGR, 0–255, no normalisation, exactly as OpenCV feeds it). The twelve outputs at strides 8/16/32 are decoded with OpenCV's FaceDetectorYN rule: score = √(cls·obj); box centre = (cell + offset)·stride; size = evalue·stride; five landmarks from the same cell; non-maximum suppression at IoU 0.3. The default confidence threshold is 0.6; "Sensitive" lowers it to 0.35.
  3. Alignment + restoration — GFPGAN v1.4. The five landmarks (eyes, nose tip, mouth corners, ordered left-to-right in image space) are fitted to GFPGAN's 512×512 FFHQ template with a closed-form least-squares similarity transform (scale, rotation, translation — the same 4-parameter model as OpenCV's estimateAffinePartial2D). The canvas warps the face into the 512 crop; GFPGAN takes the crop in [-1, 1] RGB and returns the restored crop.
  4. Upscaling — Real-ESRGAN x4plus. The whole photo is upscaled 4× on WebGPU in fixed-size tiles with 16 px of replicated context on every side; constant tile shapes keep the compiled GPU kernels cached. Tile size is picked from the adapter's storage-buffer limit and the first tile's output is sanity-checked against a bilinear upscale — above a certain window size the WebGPU kernels return noise without raising an error (measured 2026-09-16: a 352-px window is fine and a 384-px window is broken on an Apple M-series GPU), so the app halves the tile and retries instead of trusting the output. Without WebGPU a Lanczos-3 resample is used and labelled. 2× output is the 4× result resampled down; 1× skips this stage.
  5. Skin-tone matching (肤色匀化). GFPGAN can shift the overall tone of a face. The app measures the mean CIE Lab colour of the face interior (an ellipse on the 512 template) on the blurry-but-colour-faithful input crop and on the restored crop, shifts the restored crop's chroma fully and its lightness half-way toward the input, and re-measures: the stats card reports the ΔE (CIE76) before and after. On the demo the shift is ΔE 3.1 → 1.3. Below ΔE 1.5 the row reads "no shift"; the step can be switched off.
  6. Blend. Each restored crop is warped back through the inverse transform (scaled by the output factor) under a soft mask — a square inset by 26 px and feathered with an 8 px Gaussian, GFPGAN's erode-then-blur recipe — and composited over the upscaled photo. Only the face's bounding box is touched.

Models and pinned sources

StageModelSizeLicencePinned source
DetectYuNet 2023mar0.23 MB fp32MIT (Shiqi Yu)opencv/face_detection_yunet @ 3cc26e7f
EnhanceGFPGAN v1.4340 MB fp32Apache-2.0 (Tencent)facefusion/models-3.0.0 @ 728b9659 (ONNX export of TencentARC's GFPGANv1.4)
Upscale (GPU, shader-f16)Real-ESRGAN x4plus fp1636 MBBSD-3-Clause (Xintao Wang)same repository, real_esrgan_x4_fp16.onnx
Upscale (fallback)Real-ESRGAN x4plus fp3270 MBBSD-3-Clausesame repository, real_esrgan_x4.onnx
Runtimeonnxruntime-web 1.27.027 MB wasmMIT (Microsoft)npm onnxruntime-web@1.27.0, WebGPU bundle build

Every weight file is served from SkillSafe's shared browser-model registry (models.skillsafe.ai), verified against the SHA-256 recorded in the app's manifest before it reaches the runtime, and cached in the browser. SHA-256 values, copyright lines and the licence texts are in models/NOTICE.txt. The files are unmodified.

How the browser build was verified

An independent Python reference (OpenCV's own cv2.FaceDetectorYN decoder, onnxruntime on the CPU, cv2.estimateAffinePartial2D, cv2.warpAffine) was run on the bundled demo photo and compared with the browser under the production Content-Security-Policy on an Apple M-series GPU (2026-09-16):

Limits and honest caveats

Privacy

Photos are decoded, processed and exported inside the page. The app makes no network request that carries image data: its only downloads are the page itself, the model weights and the runtime. The hosting edge (Cloudflare) injects its own bot-management and web-analytics scripts into every page, as on every SkillSafe app; Face Restore adds no analytics of its own. Settings and the language choice are kept in localStorage.

Credits

YuNet — Shiqi Yu and the OpenCV Zoo contributors. GFPGAN — Xintao Wang, Yu Li, Honglun Zhang, Ying Shan (Tencent ARC Lab). Real-ESRGAN — Xintao Wang, Liangbin Xie, Chao Dong, Ying Shan. ONNX exports of GFPGAN and Real-ESRGAN published by the facefusion project. onnxruntime — Microsoft. Demo photo: portrait by Christopher Campbell (CC0, Wikimedia Commons), deliberately degraded (downscaled to 240 px, blurred, JPEG quality 32) to stand in for a low-quality input. Face Restore is not affiliated with any of them.