Today we released the Virtual Try-On Safety Classifier on Hugging Face under Apache-2.0. It looks at the photos of a product and answers one question: if a shopper tries this on, what will the generated photo show?
We built it to gate products before they reach a virtual try-on model. Any try-on service that accepts arbitrary product images faces the same problem, so we are publishing the model, the labelling policy and the evaluation.
Why NSFW detectors don't work for virtual try-on
A virtual try-on takes a photo of a person and a photo of a product, and returns the person wearing the product. The risk is in the output, but the only thing you can check before generating is the product.
Generic NSFW detectors judge the pixels of the image they are given. That is the wrong question for a product photo:
- A lingerie flat lay contains no nudity. The virtual try-on result shows the shopper in lingerie.
- Pasties on a white background are a few centimetres of fabric. The virtual try-on result is topless.
- A fetish hood on a mannequin is not explicit. It is still a sexual product.
- A model photographed nude wearing only a necklace is explicit. The necklace is a normal product.
When we tested three popular open NSFW image classifiers (NudeNet, nsfw-vit and Falconsai) on an early set of 141 labelled products, they agreed with our labels on only 48 to 62% of them. Google's ShieldGemma 2, given custom policies written for this exact question, reached 71.6%. The models work as intended. They just answer a different question.
What the classifier predicts
Every product gets a level, ordered by severity:
| Level | What the virtual try-on result shows | Examples |
|---|---|---|
ok | The shopper normally dressed | Everyday clothing, sportswear, wetsuits, cosplay that covers like clothes, accessories |
revealing | Swimwear-level coverage | Bikinis and swimsuits, men's underwear, shapewear, micro shorts, most festival wear |
lingerie | Women's underwear, nothing intimate visible | Bras, panties, lingerie sets, bodysuits, garters, corsets worn as lingerie |
adult | Intimate areas exposed, or a sexual product | Micro swimwear, pasties, see-through pieces, open-cup designs, BDSM and fetish wear, sexual role-play costumes |
Three extra outputs sit next to the level:
- subject:
human_clothing,petornot_clothing. A dog harness and a bondage harness look alike in a packshot. Wigs, cosmetics and home goods are flagged as not clothing. - swimwear:
regular,near_microormicro. Regular and near micro arerevealing, micro isadult. The word "microkini" in a product name decides nothing; the photos do. - ravewear: true or false, for festival outfits. It does not change the level, but rave wear is where the model is least sure, so it helps to know.
The levels describe the output, not the product or the shop. The right gate depends on the audience: a lingerie brand will allow lingerie, a kids' marketplace may block revealing. The model also returns a blocked flag (probability of adult at 0.3 or above), which catches more adult products than the top level alone for a small rise in false blocks.
The hard cases
Most products are easy. The policy exists for the ones that are not, and every rule comes from a real product that needed a decision:
| Product | Level | Why |
|---|---|---|
| Sports bra sold as gym wear | ok | Worn as a top in public |
| Same cut sold as a lingerie bralette | lingerie | Women's underwear |
| Lace bra, nipples clearly visible in the photos | adult | See-through over nipples |
| Men's boxers | revealing | Swimwear-level coverage |
| Wetsuit, rash guard, burkini | ok | Full coverage |
| Leather trousers, latex-look dress | ok | Fashion pieces |
| Leather or latex catsuit sold as fetish wear | adult | Fetish suit |
| Leather chest harness over a shirt | ok | Accessory over clothes |
| Strap harness on bare skin | adult | Fetish item |
| Halloween animal mask | ok | Costume |
| Pup-play mask | adult | Fetish mask |
| Body chain styled over a dress | ok | Jewellery over clothes |
| Body chain on a bare torso | adult | Worn over nothing |
| Cosplay that covers like clothes | ok | Costume |
| Lingerie-style anime outfit | adult | Sexualised cosplay |
The full policy, with the rules for each level and tag, is published as LABELLING_POLICY.md next to the weights.
The examples at the top of this post are the released model run on our own demo catalog. The long-sleeve swimsuit is a useful case: long sleeves and a zip front, but the bottom is cut like a one-piece, so the virtual try-on shows swimwear-level coverage and the model says revealing. The halter sports bra next to it in our catalog stays ok at 0.97.
How we built it
Data. About 2,600 products (roughly 5,000 photos) from online shop catalogs, chosen to cover the hard cases: lingerie, micro swimwear, fetish wear, festival wear, cosplay and pet wear, next to a large share of everyday fashion so the model does not over-block. Only product photos are used, never photos of shoppers. Any product that could show a minor in a revealing or sexual context was removed from the data for good.
Labels. Claude labelled every product after looking at every photo and reading the title and description, following the written policy. A human reviewer checked calibration examples for every level and corrected the boundary cases, and those calls went back into the policy. The policy went through several versions; the lingerie level, for example, was split out of adult once it was clear that most underwear is not explicit.
Model. The base is Google's SigLIP 2 (so400m, patch 16, 384 px), Apache-2.0. Two variants ship:
| Variant | Size | What it is |
|---|---|---|
| finetuned (default) | about 0.9 GB | Last 6 transformer blocks and the pooling head trained end to end, with four linear heads |
| frozen | about 120 KB on top of SigLIP 2 | The original vision tower, logistic heads on the mean and max of the photo embeddings |
A product is scored over all of its photos, because one photo often hides what another shows (a back view, a detail shot of the fabric).
We also tried a frontier LLM prompted with the same policy. It was close on accuracy, but each call took about 1.6 seconds and it still needed a fallback for timeouts. A vision classifier runs in tens of milliseconds per photo on a GPU, gives the same answer every time, and can be retrained when the policy changes.
How well it works
Five-fold cross-validation over 2,596 products: every product is scored by a model that never saw it during training.
| finetuned | frozen | |
|---|---|---|
| Accuracy, 4 levels | 94.6% | 94.2% |
Adult products missed, blocked flag | 34 of 534 | 39 of 534 |
False adult, blocked flag | 52 | 55 |
| Lingerie recall | 90% | 89% |
| Subject accuracy | 99.2% | |
| Swimwear tag accuracy | 90.0% |
By product group (finetuned, top level):
| Group | Products | Accuracy |
|---|---|---|
| Everyday fashion catalogs | 1,032 | 97.5% |
| Rope and bondage | 75 | 100% |
| Swimwear brands | 230 | 95.2% |
| Lingerie brands | 120 | 94.2% |
| Anime and cosplay | 63 | 93.7% |
| Fetish, lingerie and pet catalogs | 631 | 92.6% |
| Exotic and micro swimwear | 330 | 91.8% |
| Festival and rave | 115 | 85.2% |
True micro swimwear: 91 of 95 caught.
The test set is deliberately hard. On everyday fashion catalogs, the kind most stores sell, accuracy is 97.5%.
Limitations
- Sheer lingerie is the hardest boundary. Whether nipples or genitals are clearly visible decides between
lingerieandadult, and look-alike products sit on both sides. Most remaining errors are there. - Festival and rave wear is the weakest group. Fishnet, mesh and rhinestones push the model towards
adult. - Mostly western e-commerce photography. Expect lower accuracy on very different photo styles.
- It classifies products, not people. Do not use it to identify or make decisions about a person.
Use it
# pip install torch open_clip_torch safetensors huggingface_hub pillow
from huggingface_hub import hf_hub_download
import importlib.util
spec = importlib.util.spec_from_file_location(
"tryon_safety",
hf_hub_download("Genlook/tryon-safety-classifier", "tryon_safety.py"),
)
tryon_safety = importlib.util.module_from_spec(spec)
spec.loader.exec_module(tryon_safety)
clf = tryon_safety.TryOnSafetyClassifier.from_pretrained(
"Genlook/tryon-safety-classifier", revision="v1.0"
)
pred = clf.predict(["front.jpg", "back.jpg"]) # all photos of one product
print(pred.level, pred.blocked, pred.subject, pred.swimwear, pred.ravewear)
Pin revision="v1.0" so an update never changes your results without you choosing it. Inference needs only open_clip and torch, and the weights are safetensors.
False positives, false negatives and ideas for the next version are welcome in the Discussions tab of the model repo.
What this means for Genlook
The Genlook Try-On API already checks every product before a virtual try-on is returned: adult and fetish products are refused, and swimwear, underwear and lingerie need a paid account. A refused virtual try-on fails with CONTENT_POLICY_VIOLATION and its credit is refunded (error reference). This classifier is the model we built for that check.
For developers, that means safety is part of the API rather than something to build around it. One endpoint covers every product type, at $0.02 per virtual try-on on a plan and down to $0.015 at volume ($0.04 pay as you go), and new accounts start with free credits.
Summary
- The release: an Apache-2.0 classifier on SigLIP 2 that predicts what a virtual try-on of a product would show, from the product's photos.
- Why it exists: NSFW detectors judge the pixels of a product photo, which is the wrong question for virtual try-on. Three of them scored 48 to 62% on our labels.
- Outputs: four levels (ok, revealing, lingerie, adult), plus subject, swimwear and ravewear tags.
- Accuracy: 94.6% across four levels on a deliberately hard set, 97.5% on everyday fashion, 100% on rope and bondage.
- Open: weights, inference code and the labelling policy are public. The training data is not.
Get the model on Hugging Face, get an API key at platform.genlook.app, or add virtual try-on to your store.
FAQ
