Probably somewhat, and nobody selling it has measured that properly. The figures you will be shown range from 15 percent to 65 percent, they come almost entirely from companies selling the technology, and the two most-cited ones compare shoppers who used a try-on tool against shoppers who did not, which is not the same thing as measuring what the tool did.
There is also a more basic problem underneath the numbers. The biggest single cause of apparel returns is sizing, and rendering a garment onto a photograph of a shopper does not fix sizing. That does not make the technology useless. It makes it useful for something other than the thing it is most often sold for, which is worth knowing before you sign a contract priced against a returns forecast.
What actually causes apparel returns
Of the ten results on this search, exactly one is published by somebody who does not sell virtual try-on: the retail research firm Coresight. Their finding is worth quoting in full because it does the diagnostic work everything else skips. Incorrect sizing is the top reason for online apparel returns, “likely because the industry does not use a standardized sizing system and apparel brands and retailers are unable to capture what the customer really looks like in different styles and sizes of apparel.”
That sentence contains two separate failures. There is no shared sizing grammar between brands, so a medium means different things in different catalogues and a shopper cannot carry knowledge from one to the next. And the retailer has no model of the individual body it is selling to, so it cannot say what any given garment will do on that person.
Neither of those is a picture problem. A visualisation tool does not standardise a size grade across an industry, and it does not measure anyone. It shows a garment on a body, which is genuinely valuable and addresses a real source of hesitation, but it is a different source from the one Coresight identifies as the largest.
One caveat on that source, since this page is about taking claims seriously: the Coresight report is dated April 2023. On a subject moving this fast, three years is old, and it is the newest independent research the search surfaces. That absence is itself a finding.
The concession hiding in the vendor FAQs
The vendors are not hiding this, exactly. They mention it and move on.
Camclo3D publishes an FAQ headed “How accurate is AI virtual try-on for sizing?” It opens with “AI virtual try-on has become highly accurate,” and then, in the same breath: “While it is primarily a visualization tool, modern systems use body mapping technology to estimate how a garment will fit specific body proportions.” Read the subordinate clause rather than the headline. The category is a visualisation tool that estimates fit as a secondary capability.
The second result on the search makes the point more sharply, and it is a fit-technology company making it. Naiz Fit published a post titled “Virtual Try-On is great. But is it enough to fix your returns?” They go on to say they are working on making their try-on “size-responsive,” so that shoppers can see “not just how a garment looks, but how it fits across sizes.” A company building that feature is telling you the category does not currently have it. That post sits behind a login wall, so it is quoted here from the search result rather than read in full.
Reading the numbers you will be shown
Seven observations about the evidence base, in the order that matters if somebody is quoting a returns figure at you.
1. The two market-size figures do not agree
Stytrix describes returns as “a $743 billion problem.” Camclo3D describes returns as “a $890 billion problem for US retailers as of 2026.” Those are presented as the same quantity, they are roughly twenty percent apart, and neither page carries a citation for its own figure.
This is a small thing and it is a useful tell. If the number at the top of the page, the one chosen to establish that the problem is enormous, is unsourced and inconsistent with the next vendor’s, that sets a reasonable prior for the numbers further down that are harder to check. Nobody is necessarily wrong here. Different definitions of scope, geography and year produce genuinely different totals. That is precisely why the definition should be stated.
2. What the headline percentages actually compare
Stytrix carries the densest set of figures and, to its credit, attributes two of them. Snap Inc. is cited for a 2.4 times increase in purchase intent and a 36 percent reduction in return rates for fashion products. Google research is cited for shoppers who used augmented-reality try-on spending 2.7 times longer with products and being 65 percent less likely to return.
The same page also publishes its own tier table, which pairs price bands with return-reduction ranges: 15 to 20 percent at $500 to $2,000 a month, 25 to 35 percent at $2,000 to $10,000, and 30 to 40 percent at $10,000 to $50,000, with accuracy labelled moderate, high and very high respectively. Those are Stytrix’s figures for Stytrix’s tiers, and vendor pricing in this category moves faster than any article can track.
Take the two attributed ones seriously for a moment, because they are the strongest evidence on the search. Both describe a comparison between shoppers who used the feature and shoppers who did not.
3. Why that comparison overstates the effect
A shopper who opens a virtual try-on tool has already told you something about themselves. They are engaged enough to click an extra thing, deliberate enough to want a closer look, and invested enough in getting this particular purchase right to spend time on it. Shoppers like that return less than average whatever software is on the page.
So when a study reports that try-on users returned 65 percent less than non-users, that gap contains two effects welded together: whatever the tool did, and whatever kind of person chooses to use tools. Nothing in the design separates them. The measurement is real, the arithmetic is presumably correct, and the causal claim it is used to support does not follow from it.
This is not an accusation of dishonesty. It is the most natural analysis to run, because the data falls out of ordinary event tracking with no experimental setup at all. It is also the analysis that reliably flatters any optional feature, which is why the same shape of number appears for live chat, product video, reviews widgets and configurators. If a feature is opt-in, comparing users to non-users will make it look good.
4. The one page that names the right method
Uwear.ai, describing how a returns team should approach this, writes: “Where it fits, layer in virtual try-on, then track returns on try-on engaged orders versus control.”
That is the whole answer, and it is one line on one page out of ten. Hold back a randomised slice of traffic that never sees the feature. Ship it to everyone else. Compare return rates between the two groups rather than between users and non-users. Because assignment is random rather than chosen, the two groups contain the same mix of deliberate and impulsive shoppers, and the difference between them is attributable to the tool.
Worth noticing: the page that names the correct method is the one that publishes no return-reduction percentage of its own. That is internally consistent, and it is the reverse of the pattern everywhere else on the search. Also worth noticing that this test is portable. It works for any optional feature a merchant is being sold, and running it once teaches you more about your own catalogue than any published benchmark can.
5. Size bracketing has two causes and only one is visual
Bracketing is ordering the same item in two or three sizes intending to keep one. Camclo3D diagnose it well: it “is not a sign of a happy customer; it is a sign of an uncertain customer. They bracket because they do not trust the size chart. They bracket because they cannot visualize how the fabric will drape on their specific body type.”
Those are two distinct problems sitting in one sentence. Distrust of the size chart is a data problem, and it is fixed with garment measurements rather than letter sizes, per-item specifications, fit models with published dimensions, and honest guidance about which items run small. Inability to picture drape is a visualisation problem, and that one virtual try-on genuinely does address.
A merchant who buys visualisation expecting to solve both has bought half the solution and will, quite reasonably, be disappointed by the returns line. A merchant who fixes sizing data and adds visualisation has addressed both halves, and will have no idea which one worked unless they ran the test in the previous section.
6. The downside nobody else mentions
Wanna, which sells this technology, is the only source on the search to say the intervention can go backwards: “poorly executed experiences can actually hurt sales. If the virtual try-on experience is not realistic or does not accurately represent how the clothing or accessories would look on the customer, it can lead to dissatisfaction, damage to a brand’s reputation, and lower sales.”
That is worth more than most of the positive claims, because it establishes that this is a real intervention with a real failure mode rather than a free upgrade. An unconvincing rendering does not simply fail to help. It tells a shopper something inaccurate about the product, and if they buy on that basis, the resulting return is one the tool caused. Any honest returns measurement has to be able to detect that outcome, which is another argument for the control group.
7. What the technology is genuinely good at
Strip out the returns claim and a solid case remains, which the vendors make almost in passing while chasing the bigger one.
Consistent on-model imagery at catalogue scale is the real product. Photographing every item on multiple body types is expensive and slow, and for a long-tail catalogue it is simply not done, which is why so many product pages show a flat lay and nothing else. Generating that imagery instead means a shopper sees the garment on a body for items that would otherwise never have justified a shoot. Uwear.ai frame their whole offer around consistency of on-model imagery for exactly this reason, and Zakeke make the same case for visualisation generally. Tools in this group, Pic Copilot’s AI product imagery among them, are aimed at that job rather than at fit prediction.
The engagement effects are also the least contested findings on the search, and they do not depend on the causal claim that gets people into trouble. More time with a product and higher purchase intent are worth paying for on their own terms. Buy the tool for imagery cost, catalogue coverage and conversion, judge it on those, and treat any effect on returns as a hypothesis you intend to test rather than a line in the business case.
What to do if you are buying this
Three things, and the first is the one almost nobody does. Run the holdback: keep a randomised control group that never sees the feature, and compare return rates against it for long enough to clear seasonal noise. Fix your sizing data in parallel rather than instead, because garment measurements and per-item fit notes attack the cause Coresight identified as largest and cost nothing but attention. And judge the imagery on imagery grounds, comparing it to what a photoshoot would cost for the same coverage.
One note on scope. The formats differ more than the marketing suggests: a flat two-dimensional overlay, a three-dimensional garment simulation and a live augmented-reality experience are separate technologies at separate prices with separate failure modes, and the word “try-on” covers all three. This page is about clothing specifically. Eyewear and cosmetics try-on is a genuinely different problem, because glasses and lipstick have no drape and no size grade, so the fit critique above does not apply to them at all.