This week’s editionA weekly paper on machine learning and what we wear

The weekly paper of clothing, fit and machine learning

News

Return rates and the limits of recommendation

Fit systems are usually justified by the returns they prevent. That framing quietly decides what the system is allowed to be good at.

Ruled bar diagram of a returns funnel, narrowing from orders placed to the fit-related subset.
Ruled bar diagram of a returns funnel, narrowing from orders placed to the fit-related subset.

Almost every argument for automated fit advice arrives attached to the same justification: garments sent back cost money to ship, to inspect, to restock and sometimes to write off, and a system that improves the match between body and garment should reduce that flow. The reasoning is sound as far as it goes. The trouble is what happens when a single number becomes the target.

Returns are not one thing

A garment comes back for many reasons and only some of them are about fit. It can arrive damaged. It can be the wrong item. The colour on the screen can differ from the colour in daylight, which is a rendering and photography problem rather than a sizing one. The cloth can feel different in the hand than it looked. The buyer can have ordered two sizes deliberately, intending from the outset to keep one, which is a fit-related return in the data and a successful transaction in practice. And the garment can simply not suit the person, which is a matter of taste and not a measurement failure at all.

Collapsing these into one rate makes them look like a single quantity that a single system can move. It also makes the system look worse than it is when a photography problem gets counted against it, and better than it is when a change of mind gets prevented by an interface that made ordering less appealing.

The cheapest way to lower the number

Consider what a system optimised purely against returns will learn to do. Recommending the safer, looser size lowers the chance of a garment being sent back as too small. Declining to recommend at all, or hedging every recommendation, lowers the chance of being blamed for a bad one. Suppressing the products with the most variable fit removes the category that generates the most returns. Each of these reduces the metric. None of them makes anyone better dressed, and the last two reduce it by reducing what the buyer is shown.

Each of these reduces the metric. None of them makes anyone better dressed.

This is the ordinary failure mode of a proxy objective, and clothing is an unusually clear case because the thing being proxied is so obviously not the thing being measured. The goal is a garment that fits and that the wearer keeps because they want to wear it. The measurement is a logistics event.

What a better target would look like

A more honest evaluation would separate the reasons. Distinguishing size-related returns from taste-related and quality-related ones is largely a matter of asking, and the answer is usually collected at the point of return and then discarded into a free-text field nobody reads. That field is the most useful signal in the whole system and it is routinely the least processed.

It would also account for the returns that should have happened. A garment kept because sending it back is inconvenient is not a success. Retention measures friction as much as satisfaction, and a system that improves its numbers by making returns harder has improved nothing about clothing.

Recommendation has a ceiling

There is a further limit worth stating plainly. Recommendation cannot fix a garment that does not exist. If a body falls between two sizes in a range that is graded in large steps, no advice can produce a good outcome; the honest recommendation is that neither size will sit well, and that is a sentence very few systems are permitted to say. The information the system holds at that moment, that this particular body is poorly served by this particular range, is genuinely valuable, and it is valuable to the people who decide how the range is graded rather than to the person about to place an order.

Framed as a returns problem, fit work is a filter placed at the end of a process. Framed as a measurement problem, it is evidence that could be fed back into how clothes are cut. The second framing is harder to justify on a quarterly basis and is the one more likely to produce clothes that fit.

Also on the News beat

Elsewhere in this issue

Return to the News beat · Issue archive