The Score in One Sentence
Copy Opt's new score answers one question: what share of your product's real shopper demand does your listing copy actually cover?
The word doing the work in that sentence is "demand." In the previous version of Copy Opt, demand meant one thing: the search queries shoppers type. The new score measures your copy against two kinds of demand at once. The queries shoppers type, and the intents behind them, meaning what the shopper is actually trying to accomplish. A listing can echo every top search term and still never say who the product is for, what situation it serves, or what it works with. Those unstated intents are demand your copy is leaving uncovered, and Amazon's own research says its ranking systems now look for them.
Query coverage tells you whether your copy matches what shoppers type. Intent coverage tells you whether it answers why they typed it. The new score measures both, weighted by how much demand each one represents.
What COSMO Actually Is (And Isn't)
In 2024, Amazon published a paper describing COSMO, its large-scale commonsense knowledge system for e-commerce ("COSMO: A Large-Scale E-commerce Common Sense Knowledge Generation and Serving System at Amazon," SIGMOD 2024). The motivating example in the paper is a shopper searching for "shoes for pregnant women." Nothing in a shoe listing says "pregnant." What the shopper needs is slip resistance, comfort, and easy on-off. Amazon's classic matching systems could not bridge that gap, because the connection is not in the words. It is in commonsense knowledge about why someone with that need buys that product.
COSMO closes the gap by mining those connections at scale: large language models generate candidate relationships from real search and purchase behavior (this product is used for that situation, capable of serving that need), and human review plus critic models filter them into a knowledge graph that Amazon serves inside product search.
Two things are worth being precise about, because much of the seller commentary on COSMO gets them wrong:
- COSMO is an augmentation layer, not a new ranking engine. It feeds intent knowledge into the same retrieval and ranking pipeline Amazon has documented for years. The semantic embedding foundation we covered in our earlier post on how Copy Opt works is still the substrate. There was no switch flipped from an old algorithm to a new one.
- The direction is real, though. Amazon is investing in understanding what shoppers are trying to accomplish, not just what they type. Listings that state use cases, audiences, and compatibility give both the knowledge layer and the shopper more to connect with.
Copy Opt's new scoring model takes the same step from the seller's side: it stops treating queries as the whole of demand and starts scoring the intents behind them as demand in their own right.
Where Query Coverage Falls Short
Here is the failure mode the old score could not see. Take a dog bed with strong query data. The copy mentions "orthopedic," "washable," and "large dog bed," so query relevance looks healthy. But look at what those queries are actually expressing:
A listing that says "orthopedic memory foam dog bed, machine-washable cover, 42 x 30 inches" matches the queries. But it never says for senior dogs, never says fits inside a 42-inch crate (dimensions are not the same claim as fit), and makes the shopper do the work of connecting foam density to an aging dog's joints. Those are three pieces of demand the copy leaves on the table, and they are invisible to a score that only checks queries.
The Demand Map: Two Kinds of Demand, One Score
The new model represents your product's demand as a single set of demand units. A demand unit is one thing shoppers want from your product, with a weight for how much of total demand it represents.
Query units come straight from your real search data (Search Query Performance and ads search term reports). Their weights come from measured search volume, so a query driving ten times the impressions carries ten times the mass.
Intent units are generated per product by a language model that reads your top queries, your listing content, and your category, then names the intents behind demand. The vocabulary is deliberately closed. Every intent must be one of seven types:
| Intent type | What it names | Example |
|---|---|---|
| Audience | Who the product is for | "for nurses", "for pregnant women" |
| Use case | The job or situation it serves | "standing all day", "keeps feet dry" |
| Occasion | An event or moment | "wedding guest", "gift for new parents" |
| Location | Where it is used | "small apartments", "outdoor patio" |
| Body compatibility | Fit or physical constraint | "wide feet", "sensitive skin" |
| Season | Time of year | "summer", "rainy season" |
| Complement | What it works with | "fits queen frames", "pairs with scrubs" |
An intent that does not fit one of these types is dropped rather than invented. That constraint keeps the generator honest: it can only surface the kinds of demand that COSMO-style intent understanding actually models, not restatements of the product or vague marketing themes.
Each intent unit inherits its weight from the queries that express it, scaled by how directly each query states the intent (a query that says the intent almost literally contributes more mass than one that merely hints at it). The generator can also add up to three category priors: intents that are obvious for the category even though no query expressed them, like "easy to keep clean" for a dog bed. When the two populations combine, query units carry 60% of total demand mass and intent units 40%.
The intent list is an inventory you control, not disposable model output. Dismiss an intent and it stays dismissed through every future regeneration. Add one by hand and regeneration never touches it. The generator only re-runs when its inputs actually change, so the list is stable between real changes to your demand data.
Measuring Coverage in Two Tiers
With the demand map built, the question for every unit is the same: does the copy cover this? The model answers it in two tiers, cheap and broad first, expensive and precise where it matters most.
Tier 1 runs on every unit. It is the embedding approach from the previous generation of Copy Opt: your copy and each demand unit are encoded into the same semantic space, and proximity estimates coverage. It is cheap enough to run across the whole demand map on every update.
Tier 2 runs on the heaviest units. For the highest-weight units (in practice, the top few dozen units carry roughly two-thirds or more of total demand mass), a language model reads your actual copy and issues a ruling for each unit:
- Explicit (full credit): the copy contains a word or phrase that means the same thing as the demand unit. A shopper skimming would see their need named.
- Implied (70% credit): the copy names materials, components, or construction details from which a reader could infer the claim, without stating the outcome. "Orthopedic memory foam" implies support for aging joints. It does not state it.
- Absent (15% credit): the copy neither names the outcome nor names a mechanism that would produce it. The small residual reflects that embeddings still register faint topical proximity.
The distinction is modeled on Amazon's own relevance science, which grades matches on a spectrum rather than yes/no. It also matches shopper behavior: a claim the reader has to derive is a claim most readers never receive. The 30% gap between explicit and implied is deliberate. Making an implied claim explicit is usually a one-line edit, and the score rewards it accordingly.
Rulings are cached and re-run only when your copy or the demand unit changes, so tier 2 judgment is precise without being wasteful. If a ruling is ever unavailable, the unit falls back to its tier 1 estimate. You always get a score.
The Roll-Up
The final score is coverage-weighted demand mass: each unit's weight times its coverage credit, summed across the demand map, divided by total mass. There is no keyword checklist and no per-field formula. Just one question asked of every unit of demand, weighted by how much that demand matters.
The weighting has a practical consequence worth internalizing: covering the heavy head of demand moves the score a lot, and covering long-tail trivia barely moves it. That is deliberate. The score is designed to point your copywriting effort at the demand that matters, in proportion to how much of it there is.
How to Read Your Score
Three patterns show up constantly once you look at listings through this lens:
- Strong queries, weak intents. The copy matches search terms but never states who the product is for or what situation it serves. This is the signature of well-kept keyword hygiene under the old playbook, and it is exactly the demand COSMO-era search rewards. The intent gaps in your demand map are the highest-leverage edits available.
- A pile of "implied" rulings. The ideas are in the copy but stated as mechanisms rather than outcomes. Each one is usually a one-line rewrite from 70% credit to full credit. These are the fastest wins on the page.
- Heavy uncovered head. One or two high-mass units ruled absent. A single sentence addressing the heaviest uncovered unit routinely does more for the score than a full rewrite of the long tail.
Sort by weight, fix absents in the head, promote implieds to explicit, and only then worry about the tail. The score is built so that this order and "biggest score gain first" are the same order.
What's Next: Grading the Score Against Reality
A score you cannot audit is a vibe. Two more pieces are in progress:
- Calibrated bands. The thresholds for "strong" and "needs work" are being derived from the real distribution of scores across live catalogs, so a band reflects where listings actually land rather than a round number someone liked.
- An outcome loop. After you publish a copy change, we will track your impression, click, and purchase share in Amazon's Search Query Performance data across windows before and after the publish. The model's opinion of your copy gets graded against what shoppers actually did, and the grading feeds back into how we score.
- Amazon's COSMO system (published SIGMOD 2024) adds commonsense intent understanding to product search. It augments the existing semantic ranking pipeline rather than replacing it, and it rewards listings that state who a product is for, what it serves, and what it works with.
- Copy Opt's new score measures your copy against total demand: real queries (60% of demand mass) and the intents behind them (40%), each weighted by measured search volume.
- Intents live in a closed vocabulary of seven types (audience, use case, occasion, location, body compatibility, season, complement), and the intent list is an inventory you control: dismissals persist, hand-added intents are never overwritten.
- Coverage is measured in two tiers: embedding similarity on every unit, plus an LLM that reads your copy and rules each heavy unit explicit (full credit), implied (70%), or absent (15%).
- The score is coverage-weighted demand mass, so the fastest way to raise it is also the right priority order: cover the heaviest absent units first, then promote implied claims to explicit ones.
- Score bands are being calibrated against real score distributions, and an outcome loop will grade the score against post-publish impression, click, and purchase share.