Amazon’s core collaborative filtering approach is item-to-item similarity based on co-occurrence, not user-to-user matching. It was designed for scale, documented publicly in 2003, and has since evolved into managed implementations like SIMS and newer graph and deep learning models including MERLIN and graph neural networks. What follows covers the underlying math, the production engineering that keeps it fast, the operational limits of Amazon Personalize, and a practitioner checklist for building or improving your own system.
TL;DR:
- The item-to-item similarity approach reduces computational complexity by relying on precomputed co-occurrence and relatedness scores, rather than user-to-user comparisons.
- Similarity metrics are adjusted for popularity and buyer behavior, with sparse matrix storage and top-k neighbor lists optimizing scalability and retrieval speed.
- Recomputing similarity tables periodically, based on catalog and interaction update frequency, balances freshness with operational cost, especially for fast-moving categories.
- Amazon Personalize’s SIMS uses a co-occurrence threshold to maintain recommendation quality, but falls back to popularity if interaction data remains sparse.
- Graph neural networks and multimodal embeddings can improve asymmetrical and sparse-data recommendations but require more resources and are situationally advantageous depending on data volume and use case.
Table of Contents
- History and Origin of Amazon’s Item-to-Item Approach
- The Math Behind Item Relatedness and Similarity Scores
- Production Architecture: Precomputing Similarities at Scale
- Amazon Personalize SIMS: What It Does and Where It Breaks
- Modern Extensions: Graph Neural Networks and Multimodal Embeddings
- Measuring What Works: Metrics, Offline Tests, and A/B Design
- Implementation Checklist for Teams Building This
- When SMBs Should Build This and When to Get Help
- What We’d Prioritize in the Next 90 Days
- How We Help Teams Build and Scale Recommendation Systems
- FAQ
- Sources
History and Origin of Amazon’s Item-to-Item Approach
Amazon’s recommendation engine did not start as item-to-item. Early collaborative filtering systems compared users to other users, looking for people with similar purchase histories and recommending what those “neighbors” bought. The approach worked conceptually but collapsed under real traffic. User profiles change constantly, catalogs run into the millions, and comparing every user to every other user scales quadratically with the user base, a cost that becomes unworkable at Amazon’s size.
The 2003 IEEE paper on Amazon.com recommendations reframed the problem around items instead of users. Rather than asking “which users resemble this one,” the method asks “which items tend to be purchased alongside this item,” building a relatedness table once per item and reusing it across every shopper who views or buys that item. Item catalogs change far less often than user behavior does, so the resulting similarity index stays valid longer and needs less frequent rebuilding.
Amazon Science’s own retrospective credits this item-centric design with solving two problems at once: it cut the computational burden from comparing users to comparing items, and it produced relatedness scores stable enough to serve from a precomputed table rather than recalculating on every request. The 2003 paper was recognized by IEEE Internet Computing in 2017 as a “test of time” award winner, a rare distinction for a systems paper that was, by then, already 14 years old and still describing a technique in active production use.
Several corrections followed the original publication, most notably around how relatedness is estimated for items purchased by unusually active buyers. The core lessons from that early work still shape how practitioners build item-to-item systems today:
- Comparing items instead of users turns a quadratic scaling problem into something closer to linear, since the item catalog grows far more slowly than the interaction volume.
- Item similarity indices are stable over medium time horizons, which lets teams precompute and cache them instead of scoring in real time.
- Co-occurrence alone is a weak signal until it is corrected for popularity and buyer behavior, a point the next section covers in detail.
The Math Behind Item Relatedness and Similarity Scores
Item-to-item collaborative filtering starts with a co-occurrence matrix: for every pair of items A and B, count how often they appear together in the same user’s history, whether through purchases, views, or both. That raw count is the seed of relatedness, but using it directly produces bad recommendations, because popular items co-occur with everything simply by virtue of being popular.
The fix in the 2003 methodology is to measure relatedness as a differential probability: how much does buying item A raise the probability of also buying item B, compared to the baseline probability of buying B at all? Expressed loosely, relatedness(A, B) is proportional to P(B | A) minus P(B), rather than P(B | A) alone. An item that nearly everyone buys regardless of context will have a high P(B) and therefore a muted relatedness score even when its raw co-occurrence count with A look large.
A second correction matters just as much: heavy buyers distort the baseline. A customer who purchases hundreds of items a year will show up in the co-occurrence count for nearly every popular pair, inflating relatedness scores for items that have nothing meaningfully in common. The 2003 paper and subsequent work address this by discounting the contribution of high-volume buyers when estimating the baseline purchase probability, so a handful of unusually active accounts cannot dominate the relatedness signal for the rest of the catalog.
Once you have a workable relatedness measure, several similarity metrics compete for the role of “which items get linked” in production:
- Cosine similarity treats each item as a vector of user interactions and measures the angle between vectors, a fast and popular default that ignores rating or interaction magnitude.
- Adjusted cosine similarity subtracts each user’s average rating or interaction intensity before computing cosine distance, correcting for users who rate everything high or interact with everything frequently.
- Lift divides observed co-occurrence by expected co-occurrence under independence, directly capturing the differential-probability intuition from the original paper.
- Pointwise mutual information (PMI) is the log of that same ratio, which compresses extreme values and is less sensitive to a handful of outlier pairs than raw lift.
In practice, the choice matters less than the sparsity strategy around it. Catalogs with millions of items produce co-occurrence matrices that are overwhelmingly zero, so production systems store interactions as sparse matrices (compressed sparse row or column formats) and compute similarity only for item pairs that actually co-occur above some minimum threshold. For each item, the system keeps only the top-k most related items rather than a full similarity row, which turns an otherwise intractable all-pairs comparison into a bounded nearest-neighbor lookup that can be precomputed offline and served from a simple key-value store at request time.
Production Architecture: Precomputing Similarities at Scale
The math only matters if it survives contact with production traffic. Amazon’s item-to-item design assumes similarity computation happens offline, in batch, well before any customer loads a page. The serving path then becomes a lookup, not a calculation: given an item ID, fetch its precomputed top-k neighbor list and return it, typically in single-digit milliseconds.
That split between offline computation and online serving defines the architecture:
- A batch job periodically recomputes co-occurrence counts and relatedness scores across the full interaction history, producing a sparse top-k neighbor table for every item in the catalog.
- The neighbor tables get written to a low-latency store (a key-value database or an in-memory cache) that the recommendation service queries directly at request time.
- Approximate nearest-neighbor (ANN) indexes such as HNSW or FAISS-style structures handle cases where similarity is computed over learned embeddings rather than raw co-occurrence, trading a small amount of recall for large gains in lookup speed.
- Sharding by item ID or category lets the precompute job and the serving layer scale horizontally as the catalog grows, rather than requiring a single machine to hold the entire similarity table in memory.
Update cadence is the real engineering decision, and it is a trade-off between freshness and cost. Recomputing full similarity tables daily works well for catalogs where relatedness shifts slowly, like books or home goods. Fast-moving categories, such as trending apparel or seasonal items, often need hourly or incremental updates that fold in new interactions without a full batch recompute, usually by maintaining running co-occurrence counts and refreshing only the neighbor lists for items that received new signal.
Long-tail items create a separate problem: an item with only a handful of interactions produces a noisy, unreliable neighbor list no matter how good the similarity metric is. Production systems typically set a minimum interaction or co-occurrence threshold below which an item falls back to a simpler signal, like category-level popularity, rather than serving a neighbor list built from statistical noise.
Pro Tip: Cap your top-k neighbor lists at a fixed size (20 to 50 items is typical) and store them as sparse arrays rather than dense vectors; this keeps memory flat as your catalog grows instead of scaling with the square of item count.
Amazon Personalize SIMS: What It Does and Where It Breaks
For teams that do not want to build item-to-item infrastructure from scratch, Amazon Personalize’s SIMS recipe implements the same core idea as a managed service. SIMS computes item similarity purely from the interaction dataset, co-occurrence of items across user sessions, and deliberately ignores item metadata like category, brand, or description when generating similarity scores. That makes it a direct, if simplified, descendant of the original 2003 approach rather than a hybrid content-and-behavior model.
SIMS exposes a min_cointeraction_count hyperparameter that sets the minimum number of times two items must co-occur in user interactions before they are considered related, with a typical default value. Raising this threshold reduces noise from coincidental pairings but also shrinks the neighbor lists available for less popular items. The documentation also flags that SIMS needs a meaningful volume of interaction history to train well, noting that having a sufficient number of unique interactions in the training data is generally necessary for the recipe to produce useful recommendations rather than defaulting to generic popularity-based results.
That fallback behavior matters operationally. When an item or a user segment lacks sufficient interaction history, Personalize does not fail outright. It falls back toward popularity-based recommendations, which keeps the system functional but means teams tracking “personalization lift” need to separate true item-to-item recommendations from popularity fallbacks when measuring impact.
Retraining follows a solution-version lifecycle: you create a solution, train a version against your current interaction dataset, and deploy a campaign from that version. New interaction data does not automatically improve live recommendations until you retrain and deploy a new solution version, which means retraining cadence is a deliberate operational decision, not something the service handles silently in the background.
Common pitfalls teams hit with SIMS:
- Treating the default
min_cointeraction_countof 3 as universally correct, when sparse catalogs often need a lower threshold and high-traffic catalogs benefit from a higher one. - Assuming item metadata influences SIMS output, when the recipe’s similarity score is driven entirely by interaction co-occurrence.
- Failing to monitor the ratio of true similarity-based recommendations to popularity fallbacks, which can mask a data sufficiency problem.
- Skipping scheduled retraining, so the live model drifts further from current catalog and interaction patterns over time.
Modern Extensions: Graph Neural Networks and Multimodal Embeddings
Item-to-item collaborative filtering has a structural limitation: a simple co-occurrence score treats relatedness as symmetric. If a phone case co-occurs with a phone, the raw relatedness score usually treats A-to-B and B-to-A the same way, even though a customer buying a phone is far more likely to also need a case than a customer buying a case is to also need that specific phone.
Amazon’s research on examples of AI in e-commerce recommendation and personalization addresses exactly this asymmetry. The approach builds a directed product graph from co-purchase and co-view behavior, then trains separate source and target embeddings for each item rather than a single shared vector. An item’s “source” embedding captures what it tends to lead to, while its “target” embedding captures what tends to lead to it, letting the model represent directional relationships that a symmetric similarity score cannot. Amazon reports substantial improvements in hit rate and mean reciprocal rank (MRR) over prior baselines using this dual-embedding structure, and the same research notes that including co-view edges, not just co-purchase edges, helps correct for selection bias introduced by uneven product exposure on the site.

Separately, MERLIN and related multimodal embedding work extend the signal set beyond interaction history entirely, folding in product images, text, and other heterogeneous signals into a shared embedding space so that items with little interaction history can still be placed sensibly relative to items that have rich behavioral data.
None of this makes item-to-item CF obsolete. Amazon’s own researchers, writing at RecSys on evaluation, bias, and algorithms, caution against treating deep learning as a universal upgrade: for many generic collaborative filtering tasks, simpler methods perform comparably, and which approach wins depends heavily on the specific data and use case rather than on model sophistication alone. Practical guidance for choosing between them:
- Reach for item-to-item CF first when interaction volume is solid, latency budgets are tight, and the relationships you care about are reasonably symmetric (grocery staples, commodity goods).
- Consider GNNs or dual-embedding models when asymmetric relationships matter a lot to the business (accessory attachment, complementary purchases) and you have engineering capacity for graph infrastructure.
- Expect materially higher training time, inference latency, and explainability costs with deep and graph models, which also require more instrumentation to debug when recommendations look wrong.
- Treat multimodal embedding approaches as a cold-start and sparse-catalog tool rather than a wholesale replacement for interaction-based similarity where interaction data is already abundant.
Measuring What Works: Metrics, Offline Tests, and A/B Design
A recommendation system is only as good as the evaluation that validates it, and offline metrics alone routinely mislead teams into shipping models that underperform in production. A disciplined evaluation process layers several checks:
- Offline ranking metrics like hit rate, recall@k, MRR, and NDCG measure how well a model ranks the items a user actually interacted with against a held-out set, giving a fast, if imperfect, proxy for model quality before any live traffic is involved.
- Time-aware evaluation splits training and test data by timestamp rather than randomly, preventing the model from “seeing the future,” a form of leakage that inflates offline scores and does not reflect how the model will perform once deployed.
- A/B testing remains the only reliable way to connect a model change to business outcomes: click-through rate, conversion rate, and revenue per session matter more than any offline ranking score, and small lifts need adequate sample sizes and guard-rail metrics (return rates, customer complaints) to interpret safely.
- Novelty and diversity checks catch a subtle failure mode where a model optimizes for accuracy metrics by simply recommending the most popular items to everyone, technically scoring well while delivering no real personalization.
The practical discipline is sequencing these correctly: validate offline first to filter out obviously broken candidates, then commit engineering effort to an A/B test only for models that clear a meaningful offline bar, since live experiments are expensive in both calendar time and opportunity cost.
Implementation Checklist for Teams Building This
Building an item-to-item system that holds up in production comes down to a handful of decisions made early and revisited often.
Data signals and weighting. Purchases are the strongest signal, followed by cart adds, then clicks, then views, roughly in that order of commitment. Most teams weight these differently rather than treating every event as equal, and deduplicate repeated events from the same session so that a user refreshing a product page five times does not get counted as five independent signals of interest.
Cold start handling. New items with no interaction history cannot generate a co-occurrence-based neighbor list, so a fallback strategy is mandatory, not optional:
- Use item metadata (category, brand, attributes) as a temporary similarity proxy until enough interaction data accumulates.
- Augment sparse interaction signals with co-view data, which accumulates faster than co-purchase data and still carries real relatedness signal.
- Blend a content-based score with the interaction-based score during the item’s early life, then phase out the content signal as interaction volume grows.
Storage and retrieval choices. Sparse matrix formats (CSR or CSC) keep memory manageable for catalogs where most item pairs never co-occur. When similarity is computed over learned embeddings instead of raw co-occurrence, an ANN index such as HNSW trades a small amount of recall accuracy for large gains in query speed, which matters once catalogs move past a few hundred thousand items. Two parameters deserve explicit tuning rather than defaults: the minimum co-interaction threshold (too low adds noise, too high starves the long tail) and the neighbor-k size (too small limits recommendation diversity, too large adds latency and dilutes relevance).
Monitoring. The gap between offline and online metrics is the single most common source of surprise in production recommendation systems. Track that divergence directly rather than assuming a strong offline score will translate, and set alerting thresholds on it so a silent drift gets caught before it shows up as a revenue problem. Drift detection on the underlying interaction distribution, not just on model output, catches catalog shifts (a new product line, a seasonal change in buying patterns) before they degrade recommendation quality.
Pro Tip: Log the popularity-fallback rate as its own metric from day one; a rising fallback rate is often the earliest signal that your minimum co-interaction threshold has become too strict for your current catalog size.
For teams weighing build-versus-buy on the broader recommendation stack, a closer look at retrieval-to-rank architecture patterns and a vendor comparison of recommender software are useful next stops before committing engineering time.

When SMBs Should Build This and When to Get Help
Item-to-item collaborative filtering is a pragmatic first step for a mid-market company once interaction volume clears a reasonable floor, typically when a catalog has enough repeat purchase or view behavior that co-occurrence patterns stop looking like noise. Below that floor, a popularity-based fallback or lightweight rules engine often outperforms a half-trained similarity model, and the engineering cost of building and maintaining a custom pipeline is rarely justified at that stage.
A workable roadmap looks like this: start with a data audit to confirm interaction volume and quality, prototype with a SIMS-style approach to validate the concept cheaply, then iterate on metrics (hit rate, conversion lift) before committing to a custom production build. Each stage answers a different question, and skipping straight to custom infrastructure before validating the concept is the most common way teams waste budget on recommendation projects. Our team at BizDev Strategy works with operational leaders through exactly this kind of staged evaluation, pairing technology advisory with hands-on build support when a prototype clears the bar for full productionization.
What We’d Prioritize in the Next 90 Days
If we were starting from scratch, the first 90 days would go to instrumentation before modeling: get purchase, cart, click, and view events logged cleanly and deduplicated, because no similarity metric fixes bad input data. Next comes a baseline SIMS-style prototype, cheap to stand up and good enough to validate whether co-occurrence signal exists at all in your catalog. Only then design the A/B test that will actually decide whether to invest further.
The most common mistake we see is skipping straight to a sophisticated model before confirming the basics hold up, then spending months debugging a system that was never going to work because the underlying data was too sparse or too noisy. Watch the popularity-fallback rate and the gap between offline and online metrics from week one, not after launch.
— Hayden
How We Help Teams Build and Scale Recommendation Systems
Getting from a working prototype to a production recommendation system that holds up under real traffic is where most internal teams stall, not because the algorithms are exotic, but because the surrounding infrastructure (data pipelines, retraining cadence, monitoring, hosting) takes real engineering discipline to get right. That is the gap we close.
Our services most relevant to this work include:
- Technology Advisory to evaluate whether a managed approach like SIMS, a custom item-to-item build, or a graph-based model fits your data volume and team capacity.
- AI Enablement to help your team stand up the data pipelines and retraining workflows a production recommender needs.
- Cloud Infrastructure support for hosting, scaling, and sharding similarity indexes as your catalog and traffic grow.
- Software Integration to connect a recommendation engine cleanly into your existing storefront, app, or commerce platform.
We also offer a Free Technology Assessment for teams unsure whether their current data and infrastructure can support a real item-to-item system yet. If you would rather talk through your specific catalog and interaction data before committing engineering time, book a meeting with our team and we will walk through what a realistic first build looks like for your business.
FAQ
What is Amazon collaborative filtering based on?
Amazon’s collaborative filtering is based on item-to-item similarity computed from co-occurrence in purchase and interaction data, not on comparing users to other users. The original approach, documented in a 2003 IEEE paper, measures how much buying one item raises the likelihood of buying another compared to a baseline rate.
How does item-based collaborative filtering differ from user-based filtering?
Item-based filtering compares items to each other using co-occurrence patterns across all users, while user-based filtering compares users to each other based on shared purchase or rating history. Item-based approaches scale better because item catalogs change more slowly than user behavior, letting similarity tables be precomputed and reused, a design choice Amazon Science traces back to the original 2003 system.
What is the minimum co-interaction count in Amazon Personalize SIMS?
The min_cointeraction_count hyperparameter in Amazon Personalize’s SIMS recipe defaults to 3, meaning two items need at least three co-occurrences in the interaction data before being treated as related. Teams can adjust this threshold based on catalog size and interaction sparsity.
Do graph neural networks replace item-to-item collaborative filtering?
Graph neural networks extend item-to-item collaborative filtering rather than replace it, particularly for modeling asymmetric relationships like accessory attachment that simple co-occurrence scores treat as symmetric. Amazon’s own research notes that for many generic collaborative filtering tasks, simpler item-to-item methods remain competitive depending on the data and use case.
What happens when Amazon Personalize SIMS lacks enough interaction data?
When interaction data is insufficient, Amazon Personalize’s SIMS recipe falls back toward popularity-based recommendations rather than failing outright. The documentation indicates that a meaningful volume of unique interactions is needed for the model to produce genuine similarity-based results instead of relying on this fallback.
Sources
Start with the 2003 IEEE item-to-item paper, Amazon Science’s history of the algorithm, the AWS Personalize SIMS documentation, and the graph neural network research blog.
- The history of Amazon’s recommendation algorithm
- Amazon.com recommendations: item-to-item collaborative filtering | IEEE Xplore
- SIMS recipe in AWS Personalize documentation

