How collaborative filtering works: recommendations from people like you

"Customers who bought this also bought" is collaborative filtering at work. Here is how it finds people with your taste, why it needs so much data, and where it breaks down.

Written by Amili, an AI writer, from the sources listed below · 6 October 2026 · 5 min read


Collaborative filtering is one of the two main techniques behind recommender systems. It predicts what you will like from the choices of many other people, betting that people whose choices matched before will tend to match again. It needs no understanding of the items themselves, only patterns in who liked what.

In short

  • It works from behaviour, such as ratings, purchases or listening history, not from the content of the items.
  • User-based versions find people whose ratings resemble yours; item-based versions find items that tend to be liked together.
  • Similarity is often measured with Pearson correlation or cosine similarity.
  • It struggles with newcomers (cold start), huge catalogues (scalability) and mostly empty rating tables (sparsity).
  • Content-based filtering is the main alternative: it recommends items whose features resemble ones you liked.

What is collaborative filtering?


Most of the time, good advice comes from someone whose taste resembles your own. A friend who loved the same three novels you loved is a better guide to the fourth than a bestseller list. Collaborative filtering automates that instinct. It collects the preferences of a large number of people and uses them to predict the interests of one person, which is where the name comes from: the filtering is done by pooling, or collaborating on, many people's tastes.

The approach rests on a simple bet. If two people have agreed on many things before, they are more likely to agree on the next thing than two people picked at random. Notice what the method does not need: it never has to know what a film is about or what a song sounds like. It only needs to see which people chose which items.

How does it work, step by step?


A typical system runs in three moves. First, it gathers signals of preference, either explicit ones such as star ratings or implicit ones such as what people bought, watched or played. Second, it compares your record with everyone else's to find your nearest neighbours, the people whose choices line up most closely with yours. Third, it looks at what those neighbours rated highly that you have not tried yet, and offers you those items.

Comparing two people means measuring how similar their ratings are. Two common yardsticks are Pearson correlation, which asks whether two people rate the same items high and low together, and cosine similarity, which treats each person's ratings as a direction and measures the angle between them. The final prediction is usually a weighted average: neighbours who resemble you closely count for more than distant ones.

There is a mirror-image version. Instead of finding similar people, item-based collaborative filtering finds items that tend to be chosen by the same people, which is the logic behind a line like "customers who bought this also bought that." It builds a table of how strongly each pair of items is linked and then matches it against your own history.

A worked example with three readers


Suppose three readers have rated four books. Ana and Ben both loved the first two books and both disliked the third. Cleo loved the third and disliked the first two. Ana has read the fourth book and loved it; Ben has not read it yet.

A user-based system sees that Ben's ratings track Ana's closely and Cleo's not at all, so it treats Ana as Ben's neighbour and predicts that Ben will like the fourth book too. Nothing about the books themselves entered the calculation, only the pattern of agreement between readers. Scale the same reasoning up to millions of people and items, and you have the core of many recommendation engines.

How is it different from content-based filtering?


Two early music services show the contrast well. Last.fm built its stations by comparing what you listened to with the listening of other users and then playing tracks those similar listeners played often, which is collaborative filtering. Pandora instead described songs by their musical attributes, drawn from the Music Genome Project, and played tracks with similar properties to a seed song, adjusting as you liked or disliked them, which is content-based filtering.

Each has a trade-off. The collaborative approach can surface surprising picks you would never have searched for, but it needs a lot of data about you first. The content-based approach can start almost immediately, but it tends to stay close to what you already know. Many modern recommender systems combine the two.

Where does collaborative filtering fail?


Three problems come up again and again. The first is cold start: a new user has no history to match and a new item has no ratings to spread, so the system has nothing to go on. One common response is a multi-armed bandit, which deliberately tries out new items on some people to learn about them. The second is scalability: with millions of users and products, finding everyone's neighbours takes serious computing power. The third is sparsity: most people rate only a tiny fraction of what is available, so the rating table is almost entirely empty and overlaps between any two people can be thin.

There is also a subtler limit. A score averaged across everyone ignores individual taste, which is why plain popularity lists do badly in areas where tastes differ widely, such as music. Collaborative filtering does better there, but it can only reflect the behaviour it observes, so it is only as good as the signals people leave behind.

A short history


The idea of a recommender system is older than the web. Elaine Rich built one called Grundy in 1979 to suggest books, sorting users into stereotypes based on their answers to questions. A "digital bookshelf" was described in a technical report in 1990, and from 1994 onwards research groups at SICS, MIT and Bellcore developed the idea further. The GroupLens project led by Paul Resnick later received an ACM Software Systems Award for its work.

Questions people ask


What is the difference between user-based and item-based collaborative filtering?

User-based filtering looks for people whose past ratings resemble yours and recommends what they liked. Item-based filtering looks for items that tend to be liked by the same people and recommends items linked to ones you already chose. Item-based methods power suggestions like "people who bought this also bought," and are often easier to compute at scale.

What is the cold start problem?

It is the difficulty of recommending anything to a new user, or recommending a new item, when there is not yet enough data about them. Collaborative filtering depends on past behaviour, so newcomers are invisible to it at first. A commonly used remedy is a multi-armed bandit algorithm that tries new items out; content-based methods, which need little data to start, can also fill the gap.

Why is it called collaborative filtering?

Because the filtering, deciding which items to show a person, is done by combining the preferences of many people. No single person curates the list; the collective pattern of ratings and choices does. In a broader sense, the term also covers filtering information by combining many data sources or viewpoints, not only user ratings.

The thinking behind it


Pedro Domingos explains how learning machines find patterns in what people choose, including the similarity-based learners that recommender systems grew from.

Read or listen to The Master Algorithm

Hear the whole book free: start an Audible trial and your first audiobook — this one, if you like — is on the house.

As an Amazon Associate, ReadGlobe earns from qualifying purchases and Audible trials — at no extra cost to you.

Sources

How this was made: Amili, an AI writer, wrote this article in its own words from the sources above. Every link was checked before publishing. Spotted an error? Tell us and we will correct it.

More algorithms, explained


Books readers reach for

As an Amazon Associate, ReadGlobe earns from qualifying purchases — at no extra cost to you.