🎯 Finding Your Taste Twin

Imagine walking into a party where you don't know anyone.

Someone notices your t-shirt and starts a conversation: "You like Arctic Monkeys? Me too."

"You also watched Severance?"
And your favorite game is Baldur's Gate 3?"

After five minutes, you realize you share almost identical tastes.

Then they say: "You HAVE to watch Silo."

Chances are you'll trust them. Not because you know anything about Silo. But because this stranger has repeatedly proven they like the same things you do.

That my friends, is collaborative filtering.

⚙️ The Mystery: How Does Netflix Know?

It's tempting to think Netflix stores detailed notes on every film ever made. Somehow the algorithm always recommends something you end up liking. Maybe it relates actors? or genres?or cinematography?

While some recommendation systems definitely can use that information, collaborative filtering doesn't need any of it. Instead, to the algorithm every movie is just a column in a giant spreadsheet and it asks a much simpler question: Who else behaves like you?

User Iron Man Barbie Titanic Matrix
Alice ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐⭐
Bob ⭐⭐⭐⭐ ⭐⭐⭐⭐
Carol ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐

A Quick History

Collaborative filtering dates back to the early 1990s. Researchers working on a project called GroupLens realised something:

People who agreed in the past often agree again in the future.

Amazon, Netflix, and Spotify later applied similar ideas to recommendations. Today, almost every major platform uses collaborative filtering as part of a larger system.

How the Algorithm Finds Your Taste Twin

Collaborative filtering doesn't understand products. It doesn't know anything about movies or music. It's based on the idea that people with similar past behavior often make similar choices.

You and another person watched these movies:

The recommendation engine gives a high similarity score between you and Kelly because your ratings are consistently alike.

If Kelly watches Dune and rates it with 5 stars, there's a good chance you will love Dune too.

Under the hood, these systems often compare users using mathematical similarity measures like cosine similarity or Pearson correlation. You don't need to understand the formulas to understand the idea: the system compares the users history with millions of other users. The more often two people agree, the more similar they become.

Measuring Similarity

So far, we've treated "similarity" as a vague concept, but recommendation systems need a number.

There are several ways to calculate how similar two users are:

Cosine Similarity

Imagine each user's ratings are represented as a vector.

You: [5, 4, 1, 5]
Kelly: [5, 4, 1, 4]

Cosine similarity measures the angle between those vectors.

A score close to 1 means the users have very similar preferences.
A score close to 0 means they're largely unrelated.

The actual formula looks intimidating:

cos(θ) = (A · B) / (||A|| ||B||)

Fortunately, the intuition is much simpler than the equation:

users whose rating vectors point in roughly the same direction tend to enjoy the same things.

A different approach

Everything we've seen so far is called User-Based Collaborative Filtering. Another approach is Item-Based Collaborative Filtering. In our previous movie example, imagine targeting similar movies instead of users. If thousands of viewers liked Interstellar AND Arrival AND Blade Runner, there is a good chance they will love Dune too.

This algorithm groups movies that frequently appear together, so when you watch 2-3 k-dramas, now your feed is full of them. Not because the system understands the genre, but because lots of people who loved the choice A loved the choice B too.

You Don't Even Need Ratings

I don't know who even remembers this, but in the early days of Netflix you could rate movies with stars. You don't actually need them. Modern recommendation systems often learn from your behaviour. Instead of asking if you like something or you, they watch things like "did she finish the movie?", "what did you rewatch", "what do you have in your wishlist", "what you stop after two minutes".

This is called implicit feedback, and it is often more reliable than ratings because it reflects what people actually do.

Where It Breaks Down

Collaborative filtering is powerful but it has some problems:

1. The New User Problem

If you are a new user in Netflix, you haven't watched anything. So nobody is at this point similar to you. This is called the Cold Start Problem.

Solution: Ask new users to rate a few movies first, or show them the most popular content until you have enough data.

2. The New Movie Problem

A new movie released today, but this means that nobody has watched it yet, and thus nobody has rated it. So the collaborative filtering cannot recommend it.

Solution: Use content-based filtering (analyzing genre, cast, plot) until the movie has enough ratings.

3. Taste Changes

A very human behaviour we have is that our taste changes over time. Maybe in 2018 you mostly watched cheesy hallmark movies and then in 2026 you watch only animes. The problem is that your old ratings still exist.

Solution: Recommendation systems weight recent behaviour more heavily than old behaviour, so your old tastes and choices doesn't haunt you forever.

4. Niche Content

The thing with recommendation systems is that they need data. A lot of them. So popular movies which have millions of ratings are more likely to get endorsed.

Independent films might have only a few hundred, which makes niche content harder to recommend. The algorithm works best with volume. Blockbusters get pushed. Hidden gems stay hidden.

5. The Filter Bubble

Recommendation systems can accidentally reinforce your existing preferences. If Netflix keeps recommending crime dramas, you're more likely to watch them. The system then becomes even more confident that crime dramas are what you like, creating a feedback loop. Over time this can reduce diversity in recommendations and make it harder to discover completely new kinds of content.

Key Takeaways

✅ Collaborative filtering doesn't need to understand movies or products but instead looks for people with similar behaviour.
✅ Similar people often make similar future choices.
✅ More data usually means better recommendations. The algorithm gets smarter as more people use it.
✅ Modern recommendation systems combine collaborative filtering with many other techniques. Content-based filtering, deep learning, and hybrid approaches all work together to overcome individual weaknesses.

Resources