Build a Recommendation Engine with Nori in Minutes
Use Nori to predict MovieLens ratings without task-specific training, compare against SVD++, and rank candidate movies with a short Python workflow.
Why is building a recommendation baseline difficult?
An unrated movie could be a perfect match, a poor fit, or something the user has never encountered. Ratings alone do not tell you which. New users have little history, new movies have few ratings, and a popular movie can still be the wrong choice for a particular person.
Common recommendation methods use different parts of the available data:
- Popularity and average-rating baselines recommend broadly liked movies. They are easy to build but give everyone similar suggestions.
- Content-based methods match movie attributes, such as genre, to a user's preferences. They depend on useful metadata and a way to infer those preferences.
- Collaborative filtering learns from user–movie interactions. Matrix-factorization methods, including SVD++, learn user and movie representations from the rating history; sparse histories and unseen users still need attention.
These approaches can also be combined. Google's recommendation-system guide separates the wider workflow into candidate generation, scoring, and re-ranking. Even a good rating predictor still needs code that selects eligible movies and turns scores into a useful list.
For a first baseline, the modeling step adds another decision: which algorithm to train, with which settings? Nori provides a pretrained model for scoring prepared user–movie rows, so you can start testing your features without building a task-specific training loop.
Build a movie recommendation baseline without task-specific training
Each row combines user attributes, movie metadata, and rating history. Nori uses labeled examples as context to predict ratings without task-specific training.
The workflow has three parts:
- Describe the pair. Combine user attributes, movie metadata, and past ratings into features.
- Predict a rating. Give Nori historical examples and the user–movie pairs you want to score.
- Evaluate and recommend. Measure rating error, or filter candidate movies and return the K highest-scoring movies for each user: a top-K list.
Prepare MovieLens ratings and predict with Nori
MovieLens 100K contains 100,000 ratings (1–5) from 943 users across 1,682 movies. We use u1.base for context (80,000 ratings) and u1.test for evaluation (20,000).
Each user–movie row combines user attributes, movie genres and release year, and rating-history summaries. The rating is kept separate as the prediction target.
| User–movie pair | User attributes | Movie metadata | Rating history | Separate target | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| user_id | item_id | age | gender | occupation | year | Action | Animation | Comedy | item_mean | user_mean | genre_affinity | rating |
| 1 Context | 1 Toy Story (1995) | 24 | 1 | 19 | 1995 | 0 | 1 | 1 | 3.862 | 3.621 | 3.061 | 5 Known label |
| 1 Context | 50 Star Wars (1977) | 24 | 1 | 19 | 1977 | 1 | 0 | 0 | 4.373 | 3.633 | 3.371 | 5 Known label |
| 2 Context | 1 Toy Story (1995) | 53 | 0 | 13 | 1995 | 0 | 1 | 1 | 3.926 | 3.938 | 3.803 | 4 Known label |
| 2 Query | 50 Star Wars (1977) | 53 | 0 | 13 | 1977 | 1 | 0 | 0 | 4.351 | 3.770 | 3.871 | ? Predict |
See all 37 input columns
A training row's own rating must not appear in its input statistics. We construct rating-derived features using five folds: each row receives statistics computed from the other four folds. For held-out rows, we compute statistics using the full training split.
With those tables prepared, the prediction step is:
1import numpy as np
2from synthefy_nori import MemoryPolicy, NoriRegressor
3
4# Prepared feature tables; rating is kept separate from the input columns.
5rating_model = NoriRegressor(
6 model="nori-6m",
7 encode_id_cols=["user_id", "item_id"],
8 memory_policy=MemoryPolicy(cache=True, allow_subsample=False),
9)
10rating_model.fit(train_features, train_ratings)
11predicted_ratings = np.clip(rating_model.predict(test_features), 1, 5)fit supplies the labeled context; it does not update the model's weights. The memory policy keeps all context rows. In this experiment, Nori uses the full 80,000-row training split.
How does Nori compare with SVD++ on MovieLens 100K?
SVD++ is an established ratings baseline, introduced by Yehuda Koren in 2008. It learns user and movie representations from rating history, providing a trained model to compare with Nori. We use Nicolas Hug's Surprise library, which benchmarks SVD++ on MovieLens.
| Method | Rating RMSE |
|---|---|
| Movie-average baseline | 1.037 |
| SVD++ collaborative filtering (Surprise, default settings, seed 0) | 0.933 |
| Nori-6M | 0.929 |
Nori received the richer feature table described above; SVD++ received user IDs, movie IDs, and ratings. These results compare the complete approaches on one split and do not establish a statistically significant accuracy advantage.
We also held out 20% of users entirely in a separate cold-start experiment. Nori did not improve on the movie-average baseline for those users. For a product serving new users, evaluate that population separately and choose a useful fallback.
Turn rating predictions into top-K movie recommendations
The same fitted model can score candidate movies for each user. Build candidate rows with the same feature columns, using only the training history to compute their statistics. Exclude movies the user has already rated in that history, then sort the predicted scores.
Here, candidates contains the prepared features plus user_id and item_id; seen_pairs contains the user–movie pairs from the training ratings. Candidate rows are unique per user and movie.
1import pandas as pd
2
3feature_columns = train_features.columns
4seen_index = pd.MultiIndex.from_frame(seen_pairs[["user_id", "item_id"]])
5candidate_index = pd.MultiIndex.from_frame(candidates[["user_id", "item_id"]])
6unseen = candidates.loc[~candidate_index.isin(seen_index)].copy()
7
8unseen["score"] = rating_model.predict(unseen[feature_columns])
9top_10 = (
10 unseen.sort_values(
11 ["user_id", "score", "item_id"], ascending=[True, False, True]
12 )
13 .groupby("user_id", sort=False)
14 .head(10)[["user_id", "item_id", "score"]]
15)This returns up to ten recommendations per user. The scores remain unrounded when sorting, so differences between predictions can determine the ordering.
Evaluating those lists requires a ranking protocol: which movies count as relevant, which candidates are eligible, and what K should be. Recall@K measures the fraction of relevant movies recovered in the first K results. Normalized discounted cumulative gain (NDCG@K) rewards placing more relevant movies earlier in the list. The MovieLens results above measure rating error; evaluating these ranked lists is a separate step.
Build a first baseline on your own ratings data
Start with historical user–item ratings and the attributes available when each prediction would have been made. Choose an evaluation split that matches your product: predicting another rating from an existing user and recommending to a new user are different tests.
Prepare the feature tables, supply the labeled examples as context, and predict held-out ratings with Nori-6M. Compare against a simple reference, such as the item's average rating, as you refine the features. If users will see a short list, evaluate that list with a fixed candidate set and ranking metric as well.
Start with the MovieLens notebook and adapt its feature preparation to your data. The Nori quickstart covers the prediction interface. To discuss a recommendation baseline for your ratings data, contact the Synthefy team.

