Meet Nori. The open-source model for structured data.
AI & Machine LearningEngineering7 min read

Build a Recommendation Engine with Nori in Minutes

By

Use Nori to predict MovieLens ratings without task-specific training, compare against SVD++, and rank candidate movies with a short Python workflow.

#Recommendation Systems#MovieLens#Rating Prediction#Tabular Foundation Models
Ratings and item metadata become user–item features; Nori predicts scores that application code ranks into top-K recommendations.
With user–item features prepared, write the prediction and ranking code in minutes. Nori uses labeled examples as context without task-specific training.

Why is building a recommendation baseline difficult?

An unrated movie could be a perfect match, a poor fit, or something the user has never encountered. Ratings alone do not tell you which. New users have little history, new movies have few ratings, and a popular movie can still be the wrong choice for a particular person.

Common recommendation methods use different parts of the available data:

  • Popularity and average-rating baselines recommend broadly liked movies. They are easy to build but give everyone similar suggestions.
  • Content-based methods match movie attributes, such as genre, to a user's preferences. They depend on useful metadata and a way to infer those preferences.
  • Collaborative filtering learns from user–movie interactions. Matrix-factorization methods, including SVD++, learn user and movie representations from the rating history; sparse histories and unseen users still need attention.

These approaches can also be combined. Google's recommendation-system guide separates the wider workflow into candidate generation, scoring, and re-ranking. Even a good rating predictor still needs code that selects eligible movies and turns scores into a useful list.

For a first baseline, the modeling step adds another decision: which algorithm to train, with which settings? Nori provides a pretrained model for scoring prepared user–movie rows, so you can start testing your features without building a task-specific training loop.

Build a movie recommendation baseline without task-specific training

Each row combines user attributes, movie metadata, and rating history. Nori uses labeled examples as context to predict ratings without task-specific training.

The workflow has three parts:

  1. Describe the pair. Combine user attributes, movie metadata, and past ratings into features.
  2. Predict a rating. Give Nori historical examples and the user–movie pairs you want to score.
  3. Evaluate and recommend. Measure rating error, or filter candidate movies and return the K highest-scoring movies for each user: a top-K list.

Prepare MovieLens ratings and predict with Nori

MovieLens 100K contains 100,000 ratings (1–5) from 943 users across 1,682 movies. We use u1.base for context (80,000 ratings) and u1.test for evaluation (20,000).

Each user–movie row combines user attributes, movie genres and release year, and rating-history summaries. The rating is kept separate as the prediction target.

From MovieLens records to Nori inputs
80,000 context rows · 37 input columns
User attributes
Age and category codes
Age stays numeric. Gender and occupation become integer category codes.
Movie metadata
Year and genre flags
Release dates become years. Each genre gets a column: 1 means present, 0 means absent.
Rating history
Summaries of past ratings
Smoothed movie and user averages, plus the user’s affinity for the movie’s genres.
Scroll across to see the feature columns. The rating target stays on the right.
User–movie pairUser attributesMovie metadataRating historySeparate target
user_iditem_idagegenderoccupationyearActionAnimationComedyitem_meanuser_meangenre_affinityrating
1
Context
1
Toy Story (1995)
2411919950113.8623.6213.0615
Known label
1
Context
50
Star Wars (1977)
2411919771004.3733.6333.3715
Known label
2
Context
1
Toy Story (1995)
5301319950113.9263.9383.8034
Known label
2
Query
50
Star Wars (1977)
5301319771004.3513.7703.871?
Predict
Three real context rows and one held-out query show 12 of the 37 input columns; history values are rounded to three decimals. Movie titles and Context/Query labels help you read the table. The model receives the input columns, with known ratings passed separately and the query rating withheld.
Category codes in these rows: gender 0 = F, 1 = M; occupation 13 = other, 19 = technician. Context summaries use the other four folds, so even the same movie can have different means across rows. Query summaries use all 80,000 context ratings.
See all 37 input columns
user_id, item_id, age, gender, occupation, age_band, unknown, Action, Adventure, Animation, Children, Comedy, Crime, Documentary, Drama, Fantasy, Film-Noir, Horror, Musical, Mystery, Romance, Sci-Fi, Thriller, War, Western, year, item_mean, item_mean_n, item_mean_by_gender, item_mean_by_gender_n, item_mean_by_age_band, item_mean_by_age_band_n, item_mean_by_occupation, item_mean_by_occupation_n, user_mean, user_mean_n, genre_affinity

A training row's own rating must not appear in its input statistics. We construct rating-derived features using five folds: each row receives statistics computed from the other four folds. For held-out rows, we compute statistics using the full training split.

With those tables prepared, the prediction step is:

Python
1import numpy as np
2from synthefy_nori import MemoryPolicy, NoriRegressor
3
4# Prepared feature tables; rating is kept separate from the input columns.
5rating_model = NoriRegressor(
6    model="nori-6m",
7    encode_id_cols=["user_id", "item_id"],
8    memory_policy=MemoryPolicy(cache=True, allow_subsample=False),
9)
10rating_model.fit(train_features, train_ratings)
11predicted_ratings = np.clip(rating_model.predict(test_features), 1, 5)

fit supplies the labeled context; it does not update the model's weights. The memory policy keeps all context rows. In this experiment, Nori uses the full 80,000-row training split.

How does Nori compare with SVD++ on MovieLens 100K?

SVD++ is an established ratings baseline, introduced by Yehuda Koren in 2008. It learns user and movie representations from rating history, providing a trained model to compare with Nori. We use Nicolas Hug's Surprise library, which benchmarks SVD++ on MovieLens.

Internal MovieLens 100K results on the 80,000/20,000 u1.base / u1.test split. Lower root mean squared error (RMSE) is better.
MethodRating RMSE
Movie-average baseline1.037
SVD++ collaborative filtering (Surprise, default settings, seed 0)0.933
Nori-6M0.929

Nori received the richer feature table described above; SVD++ received user IDs, movie IDs, and ratings. These results compare the complete approaches on one split and do not establish a statistically significant accuracy advantage.

We also held out 20% of users entirely in a separate cold-start experiment. Nori did not improve on the movie-average baseline for those users. For a product serving new users, evaluate that population separately and choose a useful fallback.

Turn rating predictions into top-K movie recommendations

The same fitted model can score candidate movies for each user. Build candidate rows with the same feature columns, using only the training history to compute their statistics. Exclude movies the user has already rated in that history, then sort the predicted scores.

Here, candidates contains the prepared features plus user_id and item_id; seen_pairs contains the user–movie pairs from the training ratings. Candidate rows are unique per user and movie.

Python
1import pandas as pd
2
3feature_columns = train_features.columns
4seen_index = pd.MultiIndex.from_frame(seen_pairs[["user_id", "item_id"]])
5candidate_index = pd.MultiIndex.from_frame(candidates[["user_id", "item_id"]])
6unseen = candidates.loc[~candidate_index.isin(seen_index)].copy()
7
8unseen["score"] = rating_model.predict(unseen[feature_columns])
9top_10 = (
10    unseen.sort_values(
11        ["user_id", "score", "item_id"], ascending=[True, False, True]
12    )
13    .groupby("user_id", sort=False)
14    .head(10)[["user_id", "item_id", "score"]]
15)

This returns up to ten recommendations per user. The scores remain unrounded when sorting, so differences between predictions can determine the ordering.

Evaluating those lists requires a ranking protocol: which movies count as relevant, which candidates are eligible, and what K should be. Recall@K measures the fraction of relevant movies recovered in the first K results. Normalized discounted cumulative gain (NDCG@K) rewards placing more relevant movies earlier in the list. The MovieLens results above measure rating error; evaluating these ranked lists is a separate step.

Build a first baseline on your own ratings data

Start with historical user–item ratings and the attributes available when each prediction would have been made. Choose an evaluation split that matches your product: predicting another rating from an existing user and recommending to a new user are different tests.

Prepare the feature tables, supply the labeled examples as context, and predict held-out ratings with Nori-6M. Compare against a simple reference, such as the item's average rating, as you refine the features. If users will see a short list, evaluate that list with a fixed candidate set and ranking metric as well.

Start with the MovieLens notebook and adapt its feature preparation to your data. The Nori quickstart covers the prediction interface. To discuss a recommendation baseline for your ratings data, contact the Synthefy team.