Meet Nori. The open-source model for structured data.
AI & Machine LearningResearch8 min read

Nori + Jev: Combining In-Context Learning with Semantic Understanding

By

A simple Jev feature lowers Nori's restaurant-rating RMSE by 5.2–6.1% in our case study, combining semantic judgments with learning from labeled context.

#Nori#Jev#In-Context Learning#Tabular Foundation Models#Rating Prediction
Jev interprets each row and adds a feature to the context and query table. Nori uses the table and observed context ratings to predict the query rating.
The same Jev feature is added to context and query rows. Nori learns how it relates to observed ratings. Rows shown are illustrative.

Introduction

Jev, from TypeSafe AI, is a System One model designed for fast, probabilistic classification and decision-making. It takes natural-language input and returns typed choices, scores, and probabilities over predefined outcomes. That makes it useful for the kinds of judgments software needs to make repeatedly: interpreting a customer message, routing a request, or deciding whether an item meets a criterion.

You can apply Jev directly, zero-shot, to many prediction problems that we attack with tabular foundation models like Nori. But we believe that decision models like Jev and tabular models like Nori have complementary strengths. Jev contributes language understanding and task-specific semantic judgments. Nori learns how the available features relate to observed outcomes in the supplied dataset.

Our north star is a model that combines both: it understands the real-world context surrounding a prediction and learns complex relationships through in-context learning. We expect semantic knowledge to remain useful even as the amount of labeled context grows, because it supplies information beyond what examples alone can reveal. In this post, we test a simple step toward that goal: use Jev to produce one additional feature for Nori, then measure whether the combination improves on both models individually.

Comparing and contrasting Nori and Jev

Start with the inputs and outputs. What is each model trained to look at, and what is it trained to produce?

Nori learns structured relationships from context

Nori is trained for structured-to-structured prediction. Numerical and encoded categorical features go in, alongside labeled context rows; numerical predictions and predictive uncertainty come out. Its pretrained on large amounts of synthetic tables, generated through structural causal models and other synthetic processes. These generators expose the model to nonlinear relationships, correlated features, missing values, noise, and other patterns that occur in real data.

The result is the capability to infer relationships within a new table, at inference time, in a single forward pass. For a restaurant-rating task, Nori can learn how price, cuisine, location, and other features relate to ratings in the supplied examples. Those relationships can change from one dataset to another without requiring a new set of model weights.

Nori does not by itself attach ordinary language meaning to a number or column name. A value of 5 could be a restaurant rating, a waiting time, or a price. It's the examples that establish how that value relates to the prediction target. A text encoder can provide additional semantic features, but without pretraining on large amounts of text, Nori does not develop an internal "state" that combines text and numerical features.

Jev turns language into structured judgments

Jev is designed for unstructured-to-structured prediction: natural-language descriptions and application context go in; constrained choices, scores, and probabilities come out. Sebastian Raschka's overview reports that Jev is pretrained on fully curated synthetic decision examples. TypeSafe's AI primer describes its approach as adapting pretrained language models through reinforcement learning for calibrated decisions. The detailed training recipe and the data used at each stage are not public.

This gives us a useful distinction between two kinds of prior knowledge. Language understanding supplies associations between words, concepts, and their usual meanings: five out of five stars expresses a high rating; a review describing attentive service provides evidence about a customer's experience. The numerical model learns how the features we supply map onto ratings in this particular population. Both forms of knowledge can matter to the same prediction.

Jev returns a judgment without generating a written explanation or chain of thought. TypeSafe trains it to produce calibrated probabilities; that is a training objective, and calibration still needs to be evaluated for a particular use case. Our experiment measures rating accuracy rather than probability calibration.

Combining semantic knowledge with in-context learning

Consider a restaurant review that praises the food but complains about a long wait. A semantic model can interpret both parts of that description in light of a rating question. A tabular model can learn how such a judgment relates to actual aggregate ratings, together with the restaurant's other attributes. More labeled rows can improve that mapping while leaving the semantic judgment useful.

This is our hypothesis: semantic information should continue to contribute as labeled context grows. We test it by holding the query restaurants fixed and increasing the number of examples supplied to Nori.

TypeSafe also recommends using Jev probabilities as features in downstream machine-learning models. We test that composition with a model that learns from examples at inference time.

The surrounding system can maintain history and update its inputs over time. Jev's state is the information an application supplies in each request; that interface does not imply that the model itself remembers previous requests. In this experiment, Jev scores each row independently, and Nori adapts through the labeled context supplied at inference time.

Real case study: combining Nori and Jev to predict restaurant ratings

We use version 1 of the public MulTaBench Zomato restaurant dataset. The task is to predict a restaurant listing's existing aggregate rating from its structured attributes and review text. This is a regression problem: the outcome is a decimal rating, and we want an accurate numerical prediction.

We hold out 4,151 listings from 948 restaurant outlets. Listings from the same outlet stay together, so an outlet cannot appear in both context and query. From the remaining 37,426 listings, we draw nested contexts of 512, 2,048, 8,192, and 32,768 rows, using three sampling seeds. Every method is evaluated on the same query listings.

We remove numeric scores attached to individual reviews and redact explicit rating expressions from the text. The reviews still describe customer experiences that contributed to the existing aggregate rating. This evaluates reconstruction of a rating snapshot, rather than a forecast of future restaurant quality.

Four approaches with shared inputs

We use the released Nori 6M checkpoint and Jev 1.13.0, with four approaches:

  1. Nori with structured features: numerical and categorical listing attributes, with missing-value indicators.
  2. Nori + text: the same attributes plus fixed MiniLM embeddings of the prepared text, reduced to 64 dimensions. This is our primary Nori baseline.
  3. Jev alone: a compact view of one listing and a fixed question asking for its existing aggregate rating, with options 1, 2, 3, 4, and 5. It receives no labeled demonstrations or other rows.
  4. Nori + text + Jev: the primary Nori baseline with one additional column derived from that same Jev response.

Both text-aware Nori approaches receive the same prepared text as Jev. The comparison therefore measures the benefit of adding a task-specific Jev judgment to a baseline that already has a pretrained language representation.

We use a fixed budget of 1,024 input tokens per Jev request, including instructions and rating options. That is our experiment setting, not Jev's advertised context capacity; the results describe this compact, zero-shot setup.

The hybrid is one extra column

Jev's Choice response includes probabilities for the five rating options. Because the target is continuous, we use the probability-weighted expected rating as its primary prediction:

text
1expected_rating = sum(rating * probability[rating] for rating in 1..5)
2                  / sum(probability[rating] for rating in 1..5)
3
4jev_feature = (expected_rating - 1) / 4

The denominator accounts for rounding in the returned probabilities. The feature normalization is a fixed change of units. It does not fit a calibration model or use any restaurant's observed rating.

We compute this feature once for each required context and query row, using the same question and hiding that row's target from Jev. Then we append it to Nori's existing feature table:

text
1context_features = [structured attributes, text embeddings, Jev feature]
2query_features   = [structured attributes, text embeddings, Jev feature]
3
4predictions = Nori(context_features, context_ratings, query_features)

This is pseudocode for the data flow, rather than an SDK call. Nori receives the observed context ratings and learns how to use the extra feature alongside the other columns. Jev does not receive Nori's predictions or make the hybrid's final prediction. Neither model's weights are updated for this task.

The gain persists as context grows

Our primary metric is root mean squared error (RMSE) in rating points. It penalizes large rating errors more heavily than small ones. We also report mean absolute error (MAE), the average absolute distance between a prediction and the observed rating. Lower is better for both.

Internal Zomato results: mean RMSE in rating points over three fixed context draws, evaluated on the same 4,151 query listings. Lower is better; percentage reductions use unrounded means.
Labeled context rowsNori structuredNori + textJev aloneNori + text + JevRMSE reduction vs Nori + text
5120.34200.32000.78590.30056.08%
2,0480.33550.30960.78590.29185.74%
8,1920.33710.30090.78590.28465.42%
32,7680.35200.30310.78590.28725.22%

The hybrid improves on Nori + text at every tested context size and in all 12 seed/context pairs. At 32,768 labeled rows, mean RMSE falls from 0.3031 to 0.2872, a 5.22% reduction. The benefit remains present at the largest context we tested.

MAE improves as well. At 32,768 rows it falls from 0.2037 to 0.1905 rating points. Across the four context sizes, the mean MAE reduction is 0.0132–0.0153 points. The downloadable results include both metrics for every condition and the paired uncertainty intervals.

All 12 unadjusted paired 95% bootstrap intervals for the RMSE difference favor the hybrid. We resample whole restaurant outlets to keep repeated listings together. These intervals are conditional on the three fixed context draws and this query split; they do not describe uncertainty across new datasets.

Jev's standalone result is constant across context sizes because it never sees those labeled rows. Its expected-rating RMSE is 0.7859. Taking the literal integer choice instead gives 0.8644 RMSE. Both are less accurate than Nori here, yet the same Jev signal improves Nori when supplied as a feature. A signal can be useful to a predictor even when it is a poor standalone estimate of the target.

What this tells us, and what remains open

The result supports a concrete claim: a simple Jev feature improves an already text-aware Nori predictor on this rating dataset, and the improvement persists through 32,768 labeled context rows. It is evidence for useful complementary information under this setup. It does not isolate semantic world knowledge as the cause of the gain.

This is also a revised experiment on a query split we had already examined. Our original Jev question asked for sentiment and mapped its score to stars. We replaced it with the direct rating question above and reran the matched comparison; the longer question also changed some shared-text truncation. The current results are not an untouched confirmatory test. Nine query listings share identical raw review text with a different context outlet, and redaction cannot guarantee removal of every possible rating expression.

We also tested the same feature-augmentation idea on Online Retail II, selecting products using predicted next-month demand. There the hybrid helped at two context sizes and hurt at two; structured-only Nori had the best mean selection outcome at every size. The naive combination is not a universal improvement.

For the restaurant task, the combination requires just one additional feature. Jev supplies a semantic judgment for each row, and Nori learns its relationship to observed ratings. Our longer-term goal is a model that brings both capabilities together directly: understanding the meaning of the inputs while adapting to the relationships in the supplied data.