Where does the flavor actually come from?
Pick up a plate of paella. By weight it is rice, chicken, beans, water. The rosemary in it — the thing you smell from the doorway — is 0.06% of the dish’s mass. This piece is about that gap: the distance between where a dish’s mass is and where its flavor signal is, measured across every dish in Palate’s corpus.
The arithmetic is the engine’s own aggregation weight — log1p(grams) × (potency + 1)^1.5 — applied to 27,229 linked ingredient lines across 2,343 dishes (0.7% of lines were unlinked or massless and skipped; 1,665 of 1,729 distinct ingredient names link exactly). An ingredient’s signal share is its weight over the dish’s total; its mass share is grams over grams.
Three dishes, opened up
Spaghetti carbonara. The richness is 37% pancetta, 34% egg yolk, 20% pecorino. The spaghetti — 44% of the mass — contributes 0%.
| ingredient | share of richness | share of mass |
|---|---|---|
| pancetta | 37% | 22.0% |
| yolk | 34% | 12.9% |
| pecorino romano cheese | 20% | 7.3% |
| eggs | 9% | 12.9% |
French onion soup. The umami is 43% beef broth and 38% gruyère. The onions the dish is named for: 4%, from 36% of the mass. (What caramelization does to those onions is real, and is exactly what a composition model undersells — see the limits.)
| ingredient | share of umami | share of mass |
|---|---|---|
| beef broth | 43% | 37.6% |
| gruyere | 38% | 12.5% |
| dry white wine | 10% | 6.3% |
| fresh yellow onions | 4% | 36.0% |
| season salt | 2% | 0.08% |
| unsalted butter | 2% | 1.4% |
Paella Valenciana. The aromatic channel is 52% garlic and 40% rosemary — rosemary being 0.06% of the dish by mass. A 600-fold mismatch between where the weight is and where the smell is.
| ingredient | share of aromatics | share of mass |
|---|---|---|
| garlic | 52% | 0.38% |
| rosemary | 40% | 0.06% |
| sweet paprika | 9% | 0.15% |
The corpus-wide picture
Measured as the median across all recipes each appears in: cayenne punches 30× over its mass share (n=183 recipes), rosemary 35× (n=36), thyme 23× (n=132), saffron 26× (n=25), shrimp paste 39× (n=19). At the other end, white rice returns about 0.1× its mass share in signal (n=68) and water — 19% of a typical dish it appears in — carries essentially nothing (n=325).
Aggregated: in the median dish, the ingredients that are each under 5% of the mass together carry 49% of the flavor signal (interquartile range 33%–60%, n=2,343 dishes). Half of what a dish is, by this model, comes from things you could hold in one cupped hand. None of this will shock a cook — nobody seasons by the pound. What the measurement adds is how big the effect is, and which ingredients live furthest from the diagonal.
| ingredient | leverage (signal ÷ mass) | median mass share | n recipes |
|---|---|---|---|
| shrimp paste | 39× | 0.26% | 19 |
| cayenne | 36× | 0.10% | 27 |
| rosemary | 35× | 0.11% | 36 |
| cayenne pepper | 30× | 0.11% | 183 |
| marjoram | 29× | 0.08% | 15 |
| saffron | 26× | 0.08% | 25 |
| cardamom | 25× | 0.22% | 19 |
| thyme | 23× | 0.06% | 132 |
| nutmeg | 20× | 0.07% | 58 |
| dark sesame oil | 20× | 0.32% | 124 |
| pepper | 19× | 0.08% | 65 |
| coriander | 19× | 0.09% | 107 |
Honest limitations
- This is a composition model, not a perception model. Real flavor perception runs on concentration ÷ odor threshold (the odour-activity-value framework of Grosch and colleagues), not gram share with a potency correction. A mass-weighted model systematically under-weights high-impact aromatics — saffron’s true perceptual leverage is likely higher than shown, not lower. Treat the leverage numbers as floors on the phenomenon.
- The potency term is a hand-labeled partial correction, not a measurement. No human palate has ever validated one of these vectors against tasting the dish — the attribution is exact arithmetic on the engine’s model; the model itself is the untested part.
- Cooking transforms contributions. Caramelized onions give more than raw-onion labels suggest; volatiles cook off; fat carries flavor it doesn’t own. The attribution here is computed on raw composition.
- ~40% of the corpus is project-authored (934 of 2,345 dishes), which may use more regular quantities than recipes in the wild.
Reproduce this
Deterministic, offline, no model calls. Every number on this page is read from the script’s results file — the page cannot state a number the script didn’t produce.
python analysis/instrument_room/scripts/piece2_attribution.py python analysis/instrument_room/scripts/piece2_figures.py
Inputs: prototype/web/src/data/sample_corpus_v17.json + data/recipe_corpus/ingredient_vectors_v2.json (read-only). Every number on this page is in analysis/instrument_room/results/piece2_attribution.json.