AI Models

AI Model Separates Recipe Pairs from Flavor Profiles

Kaikaku.AI introduces Epicure, three nearly identical AI models trained on different data sources. One learns from recipes, another from flavor molecules, and a third blends both. The models reveal how training data choice changes answers to simple questions like what goes with chicken.

Neura News

Neura News

Neura Market Editorial

May 31, 20265 min read

Originally reported by the-decoder.com

AI Model Separates Recipe Pairs from Flavor Profiles

Chicken and Basil Get Different Answers from Different AI Models

A new research project from startup Kaikaku.AI shows that the answer to a simple question like what ingredients go with chicken depends entirely on whether an AI was trained on recipe pairs or molecular similarities. Previous models mixed both perspectives together. The new models keep them separate.

Jakub Radzikowski and Josef Chen created three nearly identical AI models under the name "Epicure." Each model differs only in its training data. The first model, called "Cooc," learned from which ingredients appear together in real recipes. The second model, "Chem," learned from the flavor molecules that ingredients share, drawing on the FlavorDB chemistry database. The third model, "Core," blends both sources.

Three Models Give Three Different Answers

The differences become clear with specific queries. When asked about chicken, Cooc returns garlic, onion, and black pepper, ingredients that frequently appear alongside chicken in recipes. Chem returns beef or pork, ingredients with a similar flavor profile. For basil, Cooc suggests parsley, olive oil, and parmesan, the typical pasta pantry lineup. Chem suggests oregano, tarragon, and rosemary, the herb relatives.

The chemistry-driven model also performed better in areas where it should not have had any information. Flavors like sweet, sour, or bitter and nutritional values like protein or fat content are not directly coded in the training data. Yet Chem classifies ingredients along these axes more clearly than the other variants. The chemical relationships appear to act as a shortcut that also tunes the model to other culinary concepts.

Multilingual Corpus Reduces English Language Bias

The most complete public ingredient model to date, FlavorGraph, is built on an English language recipe corpus. Epicure processes 4.14 million recipes from eleven sources in seven languages. These include Chinese, Russian, Vietnamese, Turkish, Indonesian, and German. A pipeline built on Claude and Gemini embeddings translates and cleans about 200,000 raw terms, such as spelling variants, brand names, and preparation instructions, into 1,790 clean ingredients.

The corpus remains unevenly distributed. About half the material comes from East Asian sources. Latin American, Eastern European, and South Asian cuisines each contribute single-digit percentages. Only about a third of the ingredients are directly anchored in the chemical database. The rest pick up the chemical signal indirectly through related ingredients.

Two Operation Modes for Ingredient Exploration

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Two modes of operation run on the finished model. The first is a simple neighbor search: which ingredients are closest to a given one? The second lets users shift a seed ingredient by an adjustable angle toward a target direction. At zero degrees, the original stays untouched. At sixty degrees, the target neighborhood takes over.

Turn rice slightly toward South Asia, and results include curry leaf, urad dal, chana dal, and fenugreek seeds. Turn chicken more toward processed Western Atlantic cuisine, and you get cream of chicken soup, crescent rolls, and ranch dressing, typical US home cooking staples.

The model choice can even decide which culture an answer comes from. Turn chocolate in the direction of sweet pastries, and Cooc and Core land on Western baking ingredients like cocoa, vanilla, and baking powder. Chem lands on an East Asian dessert cluster with red bean paste, matcha powder, and purple sweet potato.

Startup Behind the Research Runs Robot Restaurants

Kaikaku was founded in London in 2023 and operates its own robotic restaurant, Common Room, in the Brunswick Centre. The company plans to expand it into a chain. Kaikaku uses its own machine learning systems to weigh and portion ingredients. Its machine called "Fusion" can theoretically dispense 360 bowls per hour. The system also includes ML powered inventory management and 3D printed food safe components. The company raised about $1.8 million in a pre seed round in 2024.

Given that background, the interest in a machine readable map of the ingredient world makes sense. A model that switches between recipe companions and flavor relatives on demand, translates ingredients across cuisines, or shifts them along axes like fatty or fermented would be useful in several places. It could help with menu development at a bowl restaurant, suggest replacements during supply shortages, or assist when scaling to new locations.

Caveats and Open Questions

Whether this works in practice remains to be seen. Model weights and datasets are now available on Hugging Face, making independent verification possible in principle. But the examples shown in the paper are hand picked. In sparsely represented regions like South Asia or Latin America, the answers are likely far less stable than for the dominant East Asian and Western cuisines.

The vocabulary cleanup also depends on the output of language models, which carry their own cultural biases. The fact that chocolate ends up near matcha in one model variant's sweet pastry direction is a nice effect. But it says little about how reliably such rotations work beyond the cherry picked examples.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read