AI Models

Talkie: 13B Language Model from Pre-1931 Texts

Researchers Nick Levine, David Duvenaud, and Alec Radford released talkie, a 13B language model trained on historical English text before 1931. The base version used 260B tokens, while the instruction-tuned model powers a chat interface. Both models carry an Apache 2.0 license, with training data free from copyright restrictions.

Neura News

Neura News

Neura Market Editorial

April 28, 20263 min read

Originally reported by simonwillison.net

Talkie: 13B Language Model from Pre-1931 Texts

Talkie: 13B Language Model from Pre-1931 Texts

Nick Levine, David Duvenaud, and Alec Radford, known for their work on GPT, GPT-2, and Whisper, launched a new project called talkie. This includes two models: talkie-1930-13b-base at 53.1 GB and talkie-1930-13b-it at 26.6 GB. The base model consists of a 13B language model trained on 260B tokens from historical pre-1931 English text. Alec Radford brings experience from OpenAI, where he contributed to early generative models like GPT and GPT-2, and later to Whisper for speech recognition. David Duvenaud works in machine learning research, often focusing on probabilistic models and neural networks through affiliations like the University of Toronto and the Vector Institute.

Model Specifications and Access

The instruction-tuned version, talkie-1930-13b-it, comes from a checkpoint fine-tuned on a dataset of instruction-response pairs pulled from pre-1931 reference works. This setup supports a chat interface. Users can access a demo of this chat model online. Both models operate under the Apache 2.0 license. The base model's training data falls entirely out of copyright, given the U.S. cutoff date of January 1, 1931. Expectations exist for the team to release this training data later.

Training Process and Research Goals

The project's report outlines key research aims for models like this. Simon Willison, a prominent figure in data tools and LLM commentary through his Datasette project, expressed interest in what he terms "vegan models." These are large language models trained solely on licensed or out-of-copyright data. The base talkie model fits this category. However, the chat model relies partly on modern models during fine-tuning, so it does not qualify fully.

Fine-tuning started with generating instruction-response pairs from structured historical texts. Sources included etiquette manuals, letter-writing manuals, cookbooks, dictionaries, encyclopedias, and collections of poetry and fables. The base model then trained on these in a simple chat format.

To boost instruction-following, synthetic prompts covered tasks like document summarization, direct information responses, and multi-turn conversations. Online direct preference optimization applied to rollouts from these prompts, judged by Claude Sonnet 4.6.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

A final supervised fine-tuning round used rejection-sampled multi-turn synthetic chats between Claude Opus 4.6 and talkie. This addressed remaining issues in conversation skills.

Challenges with Contamination and Future Aspirations

A primary hurdle involved preventing contamination. This meant keeping post-1931 text out of the training corpus and stopping modern LLMs from adding anachronistic knowledge during fine-tuning.

Willison recently covered Mr. Chatterbox, another vintage language model project. It faced similar needs to use modern LLMs for synthetic conversations in chat training.

The talkie developers aim to overcome this. They note that reinforcement learning with AI feedback shapes behavior in modern ways. For example, the 7B version of talkie began using listicle styles after such training. Plans call for scaling up and using the vintage base models as judges. This would create a fully bootstrapped, era-appropriate post-training pipeline.

Demo Performance

In testing the talkie demo, Willison used his standard prompt: "Generate an SVG of a pelican riding a bicycle." The model replied: "An SVG of a pelican riding a bicycle was generated in 1860. The pelican was represented as perched on the saddle, with its long bill pointed forward, and its legs astride the handles. The idea was suggested by the fact that pelicans are sometimes seen fishing on horseback, on the banks of the Rhine."

This release highlights efforts to build language models rooted in historical data, avoiding modern copyrights and influences where possible.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read