AI Models

Talkie: 13B Language Model from Pre-1931 Texts

Researchers Nick Levine, David Duvenaud, and Alec Radford released talkie, a 13B language model trained on historical English text before 1931. The base version used 260B tokens, while the instruction-tuned model powers a chat interface. Both models carry an Apache 2.0 license, with training data free from copyright restrictions.

Neura News

Neura News

Neura Market Editorial

April 28, 20263 min read
Talkie: 13B Language Model from Pre-1931 Texts

Talkie: 13B Language Model from Pre-1931 Texts

Nick Levine, David Duvenaud, and Alec Radford, known for their work on GPT, GPT-2, and Whisper, launched a new project called talkie. This includes two models: talkie-1930-13b-base at 53.1 GB and talkie-1930-13b-it at 26.6 GB. The base model consists of a 13B language model trained on 260B tokens from historical pre-1931 English text. Alec Radford brings experience from OpenAI, where he contributed to early generative models like GPT and GPT-2, and later to Whisper for speech recognition. David Duvenaud works in machine learning research, often focusing on probabilistic models and neural networks through affiliations like the University of Toronto and the Vector Institute.

Model Specifications and Access

The instruction-tuned version, talkie-1930-13b-it, comes from a checkpoint fine-tuned on a dataset of instruction-response pairs pulled from pre-1931 reference works. This setup supports a chat interface. Users can access a demo of this chat model online. Both models operate under the Apache 2.0 license. The base model's training data falls entirely out of copyright, given the U.S. cutoff date of January 1, 1931. Expectations exist for the team to release this training data later.

Training Process and Research Goals

The project's report outlines key research aims for models like this. Simon Willison, a prominent figure in data tools and LLM commentary through his Datasette project, expressed interest in what he terms "vegan models." These are large language models trained solely on licensed or out-of-copyright data. The base talkie model fits this category. However, the chat model relies partly on modern models during fine-tuning, so it does not qualify fully.

Fine-tuning started with generating instruction-response pairs from structured historical texts. Sources included etiquette manuals, letter-writing manuals, cookbooks, dictionaries, encyclopedias, and collections of poetry and fables. The base model then trained on these in a simple chat format.

To boost instruction-following, synthetic prompts covered tasks like document summarization, direct information responses, and multi-turn conversations. Online direct preference optimization applied to rollouts from these prompts, judged by Claude Sonnet 4.6.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

A final supervised fine-tuning round used rejection-sampled multi-turn synthetic chats between Claude Opus 4.6 and talkie. This addressed remaining issues in conversation skills.

Challenges with Contamination and Future Aspirations

A primary hurdle involved preventing contamination. This meant keeping post-1931 text out of the training corpus and stopping modern LLMs from adding anachronistic knowledge during fine-tuning.

Willison recently covered Mr. Chatterbox, another vintage language model project. It faced similar needs to use modern LLMs for synthetic conversations in chat training.

The talkie developers aim to overcome this. They note that reinforcement learning with AI feedback shapes behavior in modern ways. For example, the 7B version of talkie began using listicle styles after such training. Plans call for scaling up and using the vintage base models as judges. This would create a fully bootstrapped, era-appropriate post-training pipeline.

Demo Performance

In testing the talkie demo, Willison used his standard prompt: "Generate an SVG of a pelican riding a bicycle." The model replied: "An SVG of a pelican riding a bicycle was generated in 1860. The pelican was represented as perched on the saddle, with its long bill pointed forward, and its legs astride the handles. The idea was suggested by the fact that pelicans are sometimes seen fishing on horseback, on the banks of the Rhine."

This release highlights efforts to build language models rooted in historical data, avoiding modern copyrights and influences where possible.

Related on Neura Market

More from Neura News

Product Launch

Acer Unveils Veriton RI110 Mini Workstation for Local Agentic AI

Acer unveiled the Veriton RI110 AI Mini Workstation on September 2, 2026, in Berlin. This compact desktop, featuring an Intel Core Ultra X7 processor and Intel Arc B390 graphics, supports local inference of AI models up to 120 billion parameters. It is designed for hybrid agentic AI workloads, combining local processing with cloud resources, and includes the Qubi Claw software suite for secure, autonomous AI tasks. The system offers up to 96 GB of LPDDR5X memory, 4 TB of SSD storage, and extensive connectivity options including OCuLink, Wi-Fi 7, and dual LAN ports. Availability begins in North America in Q4 2026 and EMEA in Q1 2027.

Sep 2·4 min read