Norway's National Library is building a large language model that understands the Norwegian language and is using 2 PB of Huawei OceanStor Dorado flash storage to handle its AI training data pipeline.
Marius Husnes, Head of IT Platform at the library, discussed the project at Huawei's ID Forum 2026 in Paris. He explained that no commercial LLM provider was developing a local Norwegian language model. He argued that any country with its own language that lacks a sovereign LLM trained in that language would be at a disadvantage. A globally trained, English-speaking LLM would not know about that country's history, news, and culture described in the local language.
The Library's Role as AI Builder
Norway's Ministry of Culture tasked the National Library with building a sovereign AI because the library holds the largest digital collection of Norwegian books, newspapers, web pages, and other materials. Like many state libraries, it receives copies of every published book and broadcast content. Its legal deposit mandate extends beyond books, making it duty-bound to collect and preserve all of Norway's cultural heritage.
The library has an agreement with Norwegian newspapers that permits training on copyrighted content. Husnes stated, "No private company has this."
The library was well positioned for this work because it has been digitizing its collection since 2005. It has amassed 20 PB of unique data stored in a 3-2-1 format three copies, two media types, one off-site, totaling about 60 PB overall. The digitization process for raw text, sound, moving pictures, still images, and web content involved extensive OCR scanning and generated a large amount of metadata, along with APIs for online access.
The Storage and Pipeline Challenge
The bulk of the data resides in a digital disk plus tape archive optimized for preservation. Husnes's task was to move this data to the LLM training system. He said the bottleneck was not compute but data quality, cleaning, and pipeline throughput.
There are two main processing stages. The first stage is in-house computation, using an Nvidia DGX H200 system, a 384 core CPU cluster, and multiple Huawei OceanStor Dorado all-flash arrays, totaling 2 PB of flash capacity. This low-latency storage serves the data pipelines and training preparation.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
The pipeline includes steps for data ingestion, cleaning, deduplication, format normalization, validation, and preparation. Once data passes through the pipeline, it is sent to Norway's national supercomputer, the Sigma2 Olivia system, for actual training runs. The Olivia system is an HPE Cray Supercomputing EX system with 448 GPUs and 64,512 CPU cores, using a 5.3 PB Cray ClusterStor E1000 storage system.
One major challenge was overcoming the different needs of two storage systems. The 60 PB preservation system is optimized for durability and cost, not fast IO, and has high read latency because it is designed for infrequent access. The AI pipeline storage is designed for high-throughput, low-latency, parallel data IO. Husnes said he learned that nobody was talking about the problems involved in moving PB-scale datasets from an archive to and through an AI data pipeline. His team had to figure out how to do it themselves.
Ongoing Learning and Future Issues
The LLM training is still ongoing. Husnes summarized what his team continues to learn about:
Evaluation: There are no standard evaluation tools to assess a sovereign Norwegian LLM. The language has two written forms, multiple dialects, and historical changes. The team is building their own evaluation tool on the fly.
Governance: Who controls access to a sovereign LLM? Who decides what it can be used for? These are institutional and political questions with no easy answers.
Orchestration: Making three systems the preservation archive, the on-prem AI environment, and the national Sigma2 supercomputer work smoothly together is an ongoing project.
The takeaways here are that Huawei storage is playing a serious and significant role in the European market, and that any country developing a sovereign, local language LLM would do well to consult with Husnes and learn what is involved.
As Husnes put it, Norway is a small country solving a problem every non-English-speaking nation will face: how do you build AI that reflects your language, your culture, and your history? AI needs custodians, not just builders.
Related on Neura Market
- AI Models Directory, Browse other large language models and training resources.
- AI Tools Marketplace, Discover tools for AI data pipelines and storage management.
- Automation Workflows, Explore automation solutions for data ingestion and preprocessing.

