A Single Open World Corpus of Knowledge

by dnb66 29.07.2026 12:00

✎ Opinion column — the author's view, not a news report.

The race to train the best model has a limit — everyone arrives at the same library of human knowledge. Instead of each maker paying anew for the same data, we need a single, legally protected and freely usable world corpus of knowledge, funded by a permanent international fund.
The race among models for the best training has a limit. Gradually, every maker will assemble the full set of data that a maximally well-informed model needs. A road on which all human knowledge is gathered always ends in one place — a library of that knowledge. There will be no winner by outcome; there will only be temporary winners by timing. States are already funding their laggards, and this process is impossible to miss. Links to several national programs are at the end of this note. What we end up with. Every maker pays anew for the right to access the very same thing, builds anew its own parsing pipeline for the very same texts, and arrives anew at the very place everyone will arrive at. Repeated costs for a single result. What is needed instead. A single worldwide corpus of knowledge. Legally protected. Free for use by anyone. And funded by an international fund — not once, but continuously, because the flood of new knowledge rolls in without pause and the corpus must be replenished without pause. Is there any sense in worldwide competition on this field, or is it already time to choose cooperation among nations? Billions upon billions would be saved and could be invested into competition on the field of AI adoption. In what form should a world fund of knowledge be created? Not weights are needed, but a formalized single source of knowledge, backed by source materials. Weights grow obsolete within a year, sometimes sooner; architectures — within two. Formalized knowledge does not grow obsolete — it accumulates. Formalization is needed so that knowledge can be compared, verified and delivered in any language: language becomes a projection, not a store. Source materials are preserved always, because extraction will never be complete, and an extracted fact is only a hypothesis about what is written in the source — and one must be able to re-check it. What this gives. In the evolution of AI for the benefit of humanity, not only ultra-large business under state support will be able to take part, but also small IT companies, individual specialists, and countries that have no large models of their own. It is there, and not at the cutting edge, that the main untapped benefit lies. Forks will be available to many and will drive the evolutionary process. Competition will not go anywhere. Let architectures, techniques and training methods compete. But not the facts themselves. The facts of humanity belong to humanity. DNB National support programs: a few examples. European Union, InvestAI (February 2025) — a mobilization of 200 billion euros, of which a new European fund of 20 billion euros for "AI gigafactories," that is, for the compute capacity to train the largest models. https://oecd.ai/en/dashboards/policy-initiatives/investai China, National AI Industry Investment Fund (April 2025) — 60 billion yuan, about 8.2 billion dollars. Founders — the Ministry of Industry and Information Technology together with the Ministry of Finance. Plus 11 national pilot zones for AI innovation. https://www.chinadaily.com.cn/a/202504/18/WS6802358ea3104d9fd38204b5.html India, IndiaAI Mission (approved by the government on 7 March 2024) — over 10,300 crore rupees for five years, around 1.2 billion dollars. Its design is telling: separate line items go to compute (over 10,000 GPUs), a center for developing homegrown foundation models — and the IndiaAI Datasets Platform, that is, a state program to collect and open datasets. Exactly what this note is about. https://www.pib.gov.in/PressReleasePage.aspx?PRID=2012375&reg=3&lang=2 South Korea, 2026 budget — 10.1 trillion won for AI, around 7 billion dollars: 2.6 trillion for adoption in industry and public services and 7.5 trillion for talent and infrastructure. Procurement of another 15,000 accelerators toward a total fleet of 35,000, a separate line for homegrown foundation models. https://www.korea.net/NewsFocus/policies/view?articleId=281702 United Kingdom, sovereign AI (June 2026) — 1.1 billion pounds on top of the 500-million-pound Sovereign AI program announced in March 2026. Of these, 750 million for a national AI supercomputer by 2030 and 400 million for procuring accelerators. The stated goal — a twentyfold increase in the country's compute capacity by 2030. https://www.computing.co.uk/news/2026/government/government-commits-more-than-one-billion-sovereign-ai France (February 2025) — 109 billion euros of announced private investment under a state umbrella, plus public money from France 2030, including 360 million euros for the "IA Clusters" program. https://www.info.gouv.fr/actualite/ia-une-nouvelle-impulsion-pour-la-strategie-nationale Note two things. First: almost all the money goes into compute and talent. Second: where a country starts from scratch, the first thing it creates is a state dataset platform — as India did. So the problem is already seen. But so far each solves it on its own. Although, to speak frankly, many do not yet even see or address the task of uniting knowledge, permitting unproductive, useless competition.
Fresh news