BusinessIssue #100 ·

North Korea Cited GPT-4 in an Official University Journal

Can Pyongyang actually build its own large language model?

North Korea Cited GPT-4 in an Official University Journal

Opening

Dear reader, an unexpected piece of material has surfaced. It’s a paper titled “A Method for Building Training Data for Query Recommendation Models in Intelligent Search Systems,” published in Kim Il-sung University’s journal Information Science, Vol. 72, No. 1, 2026. What’s interesting isn’t the body of the paper — it’s the bibliography. There, OpenAI’s “GPT-4 Technical Report” (2023) and Google researchers’ “REALM” paper (2020) are cited side by side.

Reports that North Korea was studying ChatGPT had trickled out since February 2025. But this is a different kind of discovery. In an official university journal, with proper citation format, American and Google LLM research appears as the actual starting point of the research methodology.

Let me cut to the conclusion: this isn’t a question of “how did North Korea get its hands on this.” It’s a question of “where does the moat in AI research actually remain today.” Oh, and let me state clearly upfront: our principal enemy is North Korea, and I completed my mandatory service in the Republic of Korea Army as a sergeant, finished my reserve duty, and am now enrolled in civil defense.

What Happened

The paper’s authors are two researchers, Kim Jin-beom and Han Seung-ju. The problem they set out to solve is surprisingly mundane: how do you train a model that, given a few words typed into a search box, automatically recommends a complete Chosŏnŏ (North Korea’s term for the Korean language) query sentence? It’s the same autocomplete feature we see every day on Google or Naver.

The approach the researchers proposed was based on a Transformer model, and they evaluated performance using training data built from 50,000 documents and 7 million word pairs. Academically, this isn’t a particularly novel attempt. But one sentence in the introduction stands out.

“Large language models (LLMs) such as GPT-4 and Claude-2 have been developed and are now showing high performance across most tasks in the natural language processing domain, and are being actively adopted and used in search systems as well.”

It goes on to accurately note the changes in Google, Bing, and Baidu’s search engines as well. In other words, North Korean researchers clearly recognize the global LLM landscape, cite official sources, and have publicly disclosed — at the level of an official university journal — that they are applying this to their own systems.

Two facts are worth noting here.

First, both cited papers are freely accessible to anyone. The GPT-4 Technical Report is available on arXiv (arXiv:2303.08774). The REALM paper is a peer-reviewed paper published at ICML 2020 (arXiv:2002.08909).

Second, the time lag is about 2 years. Given that the paper was submitted in November 2025 and the GPT-4 Technical Report was published in March 2023, the gap works out to roughly 1 year and 8 months. Compared to North Korea’s historical lag in AI research since the late 1990s — anywhere from 5 years at best to over 10 years at worst — that’s a substantial narrowing.