SuperCharged WikiAll of Wikipedia's languages, in one article

Which Wikipedia Knows Most About What? 122 Subjects, 332 Languages, One Map of Favourites

29 September 2026 · Blog
Cards for 48 Wikipedia language editions, each naming the subjects it writes about far more than Wikipedias in general do

Recently I started SuperCharged Wiki, a site that writes one English article from every language edition of a Wikipedia topic, because the editions do not say the same things and an English reader never sees the rest. Building it meant measuring how much every Wikipedia has written on what. Here is some of what turned up along the way.

Chinese Wikipedia has more on K-pop groups than any other language, Korean included. Danish Wikipedia gives Scotch whisky 40 times the attention Wikipedias give it on average. Polish Wikipedia's great love is chess grandmasters, Finnish Wikipedia's is dog breeds, Turkish Wikipedia's is beer, Armenian Wikipedia's is cryptocurrency, and Vietnamese Wikipedia has more on the emperors of China than any language but Chinese and English.

Those are some of the findings from measuring every Wikipedia on the same 122 subjects. I took each subject from an English Wikipedia category (Roman emperors, dog breeds, K-pop groups, Islamic scholars, sex toys, ski jumpers, the books of the Hebrew Bible, and 115 more), sampled up to 120 topics from it, 13,502 topics in all, and measured the article on every one of them in every Wikipedia that has one: 332 languages, more than 200,000 articles. Sizes are converted into a common unit, so that a Japanese article and a Vietnamese one can be compared, and the English articles on the same topics count as 1.00. English is ranked as one language among the 332.

Two questions come out of the same numbers. For each subject, which languages have the most? And for each language, which subjects does it write about out of all proportion to everyone else?

The Wikipedias with the most to say

Bar chart ranking the 40 language editions with the highest average amount across the 122 subjects

Averaged over all 122 subjects, Japanese Wikipedia has the most after English, at 0.40 of the English amount, then French 0.33, Chinese and Spanish 0.28, Russian 0.27, Italian and German 0.23, Arabic and Portuguese 0.16. Ukrainian, Korean, Persian, Polish, Catalan and Turkish follow at 0.10 to 0.13. The two leaders get there in opposite ways: Japanese has articles on 42% of the topics and writes them long, French has articles on 57% of the topics, the widest coverage of any language. German, among the largest Wikipedias by article count, is eighth here, because on these subjects its articles are short.

Who has the most on what

Table of the three leading languages for each of the 122 subjects

Ten languages take the lead on 28 of the 122 subjects:

  • Japanese leads 11 subjects: samurai (3.70 times the English amount), prime ministers of Japan (3.39), Japanese deities (3.32), sumo wrestlers (2.26), Japanese emperors and Sengoku battles (2.08), Go players (2.05, with Chinese at 1.96 right behind), shonen manga (1.67), sushi and Japanese dishes (1.63), anime series (1.35), Kurosawa and Ghibli films (1.28, Chinese 1.21) and yokai (1.01).
  • Chinese leads 4: K-pop groups (1.43, Korean 0.31), Chinese philosophers (1.41), Chinese emperors (1.29) and Korean dramas (1.17, Korean 0.94).
  • French leads 3: the kings of France (1.52), Tour de France winners (1.15) and horse breeds (1.02).
  • Spanish leads 2, the Latin American wars of independence (1.46) and Latin American presidents (1.21). Russian leads 2, the novels of Dostoevsky and Tolstoy (1.46) and Soviet battles (1.04). Arabic leads 2, Islamic scholars (1.17) and Arabic poets (1.06).
  • One each: Portuguese on Brazilian football clubs (1.76), German on the mountains of the Alps (1.26), Italian on the World Heritage Sites of Italy (1.07) and Malay on dog breeds (1.10).

Second place is often not the language you would guess:

  • Bengali is second on cricket (0.87), on Hindu deities (0.49) and on the Mahabharata and Ramayana (0.57), ahead of Urdu, Hindi, Tamil and Telugu.
  • Vietnamese is third on Chinese emperors (0.60) and on Soviet battles (0.59).
  • Chinese is second on dinosaur genera (0.97), elementary particles (0.92) and assault rifles (0.85).
  • Polish is second on chess grandmasters (0.83) and on Eurovision winners (0.46). Turkish is second on beer styles (0.69). Finnish is second on metal bands (0.47). Danish is second on Scotch whisky (0.52). Armenian is second on cryptocurrencies (0.40).
  • Russian has more on the Napoleonic battles (0.27) than French (0.25), and is second on Shakespeare's plays (0.38).
  • Arabic is second on the battles of the Crusades (0.70, French 0.36), on infectious diseases (0.22), mental disorders (0.27) and antibiotics (0.25), and on FIFA World Cups (0.51).
  • French is second on the pharaohs (0.93, Arabic 0.71), the Byzantine emperors (0.71), the Roman emperors (0.65, Italian 0.57), cat breeds, constellations and the national parks of the United States.
  • Spanish is second on Marvel superheroes (0.54), Vikings (0.48) and the Norse gods (0.39), a Nordic subject no Nordic language comes close on.
  • Japanese is second on Formula One drivers (0.63), Nintendo (0.61), the Pacific War (0.55), high-speed trains (0.55), car manufacturers (0.60), Linux distributions (0.41) and the Mongol Empire (0.74).
  • English has more on the Persian poets than Persian Wikipedia (0.89), and more on Turkish cuisine than Turkish (0.70).

English has the most on the remaining 94 subjects, as its size predicts.

Every Wikipedia's favourite subjects

Cards for 48 language editions, each with its three favourite subjects, the multiple and the language's rank on the subject

The cards answer the second question. A favourite is a subject a language gives a far larger share of its writing than Wikipedias give it on average: a score of 7 means 7 times the usual share. The rank says how the language stands on that subject against all 332.

Some favourites are plain enthusiasms. Polish Wikipedia gives chess grandmasters 7 times the usual attention and is second in the world on them. Finnish gives dog breeds 7.5 times (third in the world). Turkish gives beer styles 6 times (second) and Scotch whisky 5 times (fourth). Danish gives Scotch whisky 40 times the usual share and is the second-fullest Wikipedia on it. Galician Wikipedia's favourite is Formula One drivers, 15 times its norm and third in the world; Basque Wikipedia's is Tour de France winners.

Others follow history. Vietnamese Wikipedia gives the emperors of China 9 times the usual share and the Soviet battles 6.5 times. Uzbek gives Islamic scholars, the mosques of Istanbul and the Sufi mystics 11 to 13 times. Bengali gives the Hindu deities 9 times, the Sufi mystics 8 and the Islamic scholars 8, a religious mix no other language shows. Asturian Wikipedia's favourite, 17 times its norm, is Latin American presidents; Asturias sent emigrants to Cuba, Mexico and Argentina for a century. Greek gives the Byzantine emperors and the Crusades 5.5 times. Korean gives the Mongol Empire and the Japanese gods nearly 4 times, and, less expectedly, Mexican cuisine.

And some are strange. Armenian Wikipedia gives cryptocurrencies 17.5 times the usual share and is second in the world on them. Persian Wikipedia's favourites are pasta and the Tour de France. Italian Wikipedia favours yokai, the spirits of Japanese folklore, and Marvel superheroes. German Wikipedia favours the castles of Wales and Scotland, ocean liners and Scotch whisky. Hebrew Wikipedia favours the national parks of the United States. Romanian Wikipedia's favourite, 8 times its norm, is sexual acts and positions. Macedonian's is the constellations. Estonian's is the operas of Verdi. Malay Wikipedia's is dog breeds, 18 times its norm, and it has more on them than any other language.

Single articles that dwarf the English one

Bar chart of 28 single articles where a foreign-language Wikipedia is many times the size of the English article

The subject scores average over 100 topics. Single articles run further. German Wikipedia's article on the Mogilev offensive of 1944 is 28 times the English one, and its Vilnius offensive 17 times. Norwegian Wikipedia's article on the Gandavyuha, a Buddhist sutra, is 18 times the English article. Korean Wikipedia's article on Emperor Go-Daigo of Japan is 11 times the English one. Russian Wikipedia's article on the Crimean-Nogai slave raids is 12 times the English one, and French Wikipedia's on the Barb horse 13 times. Italian Wikipedia's article on La'eeb, the mascot of the 2022 World Cup, is 19 times the English one. Chinese Wikipedia's article on the J programming language is 15 times the English one. Sindhi Wikipedia, a small edition, has an article on Sehwan, the Sufi shrine town, 7 times the English article.

Home turf

Dumbbell chart comparing each home language's amount with the best other language on 74 subjects with an obvious home

Being the home language of a subject is worth a great deal in some places and nothing in others. Japanese Wikipedia has more than English on 11 of its own subjects, by up to nearly 4 times, and less on others, among them Nintendo, the Pacific War, its car makers, its high-speed trains and Buddhist thought. French has 1.5 times the English amount on the kings of France, Portuguese 1.76 on Brazilian football clubs, Spanish 1.46 on the Latin American wars of independence, Russian 1.46 on Dostoevsky and Tolstoy.

Then there are the home languages that lose. Korean Wikipedia has less than a quarter of Chinese Wikipedia's amount on K-pop groups (0.31 against 1.43) and less on Korean dramas too (0.94 against 1.17). Welsh Wikipedia has almost nothing on the castles of Wales (0.05), and Scottish Gaelic has nothing at all on Scotch whisky. No Nordic language reaches a fifth of the English amount on Vikings, and Swedish has 0.29 on the Norse gods, behind Spanish. Italian has 0.57 on the Roman emperors, 0.21 on the Greek and Roman gods, 0.49 on Verdi's operas and 0.58 on pasta. French has 0.25 on the Napoleonic battles and 0.20 on its own fashion houses. Turkish has 0.44 on the Ottoman sultans and 0.48 on the mosques of Istanbul. Arabic has 0.71 on the pharaohs and 0.73 on the Quran. Of the Indian languages, Bengali is the best on cricket (0.87) and on the Mahabharata and Ramayana (0.57), Tamil on Indian cuisine (0.37), and none reaches a third of the English amount on the Filmfare and National Film Award winners.

The sex test

Three bar charts of the seven languages with the most on sexual acts, paraphilias and sex toys

I put three sex subjects in on a hunch that Japanese Wikipedia would have far more than English. Across all 332 languages, nobody does. English leads all three by a wide margin; Chinese comes closest on sexual acts and positions (0.44) and on sex toys (0.27), Spanish on paraphilias and fetishes (0.30). Japanese sits at 0.26 on sexual acts, behind Chinese, Spanish and Ukrainian. The language that favours the subject most, relative to everything else it writes, is Romanian.

Who writes wide, who writes long

Scatter plot of every language's average coverage against its average amount, with the 34 largest labelled

Plotting how many topics a language covers against how much it writes separates the two ways of being big. Japanese sits high and to the left: articles on 42% of the topics, and the typical one about as long as the English article. Chinese is similar with shorter articles. French sits far to the right: an article on 57% of the topics, the typical one about six tenths of the English length. Spanish, Russian, Italian and German cover about half the topics with articles about half the English length. Below 0.10 lie more than 300 languages, and their differences are almost all coverage: Ukrainian, Korean and Persian have articles on about a third of the topics, Serbian on a fifth, Bengali on an eighth.

The map

Heatmap of 122 subjects against the 40 largest language editions, shaded by specialisation

The heatmap is the whole picture at once: every subject against the 40 largest languages, shaded by how far each language's share of a subject departs from the average. Read a column for a language's character (Bengali dark on the Indian and Islamic rows, Vietnamese on the Chinese and Soviet ones, Hebrew on the Bible and American parks, Turkish on beer, whisky and Byzantium) and a row for a subject's constituency (the Sufi mystics row lights up in Bengali, Uzbek, Urdu and Azerbaijani; the chess row in Polish, Ukrainian and Catalan; the Tour de France row in Serbian, Basque and Persian).

Method and caveats

Subjects. Each subject is an English Wikipedia category, with one level of subcategories where the category holds mostly subcategories, from which I drew a random sample of up to 120 member articles. Subjects with fewer than 30 sampled topics were dropped, leaving 122. Topics were matched across languages through Wikidata.

Languages. Every open Wikipedia was checked, 347 editions besides English, and 334 had at least one article on the sampled topics. Cebuano and Waray, whose articles are almost entirely machine-generated, and Simple English, which is English, are left out, leaving 331 languages plus English. Beyond that, any language's articles on a subject were set aside as a machine-made batch when their sizes varied by less than 25% across the batch, the signature of the geography bots; 70 such batches were removed, including the Swedish and Cebuano mountain articles. Swedish, Dutch, Vietnamese, Malay and Egyptian Arabic still carry many bot-era articles, and their figures should be read with that in mind.

Sizes. Article sizes were read in September 2026 as bytes of wikitext, which includes infoboxes and tables, and converted to a common unit in two steps: bytes per character, measured on a sample of each edition's own text, and a length factor for the language, the length of the Universal Declaration of Human Rights in that language against English. A language that needs more characters than English to say the same thing is not credited for them. The factor comes from the UDHR for 147 languages, from an earlier measurement for 20, from a closely related language for 32, and from a script default for 135 small editions. The English articles are the unit: 1.00.

Amount is coverage (the share of a subject's topics a language has an article on) multiplied by the median size ratio to the English article. A favourite's specialisation is the share of a language's writing across the 122 subjects that goes to one subject, divided by the average share that established languages (those with articles in at least 25 subjects, 172 of them) give it; a favourite must also be made of real articles, at least 3,000 bytes at the median, which keeps batches of stubs off the cards.

Limits. The topics come from English categories, so a subject that lives mainly in another language is under-sampled and the gaps in the other direction are larger than these numbers show. Wikitext bytes reward long infoboxes and tables as well as prose, and the conversion corrects for the language, not for the writer: a padded article counts as much as a dense one. Rankings between two small languages can move with the conversion factor; the leading positions do not. And size is not quality: this measures how much has been written, which says how much attention a subject gets in a language community and nothing more.

Data: the full results (JSON: for each subject the categories, home languages, sample size, and for every language the amount, coverage, median size ratio, share of topics fuller than English, specialisation score and bot flag; plus every language's conversion factor and its source).