Russian word list data, Free Tool Explorer ========================================== These files are the Russian word lists used by the Free Tool Explorer word tools (language code ru). They are open data, built offline by "npm run words:build -- --lang=ru" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters а б в г д е ж з и й к л м н о п р с т у ф х ц ч ш щ ъ ы ь э ю я. Yo is folded to ye (the display spelling keeps it); the short i, hard sign and soft sign stay distinct letters; stress marks are removed. Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. OpenCorpora dictionary (dict.opcorpora.txt.bz2), 2024-04-19 build (Internet Archive capture of 2024-04-23) Used for: validity list (main tier) Source: https://web.archive.org/web/20240423122601/https://www.opencorpora.org/files/export/dict/dict.opcorpora.txt.bz2 SHA-256: e09b29b17864c50c805eb2650d9cb0d9b8b7cfe1c6b93bd28c5e252b3bc1b3da Licence: CC-BY-SA-3.0 Credit: Russian word forms: OpenCorpora (https://opencorpora.org/), the dictionary of the OpenCorpora project, licensed CC BY-SA 3.0. main/*.txt, answers/*.txt, display.json and offensive.txt are modified, filtered derivatives of it (only the headwords, with the proper-noun, abbreviation and error entries removed) and are distributed under CC BY-SA 4.0, which section 4(b) of CC BY-SA 3.0 allows for adaptations. The legal code of both licences is shipped as CC-BY-SA-3.0.txt and CC-BY-SA-4.0.txt. Changes: - Read the dump (one lexeme per block, one form per line with its grammemes, UTF-8) as a stream. Kept only the first line of every block, which is the headword (lemma) of the lexeme: the other inflected forms are not shipped. - Kept only the lexemes whose part of speech is a dictionary headword (nouns, full adjectives, infinitives, adverbs, numerals, pronouns, predicatives, prepositions, conjunctions, particles, interjections) and dropped the lexemes of the finite verb forms, gerunds, participles, short adjectives and comparatives, which are forms of another headword, except 14 ordinary headwords that the dump files only that way (rad, dolzhen, gorazd, kakov, takov, nameren, naslyshan, vsyak, lyub and five comparative adverbs). Every short adjective or comparative that stands alone in the dump, with no full adjective before it (59 of them), was decided one by one: 8 are headwords of the dictionaries and are kept, 51 are forms and stay out. Dropped every lexeme whose lemma line has the grammeme Name (given names), Surn (surnames), Patr (patronymics), Geox (place names), Orgn (organisations), Trad (trademarks), Abbr (abbreviations), Init (initials), Erro (errors) or Dist (distortions). Lowercased the words (the dump writes them in capitals) and kept only entries of 2 to 15 letters of the 33-letter Russian alphabet (so entries with a hyphen, apostrophe, digit or space are dropped). Folded yo to ye for matching (display.json keeps the spelling with yo) and removed the duplicates that this folding creates. - Chose the answers tier among the headwords: common nouns in the nominative singular (and pluralia tantum, in the nominative plural) that are also among the 50,000 most frequent FrequencyWords tokens, without the colloquial-register (Infr) and slang (Slng) entries, without the offensive words and without a policy list of mild insults, ethnic terms, sensitive vocabulary and names (scripts/words/langs/ru.ts). This is a curation choice, not a rule of any word game. - Used the paradigms of the dump (every form of a lexeme with its lemma) to expand the blocked words: every lexeme that has a blocked word among its forms is blocked in full, all its forms and its headword. The forms are not shipped except as entries of offensive.txt. The dump files the forms of one verb as several lexemes (finite forms, infinitive, participles, gerunds) and the short forms and comparatives of an adjective as lexemes of their own, and links none of them; they stand next to each other, so the lexemes of the family of a blocked infinitive or full adjective are blocked in full as well, after a check that they share their first letters with the family, and, the other way round, every verb or adjective family that has a blocked lexeme or a blocked form has its infinitive or full adjective decided (blocked, left out on purpose, or accepted with a reason). Left out of the blocklist the lexemes that are ordinary words, a reviewed list in scripts/words/langs/ru.ts. Note: The project site opencorpora.org returned HTTP 521 when this recipe was written, so the dump is taken from the Internet Archive capture of its file; the newest dump on the project's page (2026-02-01) is not archived. The file is the 2024-04-19 build (last modified 19 April 2024 on the project server): 391,843 lexemes and 5,141,279 forms by the recipe's own count (the project's later page lists a build of 2024-07-06 with 391,844 lexemes). The page footer at the project site names CC BY-SA 3.0 as the licence of the dictionary. Secondary sources say the dictionary is derived from the AOT dictionary (LGPL) and Zaliznyak's grammatical dictionary; that chain of title is not confirmed and is recorded here as a risk for counsel (plan risk table). 2. FrequencyWords, Russian (content/2018/ru/ru_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/ru/ru_50k.txt SHA-256: 6095f507cc167488ec66ada5a85ac50433503a08ad24a07c6eabdf54352c4e7f Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form (yo to ye, so that the counts of the two spellings are added together) and dropped those with letters outside the Russian alphabet or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list. - Used the tokens to restrict the nominative-singular common nouns among the OpenCorpora headwords to common words for the answers tier; because that selection is made with this list, answers/*.txt is also shared under CC BY-SA 4.0 (all the Russian word data is). Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. wordfreq is not used: its author says extracting its data into another form is not compliant. 3. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), Russian, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/ru SHA-256: 2598adabd2b761733d72a68d8ee4fcc1d2c38f2f85b6ff489985c6ce1d93bf75 Licence: CC-BY-4.0 Credit: LDNOOBW, the Russian list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries, folded to the stored form, and expanded them through the OpenCorpora paradigms: a lexeme that has a listed word among its forms is blocked with all its forms. The ordinary words that coincide with a listed form are removed from the result (see the notes of the OpenCorpora entry). Phrases were dropped. 4. safetext (DeepSafe), Russian profanity list (safetext/languages/ru/words.txt), commit 215a6a52d8bca9dd36a10dc18ee80e70d32afeb7 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/viddexa/safetext/215a6a52d8bca9dd36a10dc18ee80e70d32afeb7/safetext/languages/ru/words.txt SHA-256: c6eaf89a9855cb76cd17faf2e886b6e9610f194c9a24396dafcd564a664acbee Licence: MIT Credit: safetext by DeepSafe (https://github.com/viddexa/safetext), MIT licence, Copyright (c) 2023 DeepSafe. The licence text is shipped as MIT-safetext.txt. The list is an aggregate whose upstream sources are not stated; LDNOOBW is credited as well. Changes: - Kept the single-word entries, folded to the stored form, and expanded them through the OpenCorpora paradigms like the LDNOOBW entries (a lexeme that has a listed word among its forms is blocked with all its forms, unless it is on the reviewed list of ordinary words). Phrases were dropped. Entries that are not words of the dictionary (they match no OpenCorpora form) are blocked as written, as whole words only. 5. Creative Commons Attribution-ShareAlike 3.0 Unported, legal code, 3.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/3.0/legalcode.txt SHA-256: 3f941b3b89cf7b8370ceb83cc76d2120d471b58735d8ca60238a751a48d7f72f Licence: CC-BY-SA-3.0 Credit: Licence text of the OpenCorpora dictionary, as published by Creative Commons. 6. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data and of the modified OpenCorpora derivative, as published by Creative Commons. 7. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. 8. MIT licence of safetext (the LICENSE file of the repository), commit 215a6a52d8bca9dd36a10dc18ee80e70d32afeb7 Used for: licence text Source: https://raw.githubusercontent.com/viddexa/safetext/215a6a52d8bca9dd36a10dc18ee80e70d32afeb7/LICENSE SHA-256: 4541b3066d5a63fcd260245287077d8a1db2bfe3e99fd09c2947525919956d95 Licence: MIT Credit: Licence text and copyright notice of safetext. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is the headwords of the OpenCorpora dictionary (the 2024-04-19 build, 391,843 lexemes): the first form of every lexeme whose part of speech is a dictionary headword (nouns, full adjectives, infinitives, adverbs, numerals, pronouns, predicatives, prepositions, conjunctions, particles, interjections), without the lexemes marked as names, surnames, patronymics, place names, organisations, trademarks, abbreviations, initials, errors or distortions. The dump files every other form of a verb (finite forms, gerunds, participles), the short adjectives and the comparatives as lexemes of their own: those first forms are forms of another headword, not headwords, so they are not in the list, except 14 ordinary headwords (rad, dolzhen, gorazd, kakov, takov, nameren, naslyshan, vsyak, lyub and five comparative adverbs) that the dump files only that way: every short adjective or comparative that no full adjective stands before is decided one by one (8 headwords kept, 51 forms left out), so the build stops on a new one. Only the headword is kept: the other inflected forms (5,141,279 in the dump) are not shipped, so a list search finds "стол" and not "столами". Forms are written in capitals in the dump and are lowercased here; entries with a hyphen, an apostrophe or a digit are dropped; the letter yo is folded to ye for matching and the spelling with yo is kept in display.json. - The answers tier is the nouns among those headwords (nominative singular, and the nominative plural of the plural-only nouns), without the colloquial (Infr) and slang (Slng) entries, that are also among the 50000 most frequent FrequencyWords tokens (the frequency of the yo and ye spellings added together), minus the offensive words and minus 320 valid words that are not offered unasked (mild insults, ethnic terms and sensitive vocabulary: they stay valid and unblocked). This is a curation choice of this project, not a rule of any word game. - The offensive list is made from the single Cyrillic words of LDNOOBW and safetext. An entry that is a form of an OpenCorpora lexeme is decided for the whole lexeme by a person: 197 lemmas are blocked with every form of their paradigm (3015 forms), and 297 lemmas that the lists flag are ordinary words and are never blocked because of a list (a ball, a handle, a socket, a branch, the English "who" and "sick" written in Cyrillic, euphemisms made on the fig and the horseradish). The dump splits a verb over several lexemes (the finite forms, the infinitive, the participles, the gerunds) and an adjective over a full-form lexeme and the lexemes of its short forms and comparatives, and links none of them; they stand next to each other, so every lexeme of the family of a blocked infinitive or full adjective is blocked with all its forms too (158 lexemes, 1402 forms), after a check that it shares its first letters with the family. Entries that are no form of any lexeme a list names (3956, mostly spelled-out derivations of the vulgar roots) are blocked as written, as whole words only, plus 53 words found by the audit. The lists carry no severity and no topic that is used as a blocklist; identity terms are not blocked. - Seven audits run inside the build and stop it on an undecided word: every lexeme with a form on a source list has a decision; every blocked form that is also the headword of an unblocked lexeme is protected or accepted on purpose; every blocked key that only a name, a surname or a distortion matches is accepted on purpose; the in-house stem rules (a stem plus a Russian ending, behind at most one prefix, or a long root anywhere) block every main word or frequency token that holds a root of the vulgar families unless a person reviewed it as an ordinary word, and a dictionary headword that only the rules would block stops the build until it is decided one by one; and the roots of the blocked slurs, insults, sexual and excrement words (a long root plus anything, a short root plus a closed set of endings) make every main word or frequency token that holds one a decision: blocked, or reviewed as an ordinary word with the reason (51 reviewed words in all: look-alikes, clinical vocabulary, violence in the general sense, Sodom, carrion and vultures). Two audits read the dump a second way: in every verb or adjective family that has a lexeme blocked by its key or a blocked form, the infinitive or the full adjective is decided too (65 families are flagged, 61 have their anchor blocked or protected and 4 are accepted on purpose with the reason: the participle "dolbannyi", the short adjective "kalov", the imperatives "svoloch'i" and "sosi" that are blocked words or forms of one), and every short adjective or comparative that no full adjective stands before is a headword or a form. The stem-rule and family-root audits also run on the committed files. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 127813 words extended 0 words answers 5985 words frequency 47009 tokens offensive 6976 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt CC-BY-SA-3.0.txt CC-BY-SA-4.0.txt MIT-safetext.txt