German word list data, Free Tool Explorer ========================================= These files are the German word lists used by the Free Tool Explorer word tools (language code de). They are open data, built offline by "npm run words:build -- --lang=de" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters a ä b c d e f g h i j k l m n o ö p q r s t u ü v w x y z. The sharp s becomes ss; a, o and u with umlaut are kept as separate letters; every other accent is removed and the oe and ae ligatures become two letters. Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. enz/german-wordlist, the German word list for word games (words), commit 1071d87d9535b57b80ec076c699f2578a28ee9dc (branch main, 2026-09-13) Used for: validity list (main tier) Source: https://raw.githubusercontent.com/enz/german-wordlist/1071d87d9535b57b80ec076c699f2578a28ee9dc/words SHA-256: 20095928ca70974fb5d86e44f10181cee16e2831b0d4b1a4ddd12c261aebc5e9 Licence: CC0-1.0 Credit: German word list by Markus Enzenberger and contributors (https://github.com/enz/german-wordlist), released under CC0 1.0 (public domain dedication). No credit is required; we name the authors as a courtesy, and the CC0 text is reproduced verbatim as CC0-1.0.txt. The list's own README says it is far from complete and still holds mistakes. Changes: - Dropped the entries that were judged by hand, spelling by spelling, to be proper nouns the list's own rules exclude but that leaked in (countries, regions, rivers, cities, first names, surnames and a few trademarked names, named one by one in the build recipe together with their genitive in -s; a word that is also an ordinary word, the form of one or the plural of a person noun, such as Hessen, Polen, Essen, Roman, Erika, Rosa, Jochen, Mosel, Wade, Luke, Tatort or Android, was kept) and the entries with a capital letter inside the word (brand names and abbreviations such as BitTorrent, WiFi and GroKo). - Folded the remaining entries to the stored form: lower case, the sharp s spelled ss, the ligatures oe and ae spelled out, every accent other than the umlauts of a, o and u removed (e, c, n, a with acute or grave and similar), a, o and u with umlaut kept as letters of their own. Entries with a letter that has no such fold (for example l with stroke) were dropped; entries of 2 to 15 letters were kept. - Merged the spellings that fold together and wrote the capitalised noun spellings to display.json (stored form to the spellings of the source), so a capital Haus and a lower case haus are one stored word with both display forms. Note: The repository's own blacklist file (words it rejected) is not used. 2. FrequencyWords, German (content/2018/de/de_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/de/de_50k.txt SHA-256: d9e50546fd7e8b6fe6542a2b33c51d1331092b2a3916ec09f80d97856068705b Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form and dropped those with letters outside the German alphabet or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list. - Used the 20,000 most frequent lines to restrict the main list to common words for the answers tier. The source is written in lower case, so the join is made on the folded form; because the selection is made with this list, answers/*.txt is also shared under CC BY-SA 4.0. Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. 3. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), German, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/de SHA-256: ebe44d7772897390e27f7bec412fdcdf0efc294e4db157f8487c4cd8d4fabd55 Licence: CC-BY-4.0 Credit: LDNOOBW, the German list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries and folded them to the stored form, removed the ordinary, religious-title and mild words the project does not block (naked, a nipple that is also a pipe fitting, a rosette, a notch, a rascal, a booger, a grimace, a bigwig, a mufti and similar), and expanded every remaining entry with the inflected forms that are a word of the main list or a token of the frequency list, the way the build's own rules inflect a German noun, verb or adjective (endings, umlaut plurals, participles). Phrases were dropped. 4. Creative Commons CC0 1.0 Universal, legal code (the COPYING file of the enz/german-wordlist repository), commit 1071d87d9535b57b80ec076c699f2578a28ee9dc Used for: licence text Source: https://raw.githubusercontent.com/enz/german-wordlist/1071d87d9535b57b80ec076c699f2578a28ee9dc/COPYING SHA-256: a2010f343487d3f7618affe54f789f5487602331c0a8d03f49e9a7c547cf0499 Licence: CC0-1.0 Credit: Licence text of the enz/german-wordlist word list. 5. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data, as published by Creative Commons. 6. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is enz/german-wordlist (CC0), every inflected form already written out in the source. Before the shared steps 301 entries were dropped that were judged by hand, spelling by spelling, to be proper nouns (countries, regions, rivers, cities, first names and a few trademarked names that the list's own rules exclude but that leaked in; a name that is also an ordinary word, the form of one or the plural of a person noun, such as Essen, Roman, Rosa, Hessen, Polen, Jochen, Mosel, Wade, Luke, Tatort or Android, was kept) and 23 entries with a capital letter inside the word (brand names and abbreviations such as BitTorrent, WiFi and GroKo). The shared steps then fold every entry to the stored form (lower case, ss for the sharp s, oe and ae for the ligatures, every accent except the umlauts of a, o and u removed) and keep 2 to 15 letters. The capitalised noun spellings are kept for display.json. - The answers tier is the 20000 most frequent FrequencyWords tokens that are also main words (FrequencyWords is lower case, so the two are joined on the folded form), minus the offensive words and minus 59 tokens of a policy list (ethnic terms and insults many people find offensive, crude words too mild for the blocklist, and explicit sexual vocabulary and words of sexual violence that no source list names; they stay valid and unblocked). The frequency source is not lemmatised, so conjugated and declined forms (macht, sagte, Hauses) are answers too. - The offensive list is the LDNOOBW German list (minus 17 ordinary words it flags: naked, a nipple, a rosette, a mufti, the plural of Mops and similar) plus 289 in-house lemmas, derived words and compounds (slurs and their variant spellings, the Latin plurals Orgasmen and Penes, and the English and German profanity and slurs that the subtitle frequency list carries), expanded with the inflected forms of each entry. The source has no lemmas, so the expansion uses the regular German endings (noun, umlaut plural, adjective and comparative, verb, participle and separable-prefix forms) and a form counts only when it is a word of the main list or a token of the frequency list. A few ordinary words that these rules would block (vögeln, the dative plural of Vogel, and 13 others) are protected. Identity terms (schwul, lesbisch) are not blocked, and the filter is a list of words, never a topic. - A derivation audit runs in the data test: every main word that is an in-house lemma plus a derivational ending, or a compound that starts or ends with a blocked word, is blocked, protected, or reviewed as ordinary by hand. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 590591 words extended 0 words answers 16139 words frequency 48182 tokens offensive 1347 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt CC-BY-SA-4.0.txt CC0-1.0.txt