English word list data, Free Tool Explorer ========================================== These files are the English word lists used by the Free Tool Explorer word tools (language code en). They are open data, built offline by "npm run words:build -- --lang=en" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters a b c d e f g h i j k l m n o p q r s t u v w x y z. Accents and other combining marks are removed (cafe for the accented spelling). Nothing else is folded: letters outside a-z make an entry invalid. Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. SCOWL (Spell Checker Oriented Word Lists), version 1, 2020.12.07 Used for: validity list (main tier) Source: https://downloads.sourceforge.net/project/wordlist/SCOWL/2020.12.07/scowl-2020.12.07.tar.gz SHA-256: 5587667caa20c4891390c2d42dbb4d5c4c3f41bee77af1457ece3ba23fb859cc Licence: LicenseRef-SCOWL Credit: SCOWL, the collective work Copyright 2000-2018 by Kevin Atkinson (http://wordlist.aspell.net/), together with the component copyrights named in the tarball's own Copyright file (UKACD, WordNet 1.6, Ispell, VarCon and others). That Copyright file is reproduced verbatim as SCOWL-Copyright.txt. Changes: - Read the word lists of sizes 10 to 80 for the english, american, british, british_z, canadian, australian and variant_1 spellings (ISO-8859-1 converted to UTF-8), removed accents, and kept only entries of 2 to 15 letters a-z. - Added the word "app" by hand: SCOWL files it under abbreviations. - Used the sizes up to 50 as the candidate set of the answers tier. Note: Size 80 is described by its authors as the level with the strange and unusual words people like to use in word games. 2. ENABLE (Enhanced North American Benchmark LExicon), enable1 word list, enable1.txt Used for: validity list (main tier) Source: https://norvig.com/ngrams/enable1.txt SHA-256: 61ba1392a5b6199dd161fae7b483cd262fb8a88c0d49bd2c9ecb6753e755697f Licence: LicenseRef-PublicDomain-ENABLE Credit: ENABLE word list, released into the public domain by its authors (research and compilation by Alan Beale, maintained by M. Cooper, with contributors). We credit them as the originators of the list, as they ask, and we do not restrict redistribution of the list in any way. Changes: - Kept only entries of 2 to 15 letters a-z and united them with the SCOWL list. Note: The authors' public-domain notice is part of the SCOWL Copyright file shipped next to this file. 3. FrequencyWords, English (content/2018/en/en_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/en/en_50k.txt SHA-256: 5351ff405b1126ef555791dd4d9798a48e3e9a501a9fc481a9da957752cfb458 Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form and dropped those with letters outside a-z or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list. - Used the tokens to restrict the SCOWL size-50 candidates to common words for the answers tier; because that selection is made with this list, answers/*.txt is also shared under CC BY-SA 4.0. Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. 4. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), English, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/en SHA-256: af851ecef1d5f212caba17339b12ac39cc2fef7d78c74876f67237644fcee8bd Licence: CC-BY-4.0 Credit: LDNOOBW, the English list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries, folded to the stored form, and removed the few ordinary, religious-name and identity words the project protects. Phrases were dropped. 5. profanity-list, English (en.json), commit c27924319aa9bd6f917e3782b4f4b6604a50b652 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/dsojevic/profanity-list/c27924319aa9bd6f917e3782b4f4b6604a50b652/en.json SHA-256: bcf8bfc09dd3f7480972cb041a4cba291a6ff16a6fbf29022d5d679e3c39502c Licence: MIT Credit: profanity-list by David Sojevic (https://github.com/dsojevic/profanity-list), MIT licence, Copyright (c) 2021 David Sojevic. The licence text is shipped as MIT-profanity-list.txt. Changes: - Split every match into its alternatives, kept the single words, expanded the wildcard patterns (a starred letter repeats one or more times) against the word lists as whole words, and ignored the severity levels and topic tags. 6. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data, as published by Creative Commons. 7. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. 8. MIT licence of profanity-list (the LICENSE file of the repository), commit c27924319aa9bd6f917e3782b4f4b6604a50b652 Used for: licence text Source: https://raw.githubusercontent.com/dsojevic/profanity-list/c27924319aa9bd6f917e3782b4f4b6604a50b652/LICENSE SHA-256: 33962cb4fc7b987408d5ea301ad46a2f54f9ba11c56d7a039459f931736363aa Licence: MIT Credit: Licence text and copyright notice of profanity-list. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is the union of SCOWL sizes 10 to 80 (english, american, british, british_z, canadian, australian and variant_1 spellings) and ENABLE. Entries are read as ISO-8859-1, reduced to plain letters (accents removed) and kept only when they are 2 to 15 lowercase letters a-z, so proper nouns, capitalised abbreviations, contractions and hyphenated entries are dropped. The word "app", which SCOWL files under abbreviations, was added back by hand. - SCOWL also files abbreviations such as ln, ls and ts among its ordinary words, so its entries are checked against ENABLE, which lists no abbreviations: 298 SCOWL entries of 2 or 3 letters that ENABLE does not list were dropped, and so were 4 entries of 4 or more letters with no vowel (a, e, i, o, u or y: csch, pssts, psts, pwns) and 12 initialisms with vowels (dsos, flir, kcal, milf, mips, mirv, mirvs, psia, psid, terf, ufos, vuln). The 8 real two-letter words that post-date ENABLE (da, fe, ki, oi, po, qi, te, za) were reviewed and kept. - The answers tier is the SCOWL words of sizes up to 50 that also appear among the 50000 FrequencyWords tokens, minus the offensive words and minus 13 valid words that are not offered unasked (a few slang, crude and ethnic terms that many people find offensive). - The offensive list is LDNOOBW plus profanity-list plus 252 in-house words, expanded with their inflected forms that are real words. A protected list of 39 words is never blocked: ordinary, religious-name and identity words that those topic-based lists flag, and ordinary inflected forms (cocked, damning, sexes) that the inflection rules would otherwise generate. Anatomical and sexual vocabulary that a list names as a word stays offensive. Severity levels and topic tags of profanity-list are not used. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 247850 words extended 0 words answers 29856 words frequency 46372 tokens offensive 1137 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt CC-BY-SA-4.0.txt MIT-profanity-list.txt SCOWL-Copyright.txt