French word list data, Free Tool Explorer ========================================= These files are the French word lists used by the Free Tool Explorer word tools (language code fr). They are open data, built offline by "npm run words:build -- --lang=fr" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters a b c d e f g h i j k l m n o p q r s t u v w x y z. Accents and cedillas are removed; the oe and ae ligatures become two letters (coeur for the ligature spelling). Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. Grammalecte, Lexique des formes fléchies du français (Dicollecte), 7.7 (December 2025) Used for: validity list (main tier) Source: https://grammalecte.net/dic/lexique-grammalecte-fr-v7.7.zip SHA-256: bdccf2065e252010253171000a38de84581407896b6bc60aca7d4d69dab7e4d2 Licence: MPL-2.0 Credit: French word list © Olivier R. / Grammalecte (Dicollecte), MPL 2.0, https://grammalecte.net/. Lexique des formes fléchies du français, version 7.7. The files of this folder and data/words/fr are derived from it (Covered Software under the Mozilla Public License, v. 2.0) and are modified: the entries were filtered, folded to plain letters and deduplicated, split by length, and the frequency columns were reduced to a ranked list. The unmodified source files are available from the URL above. This Source Code Form is subject to the terms of the Mozilla Public License, v. 2.0; if a copy of the MPL was not distributed with this file, you can obtain one at http://mozilla.org/MPL/2.0/. The licence text is shipped as MPL-2.0.txt and the lexique's own README as Grammalecte-README_lexique.txt. Changes: - Kept the rows of the sub-dictionaries *, C, M and R (the classical and the 1990-reform spellings) as the main list, and the rows of sub-dictionary X as the extended tier (words that are not in the main list). Dropped the rows of sub-dictionary A (spellings only the all-variants dictionary accepts). - Dropped, row by row, the rows tagged npr, prn, patr, sign, ponc, pfx, sfx, err or div, the rows whose notes include symb (unit symbols such as km and kg) or abty (typographic abbreviations such as etc and cf), and the rows for the ligature letters oe and ae written as one sign. The clipped words (abr) and the unit names (unit) were kept. - Kept only forms made of lowercase letters, folded accents and cedillas away and spelled the oe and ae ligatures out, dropped forms with a hyphen, apostrophe, digit or other character, and kept words of 2 to 15 letters. - Used the frequency index (0 to 9) for the answers tier (index 6 and up, never a word that has a sub-dictionary X row, minus the offensive words and a short list of ordinary words that are also crude or explicitly sexual) and for freq.txt (the forms with index 5 and up, ranked by their total occurrence counts, written without the counts). - Added every inflected form the lexique gives to each in-house blocked lemma (the lemmas of the LDNOOBW entries, the words reported removed from the official French tournament word list as racist, homophobic or pejorative, and the real slurs among the lexique's pejorative entries) to offensive.txt, after checking each derived lemma by hand. Note: The lexique was built from the Dicollecte database; its frequency columns come from Google Books 1-grams, Wikipédia, Wikisource and a literature corpus. It contains both the 1990-reform and the older spellings. 2. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), French, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/fr SHA-256: 79763a705e08256b393952852bb63aa903f4485e0ef02f70735f5e50efff213c Licence: CC-BY-4.0 Credit: LDNOOBW, the French list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries, folded to the stored form, and removed the ordinary words the project does not treat as offensive (everyday vocabulary around sex and the body, such as bander, jouir and gerbe). Phrases were dropped. Entries that are lemmas of the lexique were expanded to every inflected form the lexique gives them. 3. Mozilla Public License, version 2.0 (plain-text legal code published by Mozilla), 2.0 Used for: licence text Source: https://www.mozilla.org/media/MPL/2.0/index.txt SHA-256: 3f3d9e0024b1921b067d6f7f88deb4a60cbe7a78e76c64e3f1d7fc3b779b9d04 Licence: MPL-2.0 Credit: Licence text of the Grammalecte lexique, as published by the Mozilla Foundation. 4. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the LDNOOBW list. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is the Grammalecte lexique of inflected forms (version 7.7) restricted to the sub-dictionaries * C M R (the classical and the 1990-reform spellings are both valid). Rows tagged as proper nouns (places, first names, surnames), signs, punctuation, affixes or elision fragments are dropped, and so are rows noted as unit symbols (km, kg) or typographic abbreviations (etc, cf), each row on its own: a word with another, ordinary row stays. Clipped words (ado, apéro) and unit names (ampère) are kept; the rows for the letters oe and ae written as one sign are not read (folded they would be the non-words oe and ae). Only forms made of lowercase letters are read, so capitals, hyphens, apostrophes, digits and periods drop out; accents and cedillas are removed and the oe and ae ligatures are spelled out. - The extended tier holds the forms of sub-dictionary X (technical, rare and variant words) that are not in the main list; it only answers "is this a word". The forms of sub-dictionary A, which only the all-variants dictionary accepts, are in neither tier. - The answers tier is the main-list words whose lexique frequency index is 6 or more (index 9 is the most frequent), never a word that has a sub-dictionary X row, minus the offensive words and minus 15 valid words that are not offered unasked (20 ordinary words that are also crude or are an ethnonym also used as a slur (chleuh), and every form of 24 explicit sexual and anatomical lemmas, which no source list names and which therefore stay valid and unblocked). freq.txt lists the lexique's own forms with index 5 or more, ranked by their counts in the lexique's corpora (Google Books 1-grams, Wikipédia, Wikisource, literature). - The offensive list is the LDNOOBW French list (minus 17 ordinary words it flags) plus 175 in-house lemmas, each expanded to every inflected form the lexique gives it (1289 stored forms). The in-house lemmas were compiled by hand from the LDNOOBW entries and their families, the lexique's pejorative notes (only the real slurs), and the words the French Wikipedia reports removed from the official French tournament word list as racist, homophobic or pejorative (words taken as facts). 18 ordinary words that share a form with a blocked verb or noun (a swing, a rod, a seed, the plural of fou) are never blocked. The lexique's sexe, arg and péj notes are topics, not lists, and are not used as blocklists; identity terms (homosexuel, gay, lesbienne, trans, and chleuh, the name of the Shilha people) are not blocked. - Five audits run inside the build and stop it on an undecided word: every main word of a blocked lemma is blocked or protected; every blocked form that an unblocked lemma shares is protected or accepted on purpose; every lemma that the lexique files under the same spelling key (its "Métagraphe" column) as a blocked lemma, which is how it records a variant spelling such as polaque next to polack, is blocked or reviewed as ordinary; every lemma that derives from a blocked word (by suffix, by a shared stem of four letters or more, by the infinitive stem of a verb plus an ending such as niqu + eur, as a compound, or, for the lemmas the lexique marks as colloquial, pejorative or slang, by holding a root of four letters or more anywhere inside) is blocked or reviewed as ordinary; every LDNOOBW entry has all its lemmas blocked or reviewed. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 406193 words extended 1263 words answers 28876 words frequency 74034 tokens offensive 1280 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt Grammalecte-README_lexique.txt MPL-2.0.txt