Portuguese word list data, Free Tool Explorer ============================================= These files are the Portuguese word lists used by the Free Tool Explorer word tools (language code pt). They are open data, built offline by "npm run words:build -- --lang=pt" from the sources listed below, and served unchanged to the browser. They are not an official game dictionary and may differ from any tournament or board-game word list. STORED FORM ----------- Every word is lowercase, Unicode NFC, after folding, and made only of the letters a b c d e f g h i j k l m n o p q r s t u v w x y z. Every accent, tilde and the cedilla are removed. Words have 2 to 15 letters and each file is sorted in UTF-16 code-unit order. SOURCES ------- 1. IME-USP Brazilian Portuguese word list (br-utf8.txt), curated by Paulo Feofiloff from br.ispell, br-utf8.txt, 261,788 lines, file last modified 2024-09-03 (page updated 2025-05-17) Used for: validity list (main tier) Source: https://www.ime.usp.br/~pf/dicios/br-utf8.txt SHA-256: 36ae3cead899c9746d116d2332422bc5ba23476132449d9f06911ef396bf3558 Licence: CC-BY-4.0 AND GPL-2.0-or-later Credit: Brazilian Portuguese word list: IME-USP (Instituto de Matematica e Estatistica, Universidade de Sao Paulo), curated by Paulo Feofiloff (https://www.ime.usp.br/~pf/dicios/), licensed Creative Commons Attribution 4.0 (CC BY 4.0). It is derived from br.ispell, Copyright Ricardo Ueda Karpischek, distributed under the GNU General Public License. The two licences are not reconciled on the IME page, so this project complies with both: the CC BY 4.0 text is shipped as CC-BY-4.0.txt and the GPL text as GPL-2.0.txt. main/*.txt and the tiers derived from it are shared under both. Changes: - Dropped every entry that is not written entirely in lowercase letters: 770 capitalised entries (proper nouns of people, places and books of the Bible, and acronyms) and the two hyphenated entries (micro-onda, micro-ondas). - Dropped, after a review of every candidate, the lowercase spellings that are nothing but a proper noun (first names and surnames, countries, cities, regions, rivers, saints and gods: viena, egito, bahia, joao, maria). A spelling that is also an ordinary word (rosa, mar, sul, a month, a feminine nationality) was kept. The candidates came from two detectors, an entry whose capitalised spelling is also an entry and an entry that the Natura dictionary marks as a proper noun, and from the names that the frequency list shows. - Dropped the fragments that the list's own clitic splitting left behind: a first person plural verb form without its final s (abafamo for abafamos, lancemo, admitemo), which the list keeps only because the clitic nos follows them in the original text. An entry is dropped only when it ends the way such a form ends (-amo, -emo, -imo, -rmo, -pomo; superlatives and ordinals are spared), its s-form is also an entry (not required for -emo), and neither Hunspell (Natura) nor the frequency list knows it as a word. - Dropped the second kind of fragment of the clitic splitting: an infinitive with its r cut off before the clitic -lo or -la (abafá from abafá-lo, bloqueá, compô, havê, caçá, 3,646 entries; folded they are non-words such as bloquea, passea, compo, supo and have, and false display spellings such as abafá for abafa). An entry is dropped only when it ends in á, ê or ô, the infinitive it was cut from (-ar, -er or -or) is also an entry, and neither Hunspell (Natura) nor the frequency list knows the exact accented spelling as a word, which spares vê, crê, lê, dá, está, pá, bebê and babá. - Dropped a few misspellings (asssegurado with three s, artesaao, raas), the whole paradigm of fugir written with a j (fujiu, fujimos, fujir and 39 more forms, plus the two cut stems fujimo and fujirmo; the forms that really have a j, fuja, fujo and fujam, stay), the abbreviation etc and the Roman numeral vii. - Dropped, after a word-by-word review, 113 wrong spellings of words whose right spelling is also an entry (lampâda, mediócre, aparicão, heróina, armísticio and the other accents on the wrong syllable; the 13 participles written without their accent, atraido, caido, saido, traido; the esforçar paradigm without its cedilla, esforco, esforcou, esforcar and 45 more, while esforce and esforcei, which are right, stay; viuva, viuvo, familia, espetaculo; the spellings from before 1971, corôa, fôrça, pêlo; a few subtitle artefacts). They were found by comparing each spelling with the Natura dictionary (hunspell -l) and the frequency list, and nothing is lost from the word list: each folds to a stored form that its right twin keeps, so only the display spellings change. 49 entries that are wrong and have no right twin are replaced by their right spelling, which folds to the same stored form: viuvam, viuve, viuves, viuvem and ruiram (viúvam, ruíram), the answers words with an accent of the pre-1971 spelling or on the wrong syllable (ítem, ítens, crú, crús, nús, matrimônial, matrimôniais, polonesês), and the 36 verb forms in -ínge and -íge (atínge, tínge, redíge, erige) that carry a stem accent they never take. - Folded the remaining entries to the stored form (every accent, tilde and the cedilla removed), deduplicated them, and kept only words of 2 to 15 letters a-z. Note: The IME page calls the licence CC BY and links the 4.0 deed. The upstream br.ispell README says only GNU GPL, and the GitHub mirror of br.ispell carries the GPL version 2 text, which is the text shipped here; no version is named, so the licence is recorded as GPL-2.0-or-later (GPL version 2, section 9: a recipient may choose any published version when none is named) and the version 2 text is shipped. The list follows the 2003-era spelling and is only partly updated to the 1990 spelling agreement (ideia, voo, joia are updated; heroico, paranoico, asteroide, polo, pelo are not), so the accept-only Natura list (extended tier) covers the compounds and the European spellings it lacks. 2. Natura Portuguese (Portugal) dictionary for Hunspell, pt_PT.dic and pt_PT.aff (Projecto Natura, Universidade do Minho), hunspell-pt_PT-20251001 (1 October 2025), 1990 spelling agreement Used for: accept-only list (extended tier) Source: https://natura.di.uminho.pt/download/sources/Dictionaries/hunspell/hunspell-pt_PT-20251001.tar.gz SHA-256: 2066157087e83264484a6d564e1d85258fe2c5eddb9713a42d6f6cf06b0a2ed9 Licence: GPL-2.0-only OR LGPL-2.1-only OR MPL-1.1 Credit: European Portuguese word list derived from the Hunspell dictionary of Projecto Natura, Universidade do Minho (Copyright 2006-2009 Jose Joao de Almeida, Rui Vilela, Alberto Simoes; Departamento de Informatica, Universidade do Minho, Portugal; http://natura.di.uminho.pt/), covered by the GPL, the LGPL and the MPL (GPL version 2, LGPL version 2.1, MPL version 1.1), expanded and modified as described below. extended/*.txt is derived from it and is shared under the GNU General Public License, version 2; the README of the dictionary and its COPYING file, which carries the three licence texts, are reproduced verbatim as Natura-README.txt and Natura-COPYING.txt. Changes: - Dropped, before the expansion, the entries that the dictionary itself marks as proper nouns (CAT=np: places, people, brands), as an acronym (SEM=sigla), as an abbreviation (ABR=1: cm, kg, dr, sr) or as a Roman numeral (subcat=rom), every entry that is not made only of lowercase letters, and the entries with a hyphen or a space; the tab-separated morphology fields ([CAT=...], [PREAO90=...] and the rest) were stripped from every entry, so the pre-1990 spellings they carry are not used. - Expanded the remaining stems into every inflected, derived and prefixed form with unmunch from the hunspell tools (the dictionary has one-byte affix flags and no continuation flags, so one pass is complete); the forms of each stem were kept together. unmunch matches the condition of an affix rule byte by byte, so the plural rule never fires for a form that ends in an accented vowel (café, pé, chá, rã, maçã): the plural of every such generated form was asked of hunspell -l with the original dictionary, and the accepted ones (cafés, pés, chás, rãs, maçãs, vocês, robôs and others) were added to the entry that generates the singular. Every generated form was accepted by hunspell -l with the original dictionary, and every frequency token that hunspell accepts is either a generated form or a form of a dropped entry; the dropped entries were expanded too to check this, and the build stops on a token that nothing explains. - Folded the forms to the stored form (every accent, tilde and the cedilla removed), kept only words of 2 to 15 letters a-z, and removed every word that is in the main list: what remains is the extended tier (accepted by the word checker only, never in answers). The accented spellings of these words are not kept (display.json lists main words only): the accept-only tier is never shown as a word list. Note: The README says the files are covered by 'the (GPL/LGPL/MPL), by this order' and names GPL version 2, LGPL version 2.1 and MPL version 1.1; whether 'by this order' is a choice or a conjunction is not stated (the wooorm index reads it as a choice). LibreOffice's LICENSES.txt for the identical pt_PT dictionary says 'GPL and BSD' instead, so the licence statements disagree. The extended tier is therefore shipped under the most restrictive reading, the GPL, with all three texts. The dictionary is identical to LibreOffice's pt_PT. 3. FrequencyWords, Brazilian Portuguese (content/2018/pt_br/pt_br_50k.txt), commit 525f9b560de45753a5ea01069454e72e9aa541c6, from OpenSubtitles 2018 Used for: word frequencies (freq.txt) Source: https://raw.githubusercontent.com/hermitdave/FrequencyWords/525f9b560de45753a5ea01069454e72e9aa541c6/content/2018/pt_br/pt_br_50k.txt SHA-256: a61d6f2ede97c5daad5fb3907b72a228f0aff12668be6d24da590d20804fa611 Licence: CC-BY-SA-4.0 Credit: Word frequencies: FrequencyWords by Hermit Dave (https://github.com/hermitdave/FrequencyWords), built from OpenSubtitles 2018 (https://www.opensubtitles.org/, http://opus.nlpl.eu/OpenSubtitles2018.php), licensed CC BY-SA 4.0. freq.txt is a derivative of it and is shared under the same licence. Changes: - Folded the tokens to the stored form and dropped those with letters outside a-z or fewer than 2 or more than 15 letters, dropped offensive tokens, merged tokens that folded together, and wrote them most frequent first without the counts (the line number is the rank). The list is not joined to the word list. - Used the 20,000 most frequent lines to restrict the main list to common words for the answers tier (IME words only, never the Natura words); because that selection is made with this list, answers/*.txt is also shared under CC BY-SA 4.0. Note: The repository states: MIT licence for code, CC BY-SA 4.0 for content. The European Portuguese list (content/2018/pt) is not used: the answers tier follows the Brazilian list. 4. List of Dirty, Naughty, Obscene, and Otherwise Bad Words (LDNOOBW), Portuguese, commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/pt SHA-256: 05ef524b0bb7f83f87257f6ec4dabf70e16b8bb6930f769fb93deaf608ff196d Licence: CC-BY-4.0 Credit: LDNOOBW, the Portuguese list of Shutterstock and contributors (https://github.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words), licensed CC BY 4.0. Changes: - Kept the single-word entries and folded them to the stored form, removed the ordinary, topic and identity words the project does not block (abortion, beer, hell, a spider, a donkey, a tap, to eat, heroin, a homosexual, a lesbian, a condom and similar), and expanded every remaining entry to the forms the Natura dictionary's own affix classes give the stem it names and to the gender and number forms that are real words, without touching a form that an unrelated, unblocked stem shares unless a collision was accepted on purpose. Phrases were dropped. 5. cuss, Brazilian Portuguese (pt.js), commit 6bab3fef250481e34ba55bc400fac5c6d25f1429 (cuss 2.2.0) Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/words/cuss/6bab3fef250481e34ba55bc400fac5c6d25f1429/pt.js SHA-256: d7eb3d6c1ba4efa9c01ad2c354c7d5f46946342a3bb295c17204fffc6b4b54ca Licence: MIT Credit: cuss by Titus Wormer (https://github.com/words/cuss), MIT licence, Copyright (c) 2016 Titus Wormer. The licence text is shipped as MIT-cuss.txt. The Portuguese list was compiled by its author from a public site of Brazilian words (aprenderpalavras.com). Changes: - Read the file as data (it is never run), kept only the entries with the top sureness rating of 2, reviewed every one of them by hand, blocked the words that are vulgar, obscene or slurs in everyday use, and declined the ordinary words and the mild insults it also lists (a spider, a dog, a hen, a fool, a sucker, a tail). Every blocked word was expanded to all the forms of the dictionary stem it names and to its own gender and number forms that are real words. Entries with a lower rating and phrases were not used. 6. cuss, European Portuguese (pt-pt.js), commit 6bab3fef250481e34ba55bc400fac5c6d25f1429 (cuss 2.2.0) Used for: offensive-word blocklist (offensive.txt) Source: https://raw.githubusercontent.com/words/cuss/6bab3fef250481e34ba55bc400fac5c6d25f1429/pt-pt.js SHA-256: 30d4571120e7f6d1f1743e7b3624120e7087e5bdb9971dc5b38108057bc04994 Licence: MIT Credit: cuss by Titus Wormer (https://github.com/words/cuss), MIT licence, Copyright (c) 2016 Titus Wormer. The licence text is shipped as MIT-cuss.txt. The European Portuguese list was compiled by its author from Portuguese Wikipedia. Changes: - Read the file as data (it is never run), kept only the entries with the top sureness rating of 2, reviewed every one of them by hand, blocked the words that are vulgar, obscene or slurs in everyday use, and declined the ordinary words (a brooch, an orgasm) it also lists; every blocked word was expanded like the Brazilian entries. Entries with a lower rating and phrases were not used. 7. Creative Commons Attribution-ShareAlike 4.0 International, legal code, 4.0 Used for: licence text Source: https://creativecommons.org/licenses/by-sa/4.0/legalcode.txt SHA-256: 28a9529c7d0bb4dc51f4bf5c116a3d16ef247a052f7591466768ddf563fd1cf5 Licence: CC-BY-SA-4.0 Credit: Licence text of the FrequencyWords data, as published by Creative Commons. 8. Creative Commons Attribution 4.0 International, legal code (the LICENSE file of the LDNOOBW repository), commit 5faf2ba42d7b1c0977169ec3611df25a3c08eb13 Used for: licence text Source: https://raw.githubusercontent.com/LDNOOBW/List-of-Dirty-Naughty-Obscene-and-Otherwise-Bad-Words/5faf2ba42d7b1c0977169ec3611df25a3c08eb13/LICENSE SHA-256: 55bfbc1759943f7d19acfc68e35644b7fbe10cd8d39ab3551ce5ecff5dda9979 Licence: CC-BY-4.0 Credit: Licence text of the IME-USP word list and of the LDNOOBW list. 9. GNU General Public License, version 2 (the LICENSE file of the GitHub mirror of br.ispell), commit 1fe7b79f4746e36dfbfb0feb578960efc94088e1 of fititnt/br.ispell-dicionario-portugues-brasileiro Used for: licence text Source: https://raw.githubusercontent.com/fititnt/br.ispell-dicionario-portugues-brasileiro/1fe7b79f4746e36dfbfb0feb578960efc94088e1/LICENSE SHA-256: f9c375a1be4a41f7b70301dd83c91cb89e41567478859b77eef375a52d782505 Licence: GPL-2.0-or-later Credit: Licence text of br.ispell, the upstream of the IME-USP word list, as carried by its GitHub mirror. 10. MIT licence of cuss (the license file of the repository), commit 6bab3fef250481e34ba55bc400fac5c6d25f1429 Used for: licence text Source: https://raw.githubusercontent.com/words/cuss/6bab3fef250481e34ba55bc400fac5c6d25f1429/license SHA-256: ca4662cb5d1b738fbe5350c0d5485ba11773b4b7208974082ae6e129a52d631d Licence: MIT Credit: Licence text and copyright notice of cuss. WHAT WE CHANGED --------------- Everything below was done by the build, in this order: the source lists were filtered (proper nouns, abbreviations, entries with spaces, hyphens, apostrophes or digits, and entries with letters outside the alphabet were dropped), folded to the stored form, deduplicated, sorted and split into one file per word length. The answers tier is a subset of the main list with the offensive words removed. The frequency file holds the frequency source's own tokens, filtered by alphabet and the offensive list, and is not joined to the word list. The offensive list is an explicit whole-word list (never a substring or a topic tag) with its inflected forms. - The main list is the IME-USP word list br-utf8.txt (261788 lines, CC BY 4.0, derived from the GPL list br.ispell, so both licences are honoured), a flat list in which every form is written out, in the 2003-era spelling. Dropped before the shared steps: 772 entries that are not made only of lowercase letters (capitalised proper nouns and acronyms, two hyphenated entries), 318 lowercase spellings that were judged by hand, one at a time, to be nothing but a proper noun (first names, surnames, countries, cities, saints and gods: viena, egito, bahia, joao; a spelling that is also an ordinary word, such as rosa, mar, a month or the feminine of a nationality, was kept), 164 misspellings (asssegurado with three s, artesaao, raas, the whole paradigm of fugir written fujir: fujiu, fujimos, fujir, and 113 wrong spellings of words whose right spelling is also an entry, found by comparing the list with the Natura dictionary and the frequency list and then reviewed word by word: lampâda for lâmpada, aparicão, saido, esforco and the esforçar paradigm without its cedilla, viuva, familia, pêlo; each is dropped from display.json and loses nothing from the word list), 49 entries with a wrong or a missing accent that have no right twin and are replaced by the right spelling (viuvam for viúvam, ruiram for ruíram, ítem for item, crú for cru, atínge for atinge), 3 abbreviation and Roman numeral entries (etc, vii), and two kinds of fragment that the list's own clitic splitting left behind: 15983 first person plural verb forms without their final s (abafamo for abafamos), dropped only when the s-form is also an entry and neither the Natura dictionary nor the frequency list knows the entry as a word, and 3646 infinitives with their r cut off before -lo and -la (abafá for abafar, from abafá-lo; bloqueá, compô, havê, caçá), dropped only when the infinitive is also an entry and neither the Natura dictionary nor the frequency list knows the exact spelling as a word (vê, crê, lê, dá, está, pá and bebê are words and stay). The shared steps then fold every accent, tilde and the cedilla away, keep 2 to 15 letters and keep the accented spellings for display.json. - The extended tier is the European Portuguese dictionary of Projecto Natura (Universidade do Minho, 44476 entries, 1990 spelling agreement), expanded with unmunch into every form (372050 stored forms) after 2969 proper nouns (CAT=np), 1 acronym (SEM=sigla), 48 abbreviations and Roman numerals (ABR=1, subcat=rom: cm, kg, dr, xx) and 775 entries that are not made only of lowercase letters were dropped and the morphology fields were stripped (the pre-1990 spellings they carry are not used). unmunch matches the condition of an affix rule byte by byte, so the plural rule never fires for a form that ends in an accented vowel (café, pé, chá, rã, maçã); 86 plurals of that kind (cafés, pés, chás, rãs, maçãs, vocês, robôs) were therefore asked of hunspell -l with the original dictionary and added to the entry that generates the singular. Every generated form was accepted by hunspell -l with the original dictionary, and each of the 33631 frequency tokens that hunspell accepts is either generated by a kept entry (33584 tokens) or a form of one of the 3793 dropped entries (the other 47 tokens: names of months and seasons, abbreviations, Roman numerals), which the build checks by expanding the dropped entries too; it stops on a token that nothing explains. What the main list has is not repeated: the extended tier holds only the words main lacks (the European spellings and the 1990-agreement compounds), accepted by the word checker and never offered as an answer. Its accented spellings are not kept: display.json lists main words only, because the accept-only tier is never shown as a word list. - The answers tier is the 20000 most frequent FrequencyWords pt_br lines that are also main words (15277 of them; the Natura words never count), minus the offensive words and minus 122 words that are valid but not offered unasked: the policy list names 99 words (childish bodily words, mild insults, adult vocabulary and country names that are also the feminine of a nationality), which with the other forms of their stems are 495 forms. These stay valid and unblocked. The frequency source is not lemmatised, so conjugated forms are answers too. - The offensive list is the single words of LDNOOBW pt (minus 40 ordinary, topic and identity words it flags) and the 37 entries of words/cuss pt and pt-pt with rating 2 that survived a review of all 92 of them, plus 27 in-house words (derived words of the blocked roots that the audit found, and slurs the lists lack). Each is expanded to every form of the Natura stem it names, which is the dictionary's own morphology, and, for a seed that is no Natura stem, to the regular gender, number and verb forms of inflectPortuguese that are real words no Natura stem generates: 20 stems, 56 loose seeds, 251 blocked forms. 5 ordinary words that a blocked stem also generates are never blocked; 13 forms are blocked on purpose although an unrelated stem shares them. A seed names a Natura stem only when the stem is spelled the way the source spells the seed: cagado, the vulgar participle that cuss lists, is a form of cagar and does not block the turtle cágado (pairs that only fold together and are treated as forms: 1; pairs spelled differently and named on purpose: 1). Entries with a lower rating, phrases and HurtLex are not used, and no semantic tag is used as a blocklist. - Five audits run inside the build and stop it on an undecided word: every word of a blocked stem is blocked or protected; every blocked form that an unblocked stem shares is protected or accepted on purpose; every Natura stem that derives from a blocked root (the root plus an ending, or, for a root in a vowel, the root without it plus a diminutive or augmentative ending: bicha, bichinha, bichona; marica, mariquinha), or that starts or ends with a blocked form of 5 or more letters (44 roots; 36 stems derive from them without being blocked, and 33 reviewed stems record why they are ordinary words), is blocked, made of blocked or protected forms, or reviewed; every word that no Natura stem generates (the IME-only vocabulary and the frequency tokens) and that matches a root or the compound rule (16 words; 16 of them reviewed as ordinary) is blocked, protected or reviewed, so the published lists hold no vulgar word the paradigms miss; every word of the source lists is blocked or declined with a reason. FILES ----- main/.txt validity list, words of n letters extended/.txt accept-only words that are not in main (only where a language has them) answers/.txt common words (a subset of main, offensive words removed) display.json stored form -> display spellings, only where they differ freq.txt frequency source tokens, most frequent first (line number = rank) offensive.txt whole-word blocklist, one word per line manifest.json counts per tier and length, byte sizes, source versions COUNTS ------ main 224367 words extended 176064 words answers 15130 words frequency 48670 tokens offensive 251 words LICENCE TEXTS IN THIS FOLDER ---------------------------- CC-BY-4.0.txt CC-BY-SA-4.0.txt GPL-2.0.txt MIT-cuss.txt Natura-COPYING.txt Natura-README.txt