warda-lists
The lists of domains that every Warda box downloads: many public sources gathered in one place, sorted into categories, cleaned and deduplicated every day by the CI, then published at one address.
Before, each Warda box downloaded the third-party lists itself. Now the sources are chosen here, once, and the boxes download the result.
The documentation of the Warda project is in the repository
warda-dns/warda-docs, file lists/README.md.
How it works
sources.tomlnames the lists to download (the only file to edit to add one).- The CI downloads them every day, reads them the way Warda reads them,
puts each domain in its categories, adds the domains of
extra/and ofanalyst/add/, removes those ofallow.txt, ofprotect.txtand, in the protection lists only, ofanalyst/remove.txt, and removes the subdomains already covered by a parent. - The result,
dist/, is published on the branchdist. The boxes download it from there.
Add a list
Edit sources.toml, add a block, push on develop. That is all: the CI
checks the file, and once on main the list is in the next build. The
comments at the top of sources.toml explain every field.
Before adding a list, read its licence in its repository and download
its address once (curl -A "warda-lists/1.0"). The lists that were asked
for and not added (non-commercial or unknown licence, frozen list, wrong
fit for the categories, budget, download not tested) are kept in
PENDING.md, each with what would unlock it.
A list of one domain per line:
[[source]]
name = "example-malware" # unique, lower case, digits and "-"
group = "Malware List" # where it is filed (list below)
url = "https://example.org/malware.txt"
format = "domains"
licence = "MIT" # required: exact SPDX id
homepage = "https://example.org/"
categories = ["security"]
A hosts file (0.0.0.0 example.com):
[[source]]
name = "example-hosts"
group = "Hosts list"
url = "https://example.org/hosts"
format = "hosts"
licence = "CC-BY-4.0"
homepage = "https://example.org/"
categories = ["ads", "tracking"]
An Adblock list (only ||example.com^ and @@||example.com^ are kept, as
Warda does; a @@ line is an exception of its own list only: the name and
its subdomains are removed from that list, never from what other sources
give to the same category):
[[source]]
name = "example-adblock"
group = "Tracking & Telemetry List"
url = "https://example.org/filters.txt"
format = "adblock"
licence = "GPL-3.0-only"
homepage = "https://example.org/"
categories = ["tracking"]
The archive of the Université Toulouse Capitole (UT1): a folder per category; the map says where each folder goes. Folders left out are ignored (the build prints their names):
[[source]]
name = "ut1"
group = "Other Lists"
url = "https://dsi.ut-capitole.fr/blacklists/download/blacklists.tar.gz"
format = "ut1"
licence = "CC-BY-SA-4.0"
homepage = "https://dsi.ut-capitole.fr/blacklists/"
max_mb = 200
[source.map]
drogue = ["drugs"]
warez = ["piracy"]
arjel = ["except:gambling"] # removed from gambling
shopping = ["type:e-commerce"] # a site type
Every source names its group, exactly as written below (the closed
list is in taxonomy.toml, [[source_group]], and sources.toml is
sorted by group, a comment header for each). A group only files and names
the lists; where their domains go is decided by categories or the map.
The 16 names are all distinct (the only near pairs are told apart below:
Fraud/Scam, Malware/Malicious).
| Group | What goes there |
|---|---|
| Hosts list | general hosts files that mix several kinds (ads, trackers, malware) |
| Abuse List | sites known for abusive or deceptive practices |
| Drugs List | drugs and alcohol |
| Fraud List | fake shops, financial and payment fraud |
| Malware List | malware distribution and command & control servers |
| Phishing List | pages that steal credentials or bank data |
| Ransomware List | ransomware distribution, payment and command servers |
| Scam List | fake support, fake giveaways, deceptive offers |
| Tracking & Telemetry List | trackers, analytics, telemetry of devices and apps |
| Redirect List | URL shorteners and redirectors that hide the destination |
| Gambling List | gambling and betting |
| Porn List | pornography and adult content |
| Suspicious Lists | not proven bad but risky: newly registered, parked, disposable domains |
| Advertising Lists | advertising, and ad-and-tracker lists where ads dominate |
| Malicious Lists | threat feeds that mix several threats (malware, phishing, C&C, cryptojacking) |
| Other Lists | everything else: bypass, piracy, social networks, archives such as UT1 |
The group shows in the build log, in manifest.json (per source) and in
LICENSES.md.
What categories (or the map) accepts:
| Value | Meaning |
|---|---|
adult |
a blocking category (see below) |
type:blog |
a site type (classification only) |
except:gambling |
the domains are removed from that category |
Names are read as Warda reads them: lower case, no trailing dot, no IP
address, no bare TLD, no localhost. A name that is not ASCII is dropped,
as Warda drops it: international names count only in their xn-- form.
Other fields: author (for the attribution), enabled = false (keep the
block, do not use it), max_mb (biggest download accepted, 50 by
default).
The licence is the exact SPDX id of the list, one of: GPL-3.0-only,
GPL-3.0-or-later, GPL-2.0-only, GPL-2.0-or-later, CC-BY-SA-4.0,
CC-BY-SA-3.0, CC-BY-4.0, CC-BY-3.0, CC0-1.0, Unlicense, MIT,
ISC, 0BSD, BSD-2-Clause, BSD-3-Clause, Apache-2.0. Anything else
is refused (unknown or proprietary licences, NC, ND); GPL-3.0 alone
is ambiguous and refused. Each accepted licence has its text in
licenses/. The accepted ids are written in one place, licences of
taxonomy.toml: the build reads them there and publishes them in
manifest.json; an NC or ND licence is refused even there, and an id
without its text in licenses/ fails --check.
A source that cannot be downloaded fails the whole build: nothing is
published and the boxes keep the lists of the day before. A source gone
for good gets enabled = false.
The files of the owner
| File | Role |
|---|---|
extra/<category>.txt |
domains added to a category (extra/jobsearch.txt → jobsearch.txt); extra/types/<type>.txt for a site type. They win over the exceptions of the sources. |
allow.txt |
domains never listed anywhere, with their subdomains. A list that holds a parent still blocks them on the boxes: the build prints a warning. |
protect.txt |
the safety net: infrastructure that breaks devices when blocked (search engines, updates, time servers, connectivity checks, certificate checks, the domains of Warda). A protected name and its parents are removed from every list, and the build reports it (PROTECTED in the log, protected_removed in manifest.json). |
One domain per line, # starts a comment. A line that is not a valid
domain fails the build (a typo is never ignored silently).
The files of warda-analyst
The service warda-analyst writes two kinds of files, by a commit on
develop, each time its team approves or withdraws a name. They are never
edited by hand, and are not the files of the owner (analyst/README.md).
| File | Role |
|---|---|
analyst/add/<category>.txt |
names the analysis of Warda confirmed, added to that category as extra/<category>.txt does (a category that public sources may fill: never csam, never base) |
analyst/remove.txt |
names confirmed as harmless and wrongly blocked: removed with their subdomains from the protection lists only, the categories of remove_from ([analyst] in taxonomy.toml: ads, tracking, phishing, security) |
A removal never weakens a content category: the name leaves the
categories of remove_from, the bundles built from them (base) and
what includes brings into them (security includes phishing), and
stays in every other category (adult, gambling…) and in the site
types. The team clears a name because it is no threat and no tracker,
which says nothing of what the site is about. Only allow.txt removes a
name from every list.
Lines starting with # are comments; every other line is exactly one
name in its normalised form (lower case, no trailing dot, no space, no
empty line, nothing after the name), and the names are sorted (in the
order of their bytes) and unique.
--check refuses anything else; a name added and removed at once,
whichever covers the other (the same name in an add file and in
remove.txt, an added name under a removed one, a removed name under an
added one); more names than the budgets [analyst] of taxonomy.toml,
max_add (50,000, all the add files together) and max_remove
(20,000); and anything in analyst/ that is not README.md,
remove.txt or add/<category>.txt (a file the build would not read
must not look published). A missing file or directory means no name. A
refused commit is not promoted to main: the daily build goes on with
the files of main.
Files that passed the check never stop the build; what it does not apply
is said in the log and in manifest.json (analyst):
- A line of
remove.txtis applied only if it takes at mostmax_effectnames (100,[analyst]) out of the lists it applies to, itself and its subdomains, those lists together. The check cannot know it:co.ukorgithub.ioare valid names. A line over the limit is left out, with a warning, and listed inanalyst.skipped. - A removed name that one of those categories still blocks, because a
source or
extra/lists a parent of it, is listed inanalyst.covered(1,000 entries at most), with a warning. - The budgets and the guard are checked on the lists built without
these files: added names cannot hide a collapse of the sources. If a
list goes over its budget or over the size limit, or trips the guard,
only once these files are applied, the lists are built without any of
them:
analyst.statussaysignored: <why>and the logIGNORED. What fails without them fails as before.
allow.txt and protect.txt still win over an added name. The added
names are attributed to the source warda-analyst
(https://warda-dns.com) in the headers of the lists, in LICENSES.md and
in manifest.json (sources of each file): own data of this repository,
shown as maintainer as the extra/ files are, which adds no licence to
a file. No source of sources.toml may take that name.
The two taxonomies
Both are in taxonomy.toml, with French and English labels.
Blocking categories: what a Warda user switches on or off. The field
warda is the category of Warda the list feeds.
| id | Content | Warda |
|---|---|---|
ads |
Publicité / advertising | base |
tracking |
Pistage / tracking | base |
phishing |
Hameçonnage / phishing (also merged into security) |
security |
security |
Malware, phishing, command & control, cryptojacking, scams and counterfeits | security |
bypass |
Proxies, VPN, Tor, public encrypted DNS | bypass |
csam |
Child sexual abuse material: always empty, see below | csam |
hate |
Terrorism apology, hate and discrimination | hate (new) |
piracy |
Streaming, direct download, torrents | piracy |
gambling |
Illegal gambling (not approved by the French ANJ) | gambling |
social |
Réseaux sociaux / social networks | social |
video |
Streaming et vidéo | video |
games |
Jeux vidéo en ligne | games |
jobsearch |
Sites de recrutement (business policy) | jobsearch (new) |
adult |
Pornography and adult content | adult |
violence |
Weapons and violence | violence |
drugs |
Drugs and alcohol | drugs |
newdomains |
Domains registered in the last days (off by default in Warda) | newdomains |
base = ads ∪ tracking: the list every Warda box blocks by default.
Each category keeps its reason (cybersecurity, law, workplace
productivity, ethics, advertising and tracking) and the description of the
owner.
Site types: 30 kinds of sites in five groups (commercial and business,
content and information, community, directories and services, personal
and entertainment), from e-commerce to kids. They classify a domain
for the statistics; a site type is never blocked by itself.
Outputs
Published on the branch dist, at
https://gitea.tips-of-mine.com/warda-dns/warda-lists/raw/branch/dist/:
| File | Content |
|---|---|
<category>.txt |
one per blocking category, e.g. .../raw/branch/dist/security.txt |
base.txt |
ads + tracking |
types/<type>.txt |
one per site type |
manifest.json |
per file: entries (count), names (names), SHA-256, size, sources, licence of the file and exact licences of its data, address; per source: status and figures; the taxonomy, the names of the owner and the counts of warda-analyst (below) |
LICENSES.md |
the attribution of every source and the licence of every file |
licenses/<SPDX id>.txt |
the text of every licence used (GPL: full text; CC: notice and link to the legal code) |
README.md |
the files with their entries, names and addresses; the sources with their figures |
Each list: one domain per line, sorted, unique; a domain blocks its
subdomains too (a subdomain whose parent is listed is removed). The header
lines start with #: title, file, generation time (# Generated:, UTC,
such as 2026-10-04T03:21:07Z), entries (# Domains:), names
(# Names:), licence (and the exact licences of the data), sources.
Two figures for each file:
- Domains (
count): the entries of the file, what a box loads. - Names (
names): the distinct names the file stands for: the names of its sources, ofextra/and ofanalyst/add/, onceallow.txt,analyst/remove.txtandprotect.txtare applied, before the names a parent already covers are dropped. Never smaller than the entries. It is the figure to compare with a tool that counts one per name.
And for each source read, beside domains (the names read from it),
unique: how many of its names no other source of the build brings (0
for a disabled or failed source). domains - unique is what it shares.
A name that the UT1 archive holds in two folders counts twice in domains
and once in unique.
manifest.json (schema 1) also holds what warda-analyst reads before it
proposes a list or a name:
| Key | Content |
|---|---|
taxonomy.groups |
the groups of sources, in the order of taxonomy.toml |
taxonomy.categories |
per category: file, max_domains, public_sources |
taxonomy.bundles |
per bundle, its categories (base: ads, tracking) |
taxonomy.formats |
the formats a source may have |
taxonomy.licences |
the SPDX ids a source may carry |
owner |
the names of allow.txt and of protect.txt |
analyst.status |
applied, or ignored: <why> when the lists were built without the files of analyst/ (add is then empty, remove 0) |
analyst.add |
per category, the names of its analyst/add/ file (only the files that hold some) |
analyst.remove |
the lines of analyst/remove.txt applied |
analyst.skipped |
the lines not applied because they would take more than max_effect names out: {"name", "names"}, sorted by name |
analyst.covered |
the removed names a category of remove_from still blocks through a parent: {"name", "parent", "file"}, sorted, 1,000 at most |
analyst.max_add, analyst.max_remove, analyst.max_effect |
the limits of [analyst] in taxonomy.toml |
analyst.remove_from |
the categories analyst/remove.txt applies to, in the order of taxonomy.toml |
Two limits fail the build, with the file named, even with force:
- Size: Warda refuses a list over 128 MiB. A file over 120 MiB fails
(
--max-file-mbchanges the limit): the category is then split (some sources moved to a new category) or a source removed. - Budget of domains: the lists live in the memory of the boxes, a
Raspberry Pi among them. A category (or
base) may hold at mostmax_domainsnames:default_max_domainsoftaxonomy.toml(1.5 million), 3.5 million fornewdomains. Over it, disable or replace a big source, or raisemax_domainsof that category knowingly. Site types have no budget.
A limit reached only because of the names of warda-analyst does not fail the build: the lists are built without them (above).
The choices made for the budgets (comments in sources.toml): The Block
List Project malware list is disabled (2.6 million names, largely stale,
covered by HaGeZi TIF medium); the UT1 folders adult (5 million loose
names), malware, cryptojacking, stalkerware and ddos are not mapped
(security.txt = TIF medium, phishing, scam, fraud, ransomware, crypto
and three small lists: KADhosts, Spam404, uBlock badware); The Block List
Project abuse list is not used (255,000 new names would bring
security.txt to 1.48 million); newdomains uses the last 7 days only
(with days 8 to 14 it would hold 6.3 million names).
The branch dist is a single commit, replaced at each build: it never
grows.
Licences
The code of this repository is under the AGPL-3.0 (LICENSE).
The lists keep the licences of their sources, reported with their exact
SPDX id (CC BY-SA 3.0 and 4.0 are never merged in the report). A file that
contains GPL data (HaGeZi, GPL-3.0-only) is distributed under the GPL;
CC BY-SA 4.0 data (UT1) needs its attribution, and a file with CC BY-SA 4.0
and GPL-3.0 data is distributed under the GPL-3.0 (one-way compatible);
The Block List Project is public domain (Unlicense). dist/LICENSES.md
gives the details and dist/licenses/ the texts.
Mixes that cannot be distributed fail the build, before any download
(--check too): GPL-2.0-only with GPL-3.0-*, CC-BY-SA-4.0,
CC-BY-4.0 or Apache-2.0; CC-BY-SA-3.0 or CC-BY-3.0 with any GPL.
Child sexual abuse material
The category csam exists, but csam.txt is always empty: no public list
exists. Such lists are only given to vetted organisations and internet
providers, and must never be published; this repository and its lists are
public, so the build refuses any source or extra/ file for csam. A Warda
box that receives such a list configures it locally.
The guard
The CI compares the new build with the manifest.json published before:
if a category loses more than half of its domains (lists of at least 100
domains), or if base.txt has fewer than 50,000 domains, nothing is
published. The build by hand (Run workflow) has two options: force
(publish anyway) and allow_partial (go on when a source fails). The
guard counts the lists without the names of warda-analyst; removals of
analyst/remove.txt that would trip it are not published (above).
Build by hand
Python 3.11 or later, standard library only.
python3 scripts/build.py --check # validate the TOML, the owner and analyst files
python3 scripts/build.py # download and build dist/
python3 scripts/build.py --offline DIR # read each source from DIR/<name>
python3 -m unittest discover -s scripts -v # the tests (no network)
Other options: --previous URL-or-file (the guard), --force,
--allow-partial, --min-base N, --max-file-mb N, --out DIR,
--base-url URL, --no-source-stats (does not count unique, then absent
from manifest.json).
The build keeps every list in memory (millions of names for UT1 and the
newly registered domains): give it a few GB. The CI limits its container
with docker run --memory 3g --memory-swap 3g (BUILD_MEMORY); a build
killed with exit code 137 needs more. Counting unique is the only figure
that costs memory, and a few seconds. Its table grows by steps, with the
number of distinct names of the sources (printed in the log): 176 MiB at
most up to 5.6 million names, 352 MiB from there to 10 million (measured
by the tests: WARDA_LISTS_STATS_NAMES=6000000). Over 10 million it would
take 704 MiB: the build then gives unique up by itself, with a warning,
as --no-source-stats does on request.
CI
.gitea/workflows/ci.yml, on every push to develop or main, every
pull request to main, every day at 03:17 UTC and by hand:
- test: the unit tests, the validation of
taxonomy.toml,sources.toml, the files of the owner and those of warda-analyst, and an offline build fromtests/fixtures/. Python runs in a container without network. - build (pushes of
main, the daily run and the run by hand, whatever the branch they start from: Gitea runs the schedule on the default branch): it always checks outmain, builds with the memory limit, runs the guard against the previous manifest, then pushesdist/as a single commit without parent on the branchdist(forced). A concurrency group keeps one publication at a time. - Promote to main: a push of
developwhose tests pass is pushed onmain(fast-forward only).
Two secrets, in the settings of the repository:
| Secret | What |
|---|---|
RELEASE_TOKEN |
application token allowed to write the repository: the promotion to main |
LISTS_TOKEN |
application token with write access to warda-lists only: the force-push of the branch dist |
The branch dist must accept a forced push from the owner of
LISTS_TOKEN (no protection on it, or a protection that allows it).