hcornet 74c0fdf934
CI / test (push) Successful in 13s
CI / security (push) Successful in 5s
CI / Promote to main (push) Skipped
CI / Build and publish the lists (push) Successful in 3m34s
ci: every push is tracked in GLPI
2026-10-08 15:56:01 +02:00
2026-10-08 15:56:01 +02:00
2026-09-27 15:54:22 +02:00
2026-09-27 15:54:22 +02:00
2026-09-27 15:54:22 +02:00
2026-09-29 16:07:34 +02:00
2026-09-27 15:54:22 +02:00
2026-09-29 16:07:34 +02:00
2026-10-04 18:50:04 +02:00

warda-lists

The lists of domains that every Warda box downloads: many public sources gathered in one place, sorted into categories, cleaned and deduplicated every day by the CI, then published at one address.

Before, each Warda box downloaded the third-party lists itself. Now the sources are chosen here, once, and the boxes download the result.

The documentation of the Warda project is in the repository warda-dns/warda-docs, file lists/README.md.

How it works

  1. sources.toml names the lists to download (the only file to edit to add one).
  2. The CI downloads them every day, reads them the way Warda reads them, puts each domain in its categories, adds the domains of extra/ and of analyst/add/, removes those of allow.txt, of protect.txt and, in the protection lists only, of analyst/remove.txt, and removes the subdomains already covered by a parent.
  3. The result, dist/, is published on the branch dist. The boxes download it from there.

Add a list

Edit sources.toml, add a block, push on develop. That is all: the CI checks the file, and once on main the list is in the next build. The comments at the top of sources.toml explain every field.

Before adding a list, read its licence in its repository and download its address once (curl -A "warda-lists/1.0"). The lists that were asked for and not added (non-commercial or unknown licence, frozen list, wrong fit for the categories, budget, download not tested) are kept in PENDING.md, each with what would unlock it.

A list of one domain per line:

[[source]]
name = "example-malware"                # unique, lower case, digits and "-"
group = "Malware List"                  # where it is filed (list below)
url = "https://example.org/malware.txt"
format = "domains"
licence = "MIT"                         # required: exact SPDX id
homepage = "https://example.org/"
categories = ["security"]

A hosts file (0.0.0.0 example.com):

[[source]]
name = "example-hosts"
group = "Hosts list"
url = "https://example.org/hosts"
format = "hosts"
licence = "CC-BY-4.0"
homepage = "https://example.org/"
categories = ["ads", "tracking"]

An Adblock list (only ||example.com^ and @@||example.com^ are kept, as Warda does; a @@ line is an exception of its own list only: the name and its subdomains are removed from that list, never from what other sources give to the same category):

[[source]]
name = "example-adblock"
group = "Tracking & Telemetry List"
url = "https://example.org/filters.txt"
format = "adblock"
licence = "GPL-3.0-only"
homepage = "https://example.org/"
categories = ["tracking"]

The archive of the Université Toulouse Capitole (UT1): a folder per category; the map says where each folder goes. Folders left out are ignored (the build prints their names):

[[source]]
name = "ut1"
group = "Other Lists"
url = "https://dsi.ut-capitole.fr/blacklists/download/blacklists.tar.gz"
format = "ut1"
licence = "CC-BY-SA-4.0"
homepage = "https://dsi.ut-capitole.fr/blacklists/"
max_mb = 200

[source.map]
drogue = ["drugs"]
warez = ["piracy"]
arjel = ["except:gambling"]   # removed from gambling
shopping = ["type:e-commerce"] # a site type

Every source names its group, exactly as written below (the closed list is in taxonomy.toml, [[source_group]], and sources.toml is sorted by group, a comment header for each). A group only files and names the lists; where their domains go is decided by categories or the map. The 16 names are all distinct (the only near pairs are told apart below: Fraud/Scam, Malware/Malicious).

Group What goes there
Hosts list general hosts files that mix several kinds (ads, trackers, malware)
Abuse List sites known for abusive or deceptive practices
Drugs List drugs and alcohol
Fraud List fake shops, financial and payment fraud
Malware List malware distribution and command & control servers
Phishing List pages that steal credentials or bank data
Ransomware List ransomware distribution, payment and command servers
Scam List fake support, fake giveaways, deceptive offers
Tracking & Telemetry List trackers, analytics, telemetry of devices and apps
Redirect List URL shorteners and redirectors that hide the destination
Gambling List gambling and betting
Porn List pornography and adult content
Suspicious Lists not proven bad but risky: newly registered, parked, disposable domains
Advertising Lists advertising, and ad-and-tracker lists where ads dominate
Malicious Lists threat feeds that mix several threats (malware, phishing, C&C, cryptojacking)
Other Lists everything else: bypass, piracy, social networks, archives such as UT1

The group shows in the build log, in manifest.json (per source) and in LICENSES.md.

What categories (or the map) accepts:

Value Meaning
adult a blocking category (see below)
type:blog a site type (classification only)
except:gambling the domains are removed from that category

Names are read as Warda reads them: lower case, no trailing dot, no IP address, no bare TLD, no localhost. A name that is not ASCII is dropped, as Warda drops it: international names count only in their xn-- form.

Other fields: author (for the attribution), enabled = false (keep the block, do not use it), max_mb (biggest download accepted, 50 by default).

The licence is the exact SPDX id of the list, one of: GPL-3.0-only, GPL-3.0-or-later, GPL-2.0-only, GPL-2.0-or-later, CC-BY-SA-4.0, CC-BY-SA-3.0, CC-BY-4.0, CC-BY-3.0, CC0-1.0, Unlicense, MIT, ISC, 0BSD, BSD-2-Clause, BSD-3-Clause, Apache-2.0. Anything else is refused (unknown or proprietary licences, NC, ND); GPL-3.0 alone is ambiguous and refused. Each accepted licence has its text in licenses/. The accepted ids are written in one place, licences of taxonomy.toml: the build reads them there and publishes them in manifest.json; an NC or ND licence is refused even there, and an id without its text in licenses/ fails --check.

A source that cannot be downloaded fails the whole build: nothing is published and the boxes keep the lists of the day before. A source gone for good gets enabled = false.

The files of the owner

File Role
extra/<category>.txt domains added to a category (extra/jobsearch.txt → jobsearch.txt); extra/types/<type>.txt for a site type. They win over the exceptions of the sources.
allow.txt domains never listed anywhere, with their subdomains. A list that holds a parent still blocks them on the boxes: the build prints a warning.
protect.txt the safety net: infrastructure that breaks devices when blocked (search engines, updates, time servers, connectivity checks, certificate checks, the domains of Warda). A protected name and its parents are removed from every list, and the build reports it (PROTECTED in the log, protected_removed in manifest.json).

One domain per line, # starts a comment. A line that is not a valid domain fails the build (a typo is never ignored silently).

The files of warda-analyst

The service warda-analyst writes two kinds of files, by a commit on develop, each time its team approves or withdraws a name. They are never edited by hand, and are not the files of the owner (analyst/README.md).

File Role
analyst/add/<category>.txt names the analysis of Warda confirmed, added to that category as extra/<category>.txt does (a category that public sources may fill: never csam, never base)
analyst/remove.txt names confirmed as harmless and wrongly blocked: removed with their subdomains from the protection lists only, the categories of remove_from ([analyst] in taxonomy.toml: ads, tracking, phishing, security)

A removal never weakens a content category: the name leaves the categories of remove_from, the bundles built from them (base) and what includes brings into them (security includes phishing), and stays in every other category (adult, gambling…) and in the site types. The team clears a name because it is no threat and no tracker, which says nothing of what the site is about. Only allow.txt removes a name from every list.

Lines starting with # are comments; every other line is exactly one name in its normalised form (lower case, no trailing dot, no space, no empty line, nothing after the name), and the names are sorted (in the order of their bytes) and unique. --check refuses anything else; a name added and removed at once, whichever covers the other (the same name in an add file and in remove.txt, an added name under a removed one, a removed name under an added one); more names than the budgets [analyst] of taxonomy.toml, max_add (50,000, all the add files together) and max_remove (20,000); and anything in analyst/ that is not README.md, remove.txt or add/<category>.txt (a file the build would not read must not look published). A missing file or directory means no name. A refused commit is not promoted to main: the daily build goes on with the files of main.

Files that passed the check never stop the build; what it does not apply is said in the log and in manifest.json (analyst):

  • A line of remove.txt is applied only if it takes at most max_effect names (100, [analyst]) out of the lists it applies to, itself and its subdomains, those lists together. The check cannot know it: co.uk or github.io are valid names. A line over the limit is left out, with a warning, and listed in analyst.skipped.
  • A removed name that one of those categories still blocks, because a source or extra/ lists a parent of it, is listed in analyst.covered (1,000 entries at most), with a warning.
  • The budgets and the guard are checked on the lists built without these files: added names cannot hide a collapse of the sources. If a list goes over its budget or over the size limit, or trips the guard, only once these files are applied, the lists are built without any of them: analyst.status says ignored: <why> and the log IGNORED. What fails without them fails as before.

allow.txt and protect.txt still win over an added name. The added names are attributed to the source warda-analyst (https://warda-dns.com) in the headers of the lists, in LICENSES.md and in manifest.json (sources of each file): own data of this repository, shown as maintainer as the extra/ files are, which adds no licence to a file. No source of sources.toml may take that name.

The two taxonomies

Both are in taxonomy.toml, with French and English labels.

Blocking categories: what a Warda user switches on or off. The field warda is the category of Warda the list feeds.

id Content Warda
ads Publicité / advertising base
tracking Pistage / tracking base
phishing Hameçonnage / phishing (also merged into security) security
security Malware, phishing, command & control, cryptojacking, scams and counterfeits security
bypass Proxies, VPN, Tor, public encrypted DNS bypass
csam Child sexual abuse material: always empty, see below csam
hate Terrorism apology, hate and discrimination hate (new)
piracy Streaming, direct download, torrents piracy
gambling Illegal gambling (not approved by the French ANJ) gambling
social Réseaux sociaux / social networks social
video Streaming et vidéo video
games Jeux vidéo en ligne games
jobsearch Sites de recrutement (business policy) jobsearch (new)
adult Pornography and adult content adult
violence Weapons and violence violence
drugs Drugs and alcohol drugs
newdomains Domains registered in the last days (off by default in Warda) newdomains

base = ads ∪ tracking: the list every Warda box blocks by default. Each category keeps its reason (cybersecurity, law, workplace productivity, ethics, advertising and tracking) and the description of the owner.

Site types: 30 kinds of sites in five groups (commercial and business, content and information, community, directories and services, personal and entertainment), from e-commerce to kids. They classify a domain for the statistics; a site type is never blocked by itself.

Outputs

Published on the branch dist, at https://gitea.tips-of-mine.com/warda-dns/warda-lists/raw/branch/dist/:

File Content
<category>.txt one per blocking category, e.g. .../raw/branch/dist/security.txt
base.txt ads + tracking
types/<type>.txt one per site type
manifest.json per file: entries (count), names (names), SHA-256, size, sources, licence of the file and exact licences of its data, address; per source: status and figures; the taxonomy, the names of the owner and the counts of warda-analyst (below)
LICENSES.md the attribution of every source and the licence of every file
licenses/<SPDX id>.txt the text of every licence used (GPL: full text; CC: notice and link to the legal code)
README.md the files with their entries, names and addresses; the sources with their figures

Each list: one domain per line, sorted, unique; a domain blocks its subdomains too (a subdomain whose parent is listed is removed). The header lines start with #: title, file, generation time (# Generated:, UTC, such as 2026-10-04T03:21:07Z), entries (# Domains:), names (# Names:), licence (and the exact licences of the data), sources.

Two figures for each file:

  • Domains (count): the entries of the file, what a box loads.
  • Names (names): the distinct names the file stands for: the names of its sources, of extra/ and of analyst/add/, once allow.txt, analyst/remove.txt and protect.txt are applied, before the names a parent already covers are dropped. Never smaller than the entries. It is the figure to compare with a tool that counts one per name.

And for each source read, beside domains (the names read from it), unique: how many of its names no other source of the build brings (0 for a disabled or failed source). domains - unique is what it shares. A name that the UT1 archive holds in two folders counts twice in domains and once in unique.

manifest.json (schema 1) also holds what warda-analyst reads before it proposes a list or a name:

Key Content
taxonomy.groups the groups of sources, in the order of taxonomy.toml
taxonomy.categories per category: file, max_domains, public_sources
taxonomy.bundles per bundle, its categories (base: ads, tracking)
taxonomy.formats the formats a source may have
taxonomy.licences the SPDX ids a source may carry
owner the names of allow.txt and of protect.txt
analyst.status applied, or ignored: <why> when the lists were built without the files of analyst/ (add is then empty, remove 0)
analyst.add per category, the names of its analyst/add/ file (only the files that hold some)
analyst.remove the lines of analyst/remove.txt applied
analyst.skipped the lines not applied because they would take more than max_effect names out: {"name", "names"}, sorted by name
analyst.covered the removed names a category of remove_from still blocks through a parent: {"name", "parent", "file"}, sorted, 1,000 at most
analyst.max_add, analyst.max_remove, analyst.max_effect the limits of [analyst] in taxonomy.toml
analyst.remove_from the categories analyst/remove.txt applies to, in the order of taxonomy.toml

Two limits fail the build, with the file named, even with force:

  • Size: Warda refuses a list over 128 MiB. A file over 120 MiB fails (--max-file-mb changes the limit): the category is then split (some sources moved to a new category) or a source removed.
  • Budget of domains: the lists live in the memory of the boxes, a Raspberry Pi among them. A category (or base) may hold at most max_domains names: default_max_domains of taxonomy.toml (1.5 million), 3.5 million for newdomains. Over it, disable or replace a big source, or raise max_domains of that category knowingly. Site types have no budget.

A limit reached only because of the names of warda-analyst does not fail the build: the lists are built without them (above).

The choices made for the budgets (comments in sources.toml): The Block List Project malware list is disabled (2.6 million names, largely stale, covered by HaGeZi TIF medium); the UT1 folders adult (5 million loose names), malware, cryptojacking, stalkerware and ddos are not mapped (security.txt = TIF medium, phishing, scam, fraud, ransomware, crypto and three small lists: KADhosts, Spam404, uBlock badware); The Block List Project abuse list is not used (255,000 new names would bring security.txt to 1.48 million); newdomains uses the last 7 days only (with days 8 to 14 it would hold 6.3 million names).

The branch dist is a single commit, replaced at each build: it never grows.

Licences

The code of this repository is under the AGPL-3.0 (LICENSE).

The lists keep the licences of their sources, reported with their exact SPDX id (CC BY-SA 3.0 and 4.0 are never merged in the report). A file that contains GPL data (HaGeZi, GPL-3.0-only) is distributed under the GPL; CC BY-SA 4.0 data (UT1) needs its attribution, and a file with CC BY-SA 4.0 and GPL-3.0 data is distributed under the GPL-3.0 (one-way compatible); The Block List Project is public domain (Unlicense). dist/LICENSES.md gives the details and dist/licenses/ the texts.

Mixes that cannot be distributed fail the build, before any download (--check too): GPL-2.0-only with GPL-3.0-*, CC-BY-SA-4.0, CC-BY-4.0 or Apache-2.0; CC-BY-SA-3.0 or CC-BY-3.0 with any GPL.

Child sexual abuse material

The category csam exists, but csam.txt is always empty: no public list exists. Such lists are only given to vetted organisations and internet providers, and must never be published; this repository and its lists are public, so the build refuses any source or extra/ file for csam. A Warda box that receives such a list configures it locally.

The guard

The CI compares the new build with the manifest.json published before: if a category loses more than half of its domains (lists of at least 100 domains), or if base.txt has fewer than 50,000 domains, nothing is published. The build by hand (Run workflow) has two options: force (publish anyway) and allow_partial (go on when a source fails). The guard counts the lists without the names of warda-analyst; removals of analyst/remove.txt that would trip it are not published (above).

Build by hand

Python 3.11 or later, standard library only.

python3 scripts/build.py --check              # validate the TOML, the owner and analyst files
python3 scripts/build.py                      # download and build dist/
python3 scripts/build.py --offline DIR        # read each source from DIR/<name>
python3 -m unittest discover -s scripts -v    # the tests (no network)

Other options: --previous URL-or-file (the guard), --force, --allow-partial, --min-base N, --max-file-mb N, --out DIR, --base-url URL, --no-source-stats (does not count unique, then absent from manifest.json).

The build keeps every list in memory (millions of names for UT1 and the newly registered domains): give it a few GB. The CI limits its container with docker run --memory 3g --memory-swap 3g (BUILD_MEMORY); a build killed with exit code 137 needs more. Counting unique is the only figure that costs memory, and a few seconds. Its table grows by steps, with the number of distinct names of the sources (printed in the log): 176 MiB at most up to 5.6 million names, 352 MiB from there to 10 million (measured by the tests: WARDA_LISTS_STATS_NAMES=6000000). Over 10 million it would take 704 MiB: the build then gives unique up by itself, with a warning, as --no-source-stats does on request.

CI

.gitea/workflows/ci.yml, on every push to develop or main, every pull request to main, every day at 03:17 UTC and by hand:

  • test: the unit tests, the validation of taxonomy.toml, sources.toml, the files of the owner and those of warda-analyst, and an offline build from tests/fixtures/. Python runs in a container without network.
  • build (pushes of main, the daily run and the run by hand, whatever the branch they start from: Gitea runs the schedule on the default branch): it always checks out main, builds with the memory limit, runs the guard against the previous manifest, then pushes dist/ as a single commit without parent on the branch dist (forced). A concurrency group keeps one publication at a time.
  • Promote to main: a push of develop whose tests pass is pushed on main (fast-forward only).

Two secrets, in the settings of the repository:

Secret What
RELEASE_TOKEN application token allowed to write the repository: the promotion to main
LISTS_TOKEN application token with write access to warda-lists only: the force-push of the branch dist

The branch dist must accept a forced push from the owner of LISTS_TOKEN (no protection on it, or a protection that allows it).

S
Description
Base blocklist build pipeline, updated every 6 hours
Readme AGPL-3.0
178 MiB
0 Stars 1 Watchers 0 Forks
Languages
Python 98.1%
Shell 1.9%