Instructions to use SlayerLab/NERGAL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SlayerLab/NERGAL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="SlayerLab/NERGAL")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("SlayerLab/NERGAL") model = AutoModelForTokenClassification.from_pretrained("SlayerLab/NERGAL", device_map="auto") - Notebooks
- Google Colab
- Kaggle
NERGAL 1.2.0
Named Entity Recognition with Grounded Additive Labels
SlayerLab hybrid PII cleaner for Polish. Not a chat model. Not a drop-in pipeline("token-classification").
TL;DR
Python rules do the identifiers they can prove. A transformer NER head adds phone and other PII spans the regex misses. The cleaner unions the two on the original text, then replaces hits with [Telefon] or [PII]. Phone spans from both follow phone policy v3: one span per number of 7+ digits, and shorter numbers are not masked.
- Ground:
scrub_pii.pyrules - Additive labels: XLM-RoBERTa-large token classifier (epoch 5 of 7), BIO tags
phone/pii, threshold 0.95 - Throughput: about 80k chars/s on one RTX 4090 with
scrub_many+dtype="float16"and 3 processes - Changes:
CHANGELOG.md
What NERGAL detects — and what it does not
NERGAL masks contact details and selected identifiers in Polish text. It is not a general-purpose anonymizer: names, postal addresses and other personal information can remain in the output. Its corpus-masking policy also includes public, institutional and company contacts and identifiers.
Detection scope
These are target categories, not a guarantee that every occurrence or format is detected.
| Category | Values in scope | Replacement |
|---|---|---|
| Phone contacts | Phone, fax and SMS contact numbers of 7+ digits (country and area codes and keypad letters count), including foreign and vanity numbers, one span per number; an extension or alternative ending stays inside its number. Emergency, helpline, service and other numbers under 7 digits are not masked on their own | [Telefon] |
| Email addresses, including recoverable broken or incomplete addresses | [PII] |
|
| Personal identifiers | PESEL, passport and identity-document numbers | [PII] |
| Organization identifiers | NIP/VAT, REGON, KRS, GEMI, LEI and equivalent foreign company registration numbers; labelled DUNS, BDO and RPWDL register-book numbers | [PII] |
| Financial identifiers | Bank/account numbers, including Polish accounts and foreign IBANs | [PII] |
| Property identifiers | Land-register (księga wieczysta, KW) numbers | [PII] |
| Electronic contacts and access | e-Doręczenia and ePUAP addresses, explicitly labelled numeric access PINs (including pin= URL values), GG account IDs |
[PII] |
Context matters: a number that resembles a phone or identifier is not automatically in scope. Rules use labels, format checks and, for some unlabelled identifiers, checksums; the model adds contextual detections. Coverage varies by category, and the aggregate benchmark below does not establish recall for every category.
Outside detection scope
NERGAL is not designed to remove:
- Personal names, including private individuals and public officials; organization names.
- Postal/street addresses, dates of birth and ages.
- Social-media handles, ordinary URLs and filenames. An in-scope value inside a URL, such as a labelled numeric PIN, can still be masked.
- Vehicle registration plates and generic serial, model or version codes.
- Document, case, article, funding and procurement references, including procedure UUIDs; ISBNs, ORCIDs, TERYT codes, EAN/GTIN product codes and CNIL website-registration references.
- Prices, list numbers, generic labels without values, clearly fictitious examples and already-redacted placeholders.
These are intended exclusions; false positives can still mask some of this content.
Known gaps in 1.2.0
Unlabelled phones and identifiers (the rules take a bare phone only in the grouped national forms 601 234 567 and 22 123 45 67; a plain 601234567 needs a label or the model), unusual formatting and damaged text can escape detection. VINs and obfuscated emails (such as name (at) domain.pl) are approved annotation targets, but that approval alone does not establish reliable support in the released 1.2.0 model. Do not rely on it to remove them consistently. An email with a lowercase word glued onto its domain (jan@firma.plkontakt) is masked together with that word, and text glued before an address can be masked with it. A phone written as short parts joined by two slashes (12 / 345 / 678) is left as text. An identifier with a [PII] / [Telefon] placeholder inside it or right before it can be missed.
Versions
841-dev (below), union at 0.95. Same weights and API is a patch; new capability is minor; API, threshold, or weight recipe is major. 1.2.0 restated the gold to phone policy v3 and rescored 1.1.2 on it. The restatement and the mask-time cut apply the same policy, so the gain measures agreement with it, not an independent test.
| Version | Whole /315 | Residual | Rules FP | Union FP | Char P | Char R | What changed |
|---|---|---|---|---|---|---|---|
| 1.1.2 | 303 | 9 | 83 | 333 | 94.37% | 97.83% | Email boundary rules |
| 1.2.0 | 303 | 9 | 24 | 80 | 98.59% | 97.83% | Phone policy v3 at mask time, rules and model spans |
1.0.0–1.1.2 were scored on the earlier 354-span gold (whole 323 → 326, union FP 133 → 80); per-version numbers are in CHANGELOG.md.
Cue-less phone test
841-dev holds few phones without a contact label. This targeted set does: 150 web passages selected for them, labelled and frozen before any prediction. 1.1.1 added the grouped national forms and raised phones wholly masked from 37 to 66 of 111, losing none. Its benchmark half, restated to phone policy v3 (70 passages): 1.2.0 wholly masks 62 of 83 values, with 33 false characters. The set is enriched by selection and has a single reviewer, so it says nothing about how common such phones are.
841-dev
841 development passages, 215 with gold PII, 315 spans (130 phone, 185 other PII). Sources match Dynaword (EUR-Lex, HPLT, Wikipedia, parliamentary and government text, plus smaller news/literary slices). Labels mix unchanged silver with human review. Since 1.2.0 the gold follows phone policy v3 (short and emergency numbers dropped, joined numbers split) and includes reviewed corrections; before, it held 354 spans. The files contain real identifiers, so they are not released with the weights.
Compared with other systems
Same split and gold. Every model runs through the 1.2.0 recipe: its spans at 0.95, phone spans cut by phone policy v3, unioned with the 1.2.0 rules. Naked is the model alone. Residual is passages with any gold character left unmasked. Character scores are gold vs masked characters.
| System | Mode | Whole /315 | Residual | False chars | Char P | Char R |
|---|---|---|---|---|---|---|
Regex (scrub_pii) |
rules | 268 | 28 | 24 | 99.53% | 89.85% |
| GLiNER 2.5-multi zero-shot | naked | 84 | 151 | 451 | 72.96% | 21.33% |
| GLiNER 2.5-multi zero-shot | ∪ regex | 275 | 25 | 475 | 91.68% | 91.73% |
| Fine-tuned GLiNER (previous incumbent) | naked | 263 | 28 | 15 | 99.70% | 88.80% |
| Fine-tuned GLiNER (previous incumbent) | ∪ regex | 299 | 12 | 39 | 99.30% | 97.18% |
| XLM-R epoch 5 | naked | 275 | 28 | 67 | 98.72% | 90.25% |
| NERGAL 1.2.0 | ∪ regex | 303 | 9 | 80 | 98.59% | 97.83% |
Zero-shot GLiNER 2.5 is not competitive, especially on non-phone PII (7/185 whole, naked). The fine-tuned GLiNER that NERGAL replaced is the more precise union (39 false characters vs 80) and the stronger phone model (128/130 whole, naked, vs 115), but it masks 4 fewer values whole and leaves 12 residual passages vs 9; XLM-R is the stronger model on other PII (160/185 vs 135). NERGAL 1.2.0 wholly masks 124/130 phones and 179/185 other PII; exact-span precision 91.05%, recall 93.65%, F1 92.33%.
XLM-R and epoch 5 were chosen in September 2026, on the earlier gold and before the phone-policy cut. Fine-tuned GLiNER, HerBERT-large and XLM-R-large were trained on the same split; only XLM-R passed the content-preservation gate, and epoch 5 at 0.95 (seed 202609160) was the only one of 133 epoch/threshold points that covered more than the fine-tuned GLiNER union of that time while adding no false-mask characters it did not already make. Under the 1.2.0 rules, gold and phone cut, that no longer holds: the GLiNER union makes fewer false masks (table above).
Load
Regex runs on the original text, the model adds spans at 0.95, then the two are unioned and replaced with [Telefon] / [PII].
from pathlib import Path
from huggingface_hub import snapshot_download
root = Path(snapshot_download("SlayerLab/NERGAL"))
import sys
sys.path.insert(0, str(root))
from nergal import Nergal
nergal = Nergal.from_pretrained(root, local_files_only=True)
masked, counts = nergal.scrub(text)
For many texts, nergal.scrub_many(texts) (or predict_many for the raw model spans) batches windows across texts: about 2× the throughput of calling scrub in a loop on a CUDA GPU. from_pretrained(..., dtype="float16") casts the weights at load time on CUDA or MPS: another 1.9× on an RTX 4090, and 841-dev union numbers are unchanged, but scores are not bit-identical to float32 (0.95 spans can differ in rare cases). For a corpus, run 2–3 processes per GPU, because float16 inference is limited by CPU-side tokenization. Measurements are in CHANGELOG.md.
hybrid.json records the version, threshold, 841-dev eval block and weight hashes. test_nergal.py is synthetic (no corpus text): python -m unittest test_nergal.
Base weights: FacebookAI/xlm-roberta-large revision c23d21b0620b635a76227c604d44e43a9f0ee389 (MIT).
- Downloads last month
- 202
Model tree for SlayerLab/NERGAL
Base model
FacebookAI/xlm-roberta-large