- response_format is omitted by default (Infomaniak rejects the legacy
json_object → HTTP 422); opt in via the new `responseFormat` option
({type:'json_object'} or a json_schema object). BREAKING for endpoints
that relied on the forced json_object.
- parse LLM JSON leniently (tolerate markdown fences / surrounding prose)
- fold strict-coherence rules into DEFAULT_SYSTEM_PROMPT (exact
placeholder<->mapping-key identity, values are originals, strict format,
mask value not adjacent label) → reliable output across models
Verified live against Gemma 4 (google/gemma-4-31B-it, Infomaniak v2): all
demo phrases anonymize with clean round-trips, no custom provider needed.
36 tests passing.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
12 KiB
@mobiletic/anonymizer
Framework-agnostic PII anonymization & pseudonymization for TypeScript/JavaScript.
It replaces personal data in text with stable placeholders before the text leaves your trust boundary (e.g. before sending it to a third-party LLM, log sink, or analytics pipeline), and restores the real values afterwards — including across streamed tokens.
- 🔌 Pluggable LLM detection — catch free-form PII (names, addresses) via any OpenAI-compatible endpoint, or your own provider.
- 🧩 Optional regex fallback — structured identifiers (email, phone, IBAN, …) via configurable
presets (
swiss,generic) or your own. Opt in for graceful degradation, or omit it to fail closed. - 🏷️ Self-describing tokens — every result ships a
legend(PER→Personne,M→Masculin) you can hand to the downstream LLM so it understands the placeholders; the model may coin new abbreviations too. - 🔁 Deterministic coreference — the same person keeps the same id (
[PER_1]) across a question and every retrieved chunk, with automatic de-collision. - 🌊 Streaming-safe — a placeholder split across two stream chunks (
[PER_+1.NOM:M]) is never leaked partially. - 🪶 Zero runtime dependencies, ESM + CJS, fully typed.
Built by Mobiletic.
Install
npm install @mobiletic/anonymizer
Requires Node ≥ 18 (uses native fetch).
Quick start
Regex-only (no LLM, fully deterministic)
import { Anonymizer, presets } from '@mobiletic/anonymizer';
const anonymizer = new Anonymizer({ patterns: presets.swiss });
const { anon, mapping, legend } = await anonymizer.anonymize('Écris à jean@exemple.ch');
// anon -> "Écris à [EMAIL_1]"
// mapping -> { "[EMAIL_1]": "jean@exemple.ch" } (secret — keep on your side)
// legend -> { "EMAIL": "Adresse e-mail" } (safe to share downstream)
anonymizer.deanonymize(anon, mapping); // -> "Écris à jean@exemple.ch"
With an LLM (also catches names, addresses…)
import { Anonymizer, openAICompatibleProvider, presets } from '@mobiletic/anonymizer';
const anonymizer = new Anonymizer({
llm: openAICompatibleProvider({
baseUrl: process.env.LLM_BASE_URL!, // OpenAI, Infomaniak, vLLM, Ollama, …
apiKey: process.env.LLM_API_KEY!,
model: process.env.LLM_MODEL!,
timeoutMs: 3000,
}),
patterns: presets.swiss, // OPTIONAL regex fallback if the LLM is down/misbehaves
});
const { anon, mapping, legend } = await anonymizer.anonymize('Le dossier de Alain Jaccard est complet.');
// anon -> "Le dossier de [PER_1.NOM:M] est complet."
// legend -> { "PER": "Personne", "NOM": "Nom de famille", "M": "Masculin" }
Fallback is opt-in (fail-closed). If a patterns fallback is configured, a failed/timed-out/invalid
LLM call degrades to the regex engine. If you omit patterns, there's nothing to degrade to, so the call
throws an AnonymizationError (with the underlying error as .cause) — it never silently returns
un-anonymized text. With no fallback the pre-filter is also bypassed, so every non-trivial call consults
the LLM (more calls, no leaks). At least one of llm or patterns is required.
openAICompatibleProvider retries transient failures (network error, timeout, HTTP 429/5xx) before
giving up; 4xx and malformed responses are not retried. Tune with timeoutMs (per attempt, default 3000),
retries (default 1 → 2 attempts), and retryDelayMs (linear backoff, default 250). Worst-case latency
is (retries + 1) × timeoutMs, so keep retries low on latency-sensitive paths.
It works with Infomaniak and other open-model endpoints out of the box: response_format is omitted
by default (Infomaniak rejects the legacy { type: 'json_object' }), and responses are parsed leniently
(a fenced JSON code block or surrounding prose is tolerated). For endpoints that support it, opt in with
responseFormat — e.g. { type: 'json_object' } or a json_schema object.
Streaming de-anonymization
When you stream an LLM answer back to a user, restore real values without ever emitting a half-written placeholder:
const stream = anonymizer.makeStreamDeanonymizer(mapping);
for await (const token of llmTokens) process.stdout.write(stream.push(token));
process.stdout.write(stream.flush());
Anonymizing retrieved chunks consistently (RAG)
anonymizeChunks(chunks, seed) reuses the question's { mapping, legend } so the same person gets the
same id across the question and every chunk, batches the LLM call, and de-collides genuinely new entities:
const q = await anonymizer.anonymize(question);
const { anon, mapping, legend } = await anonymizer.anonymizeChunks(retrievedChunks, q);
// `mapping`/`legend` are the full question ∪ chunks tables;
// pass `mapping` to deanonymize()/makeStreamDeanonymizer(), and `legend` to the downstream LLM.
Configuration
new Anonymizer({
llm?, // LlmProvider — omit for regex-only mode
patterns?, // PatternDef[] — opt-in regex fallback; omit to fail closed
nameHint?, // RegExp flagging likely names so the LLM is consulted (has a default)
logger?, // { warn(msg) } — receives fallback warnings; defaults to no-op
});
// At least one of `llm` or `patterns` must be provided, or the constructor throws.
Presets & custom patterns
import { presets } from '@mobiletic/anonymizer';
presets.swiss; // AVS, IBAN CH, EMAIL, Swiss phone, DATE
presets.generic; // EMAIL, IBAN, credit card, IPv4, phone, DATE
// Compose / extend:
const patterns = [
...presets.generic,
{ tag: 'TICKET', re: /\bJIRA-\d+\b/g }, // patterns must use the global flag
];
Each pattern may carry an optional validate(match) => boolean second stage — a match is only redacted
if it passes. presets.generic uses it for a Luhn check
so arbitrary long digit runs aren't mistaken for credit cards:
{ tag: 'CREDIT_CARD', re: /\b\d(?:[ -]?\d){12,18}\b/g, validate: luhnValid }
The
genericpreset is a best-effort starting point — broad patterns (phone, date, card) can overlap. For production use, prefer a locale-specific preset (presets.swiss) or your own patterns.
Custom LLM provider
Implement LlmProvider to use any backend (Anthropic, a local model, a rules engine…):
import type { LlmProvider } from '@mobiletic/anonymizer';
const myProvider: LlmProvider = {
isConfigured: () => true,
async anonymize(text) {
/* return { anon, mapping, legend } */
},
async anonymizeBatch(texts, usedIds) {
/* return { segments, mapping, legend } */
},
};
Placeholder format
[PER_1.NOM:M] entity PER #1, attribute NOM, context M (rich, from the LLM)
[EMAIL_1] structured id from the regex fallback
PLACEHOLDER_RE is exported if you need to scan text for placeholders. The LLM may also coin new
abbreviations (always uppercase [A-Z_]) for entities/attributes/context it discovers — every one it
uses is described in the result legend.
How it works
The library does query-time pseudonymization: it rewrites text so that personal data never leaves your trust boundary in identifiable form, while keeping the answer fully reversible on your side.
┌───────────────────────────── your trust boundary ──────────────────────────────┐
raw text ─▶ ① pre-filter ─▶ ② LLM detect ─▶ ③ regex fallback ─▶ ④ validate ─▶ anon + mapping + legend
(skip if (names, (optional; (every │ │
clearly addresses, structured ids; placeholder ┌───────┘ │
PII-free) coins abbrevs, fail closed if is mapped) ▼ ▼
builds legend) absent) mapping legend
(SECRET, (shareable:
downstream LLM ◀── anon + legend ───────────────────────────────────────────── reversible) PER=Personne…)
answer (with [PER_1.NOM:M] tokens)
│
▼
⑤ stream de-anonymize ──▶ real values restored for the end user (placeholders never leak, even if split)
- Pre-filter — a cheap regex/name check skips the LLM round-trip for text that clearly has no PII. (Bypassed when no regex fallback is configured, so nothing slips through.)
- LLM detection — finds free-form PII a regex can't (names, addresses), keeps coreference (the
same person is always
[PER_1]), coins uppercase abbreviations for anything new, and returns a legend describing them. - Regex fallback — optional, deterministic detection of structured identifiers; used if the LLM is
unavailable. Omit it to fail closed (raise
AnonymizationErrorrather than risk a leak). - Validation — bidirectional check that every placeholder has a mapping entry and vice-versa.
- Streaming de-anonymization —
makeStreamDeanonymizerrestores real values token-by-token, buffering any placeholder split across chunks so a partial[PER_is never emitted.
Three distinct outputs, with different sensitivities:
| Output | Example | Sensitivity |
|---|---|---|
anon |
Le dossier de [PER_1.NOM:M] |
Safe to send onward — no identifiers |
mapping |
{ "[PER_1.NOM:M]": "Alain Jaccard" } |
Secret — re-identifies people; keep it on your side |
legend |
{ "PER": "Personne", "M": "Masculin" } |
Safe to share — explains the tokens to a downstream LLM |
How this helps with the nLPD
Switzerland's nLPD (and the GDPR) push for data minimisation and favour pseudonymisation when personal data is processed by third parties. This library is built around those principles:
- The third party never sees raw PII. When you send text to an external LLM (or any external service),
it receives only pseudonymised tokens like
[PER_1.NOM:M]plus the non-identifyinglegend— never the real name, e-mail, AVS number, etc. - Pseudonymisation, not loss of meaning. The re-identification key (
mapping) stays in your infrastructure; only you can reverse the tokens. Thelegendlets the downstream model still reason correctly ("a person", "male") without knowing who. - Fail-closed option. Omitting the regex fallback means that if detection can't run, the call errors instead of forwarding data that wasn't pseudonymised — no silent leak.
- Coreference & minimisation. Re-using one id per entity avoids spreading extra distinguishing detail across a prompt, and the corpus itself can stay in clear text — pseudonymisation happens only at the boundary, at query time.
This is an engineering aid, not legal advice or a certification. You remain the data controller; assess it against your own obligations (see the disclaimer below).
Compliance note
This library is a best-effort pseudonymization aid, not a guarantee of legal compliance. LLM and regex detection can miss or mis-classify data. Validate against your own requirements (nLPD, GDPR, HIPAA, …) before relying on it for regulated data.
License
MIT © Mobiletic