anonymizeTurn(text, session) threads a serializable AnonymizerSession
({mapping, legend, history}) so one entity keeps one id across a whole chat
(applyKnown reuse + usedIds + de-collision merge). Adds conversation()
in-memory wrapper and an optional LlmProvider.anonymizeInConversation(text, ctx)
for rich cross-turn context; providers without it fall back to the batch path.
openAICompatibleProvider gains includeMappingInContext (default false — only
send real values to a trusted anonymizer endpoint) + historyMaxTurns (default
10). Verified live vs Gemma 4: Nora stays PER_1 across 4 turns (name/email/AVS/
IBAN), Yanis = PER_2; 41 tests, 98% coverage.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
5.3 KiB
5.3 KiB
Changelog
All notable changes to this project are documented here. The format is based on Keep a Changelog, and this project adheres to Semantic Versioning.
[0.4.0] - Unreleased
Added
- Multi-turn conversation support. New
Anonymizer.anonymizeTurn(text, session?) → { anon, mapping, legend, session }keeps one stable id per entity across a whole chat: it seeds each turn with a running, serializableAnonymizerSession({ mapping, legend, history }), reuses known values viaapplyKnown, tells the model which ids are taken, and de-collides new ones. Persist the returnedsessionand pass it back next turn. Anonymizer.conversation(initial?)— a stateful in-memory wrapper (anonymize,deanonymize,session()) overanonymizeTurn.- Optional
LlmProvider.anonymizeInConversation(text, ctx)— providers can use prior context (anonymizedhistory,legend,usedIds, and optionallymapping) for better cross-turn coreference/attribution.openAICompatibleProviderimplements it; providers that don't fall back to the batch path automatically. openAICompatibleProvideroptions:includeMappingInContext(default false — only send real values to a trusted anonymizer endpoint) andhistoryMaxTurnsviaAnonymizerConfig(default 10).- Exported the
AnonymizerSessiontype.
[0.3.1] - Unreleased
Changed
- Default prompt hardening for lossless round-trips: (1) two different values of the same type never
share a placeholder — they must be distinguished by context (e.g.
[PER_1.IBAN:Ancien]vs[PER_1.IBAN:Nouveau]), preventing a within-message collision (old/new IBAN); (2) the model must not absorb adjacent punctuation/separators (commas, spaces, parentheses) into a placeholder. Together these took the live Gemma-4 e-learning demo from 4/5 to 5/5 exact round-trips.
[0.3.0] - Unreleased
Changed
openAICompatibleProvidernow works with Infomaniak (and other open-model endpoints) out of the box.response_formatis omitted by default instead of hard-coding{ type: 'json_object' }, which current Infomaniak rejects (HTTP 422). Pass the newresponseFormatoption (e.g.{ type: 'json_object' }or ajson_schemaobject) for endpoints that support/require it. Breaking for endpoints that relied on the previous forcedjson_object.- JSON responses are now parsed leniently — a fenced JSON code block or surrounding prose is tolerated
(the outermost
{ … }is extracted), so models without an enforcedresponse_formatdon't cause spurious parse failures. - The default prompt gained a COHÉRENCE block (exact placeholder↔mapping-key identity, values are the
original data never another placeholder, strict
[TYPE_N…]format, mask the value not the adjacent label) to improve reliability across models.
[0.2.0] - Unreleased
Added
legendon every result — abbreviation → French meaning (PER→Personne,M→Masculin), safe to forward to a downstream LLM so it understands the placeholder tokens. Backed by a built-inDEFAULT_LEGENDso coverage is guaranteed even if the model omits entries.- The default prompt now lets the model coin new uppercase abbreviations for entities/attributes/
context it discovers and return their meanings in
legende. PatternDef.meaning— optional human label for a tag, surfaced in thelegend.presets.swiss/presets.genericship French meanings.AnonymizationError(exported) — thrown when anonymization can't complete and no fallback exists; carries the originating error in.cause.
Changed (breaking)
- Regex fallback is now opt-in.
patternsno longer defaults topresets.swiss. With no fallback, an LLM failure throwsAnonymizationError(fail-closed) and the pre-filter is bypassed. At least one ofllmorpatternsis required, or the constructor throws. AnonymizationResultgained a requiredlegendfield;LlmProvider.anonymizeBatchreturnslegend.anonymizeChunks(chunks, seed)—seedis now{ mapping, legend? }(was the bare mapping) and the return includeslegend.
[0.1.0] - Unreleased
Added
- Initial public release.
Anonymizer— pre-filter → LLM → regex fallback, bidirectional validation, deterministic coreference, and de-collision across question and retrieved chunks (anonymize,anonymizeChunks,deanonymize).makeStreamDeanonymizer— streaming-safe de-anonymization that never leaks a split placeholder.openAICompatibleProvider— pluggable LLM detection over any OpenAI-compatible Chat Completions API.presets.swissandpresets.genericregex pattern sets; fully configurable custom patterns.- Regex-only mode (no LLM provider required).
PatternDef.validate— optional second-stage predicate to cut false positives;presets.genericuses it for a Luhn check on credit-card candidates, and de-overlaps its phone/date/IP patterns.openAICompatibleProviderretries transient failures (network/timeout/429/5xx) viaretriesandretryDelayMsoptions; non-transient 4xx and malformed responses are not retried.- Hardened the
nameHintheuristic:g/yflags are stripped internally so.test()is stateless.