You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit bc92226
Browse filesBrowse the repository at this point in the historyBrowse files
|**Feature / Initiative**|[UIESTRAT-229: Enable collection & upload of Observability OTel data stripped of PII/sensitive data](https://redhat.atlassian.net/browse/UIESTRAT-229)|
384
+
385
+
The three core inference spans (`/v1/query`, `/v1/streaming_query`, `/v1/responses`) carry
386
+
raw, un-hashed request/response content for evaluation, relaxing the original metadata-only rule
387
+
(**R7**; §Why "Safe observability by design"). This addendum covers how that content is
388
+
redacted of PII.
389
+
390
+
### The problem
391
+
392
+
Raw span text (prompts, responses, RAG chunks, tool I/O) can contain PII, so it must be scrubbed
393
+
before the span leaves LCORE — before OTLP export to a hosted backend such as LangFuse, and before
394
+
any cross-org sharing. This is a portfolio-wide obligation from Red Hat's AI Assessment process,
395
+
not specific to LCORE.
396
+
397
+
### The approach — Presidio from PyPI
398
+
399
+
Use the upstream [Presidio](https://github.com/microsoft/presidio) library directly. Add
400
+
[`presidio-analyzer`](https://pypi.org/project/presidio-analyzer/) and
The model version must match the resolved spaCy version, so they need to be pinned together.
436
+
437
+
At runtime, loading the model costs a few hundred MB of RSS, so the `AnalyzerEngine` and
438
+
`AnonymizerEngine` are constructed once (module-level singleton or lazy init) and reused across
439
+
requests — `redact()` never rebuilds them per call.
440
+
441
+
The `sm`/`md`/`lg` choice is the accuracy-vs-footprint knob, and it moves three things at once: image
442
+
size, pod memory, and NER recall on names, organizations, and locations. `lg` gives the best recall
443
+
at the largest cost; `sm` is small but weaker.
444
+
445
+
### Alternatives considered
446
+
447
+
-**Reuse the existing regex redaction engine** ([`redact_text`](../../../src/pydantic_ai_lightspeed/capabilities/redaction/core.py#L24) + the configurable [`RedactionConfig`](../../../src/models/config.py#L2912)
448
+
rules, already in the repo) at span emission. This is the engine under the `PiiRedactionCapability`
449
+
shield, not the shield itself — the shield is used to redact PII during a2a live message flow. Only the
450
+
"text in, redacted text out" engine would be used from here. No new dependency and no model
451
+
download, deterministic, and fast. However, regex handles structured PII well (email, IP, phone are
452
+
easy to express) but cannot reliably detect free-text entities like names, organizations, and
453
+
locations, so recall on those depends on hand-written patterns and is weak.
454
+
455
+
Comparatively, the existing data-anonymizer based on Presidio is already [compliance approved and in production](https://docs.google.com/document/d/1BNBIDUz-lLmUjT9N9cEMqgaQv9EqfLKwZrA8gtbtlY8/edit?tab=t.0).
456
+
457
+
-**Managed cloud PII services** (AWS Comprehend, Google Cloud DLP, Azure AI Language). Ruled out:
458
+
each sends span content to an external service on every request, which reintroduces the
459
+
PII-leaves-the-process problem this work exists to close, adds a network hop, and adds a cloud
460
+
dependency the LCORE binary can't assume.
461
+
462
+
Presidio wins on two counts: good name/org/location recall without building an NER pipeline by
463
+
hand, and alignment with the shared `data-anonymizer` library (also Presidio-based; see § Future),
464
+
so the later swap stays a drop-in rather than an engine change. The one cost is the spaCy model
465
+
download. If that model turns out to block image size or air-gapped builds, the regex-extend option
466
+
above is the fallback.
467
+
468
+
### Where redaction runs — redact at emission, through one helper
469
+
470
+
Content is redacted at the point it's set on the span. All content fields will go on spans
471
+
through a **single centralized helper** — e.g. `set_content_attribute(span, key, text)` — that runs
472
+
each value through the Presidio `redact()` helper before calling `set_attribute`. For any content
473
+
serialized as a JSON string (e.g. RAG chunks or tool I/O), the helper redacts each leaf string
474
+
(`content`/`args`/`source`) before `json.dumps`, so the serialized attribute never contains raw
475
+
PII. Properties of this approach:
476
+
477
+
-**The span never holds raw PII.** R11 ("redact before content leaves the process") is satisfied
478
+
by construction — there is no in-memory window where a raw value sits on a span or in the export
479
+
queue.
480
+
-**One code path to audit.** Because content can only reach a span through the helper, there is a
481
+
single place to review and test; individual emission sites cannot forget to redact (there is no
482
+
raw path to `set_attribute` for content keys). A review/lint rule reinforces "content goes
483
+
through the helper."
484
+
-**Fail-closed.** If the redactor raises, the helper omits the field rather than emitting raw
485
+
text — consistent with R9 (a tracing failure must never leak or break the request).
486
+
-**Cost.** Redaction runs on the request path (synchronous), not in the export background thread.
487
+
This is minor. It's a text pass over already-generated output, dwarfed by the LLM call, and only
488
+
incurred when raw-content capture is enabled (scoped to the three core spans). The helper is the
489
+
single place to make it async/best-effort if it ever becomes a bottleneck.
490
+
491
+
### Future: adopting the shared `data-anonymizer` library
492
+
493
+
A shared, Presidio-based redactor already exists: `data-anonymizer` (used by Ask Red Hat et al.;
494
+
see the [shared Presidio library strategy doc](https://docs.google.com/document/d/1BNBIDUz-lLmUjT9N9cEMqgaQv9EqfLKwZrA8gtbtlY8/edit?tab=t.0#heading=h.9qfphopokr7u)).
495
+
LCORE may move to it eventually, so redaction is maintained once across teams rather than
496
+
per-project. The only obstacle is distribution. `data-anonymizer` lives on Red Hat's internal GitLab
497
+
(`gitlab.cee.redhat.com/uxe-data-ai-solutions/data-anonymizer`), not PyPI, and a GitHub build/CI
498
+
can't install from internal GitLab without internal credentials — so it can't go in
499
+
`pyproject.toml`.
500
+
501
+
The swap stays small because redaction is already funneled through the single `redact()` helper
502
+
(§ Where redaction runs). To adopt the shared library without breaking the GitHub build, turn that
503
+
helper into a **redaction slot** — one interface the telemetry path calls without knowing the
504
+
backend:
505
+
506
+
-**GitHub build (default):** the slot keeps using public Presidio from PyPI.
507
+
-**Red Hat's internal build:** a downstream build fills the slot with `data-anonymizer` via an
508
+
adapter, added at the stage where internal GitLab is reachable.
509
+
510
+
Until such a move is deemed required, the PyPI Presidio helper above is the whole
511
+
implementation.
512
+
513
+
### Requirements addendum
514
+
515
+
-**R11 — PII redaction before export.** Raw content on the three core inference spans shall pass
516
+
through PII detection/redaction before leaving the process (OTLP export or cross-org sharing).
0 commit comments