OpenTelemetry tracing can connect observations across an application’s services and show where a request spends time. That evidence is useful for diagnosis, but traces can also collect identifiers, URLs, arguments, and other sensitive information. Instrumentation should answer operational questions without becoming an uncontrolled copy of business data.
This guide explains a privacy-aware tracing workflow for systems you operate. It focuses on what is recorded, where it flows, and how collection is verified. A trace identifier or sampling setting does not automatically make the entire telemetry pipeline safe.
Define the OpenTelemetry tracing purpose
Identify the failures and performance questions the team needs to investigate. A request path, dependency outcome, and duration can be useful without recording the full request body. Choose attributes that support those questions.
Classify values that must not enter ordinary telemetry, such as credentials, session tokens, private documents, or unnecessary personal data. Include values generated by libraries and automatic instrumentation, not only manually created spans.
Assign ownership for the instrumentation contract. A central collector cannot reliably correct every unsuitable attribute if developers do not know what should be emitted.
Review the complete collection path
Map application instrumentation, exporters, collectors, processors, storage, and user interfaces. Different components can retain or duplicate data. A redaction step in one path may not cover another exporter.
Check default instrumentation behavior and debug modes. URLs or exception details can contain private values even when the main span name looks harmless. Review the actual emitted records.
The OpenTelemetry sensitive-data guidance describes collection and processing considerations. Use the supported tools and version-specific configuration for your deployment.
Prefer safe attributes at the source
Use stable operation names and bounded outcome categories. Avoid placing arbitrary user input in span names or high-cardinality labels. Uncontrolled values can expose information and make telemetry difficult or costly to use.
Record a safe resource reference only when it is needed and approved. An identifier that seems anonymous can still link to a person or reveal internal structure. Treat context as part of the privacy review.
Do not record authorization headers or entire environment mappings for convenience. Their diagnostic value rarely justifies routine credential exposure.
Use processing as defense in depth
Supported collector processors can remove or transform selected attributes. Configure them deliberately and verify ordering and coverage. A processor’s presence is not proof that every sensitive value is removed from every field.
Use known-field policies and tests rather than claiming one regular expression recognizes all secrets. Values can be nested, encoded, or embedded in exception messages. Avoid collecting unnecessary material in the first place.
Restrict who can change the collection pipeline. A configuration update can broaden disclosure even when application code remains unchanged.
Choose sampling for the operational requirement
Sampling can control volume, but it does not make sampled records non-sensitive. Every retained trace still needs approved handling. Review which failures and request classes may be missed by the selected strategy.
Choose supported sampling behavior that fits diagnosis and cost requirements. Do not assume a low percentage retains every important incident, or that a large sample automatically provides representative coverage of rare events.
Keep sampling configuration with the interpretation of metrics and investigations. A missing trace may reflect collection policy rather than proof that no request occurred.
Protect exports, storage, and access
Use approved transport, credentials, and destinations. Telemetry platforms need access controls and retention appropriate to the data they receive. A private application can still leak information through an overexposed observability account.
Review administrators, support access, shared dashboards, and exported files. Obscure links are not a dependable permission boundary. Avoid placing complete traces in a public issue or unrelated external analyzer.
Keep retention purposeful. Data useful for a short diagnostic window does not necessarily need indefinite storage.
Test actual emitted telemetry
Use synthetic canary values and representative success, failure, timeout, and exception paths. Inspect every configured destination to confirm prohibited values do not appear. A unit test of one processor cannot establish the deployed pipeline’s full behavior.
Check automatic instrumentation and dependency calls too. A library update can add attributes the application never explicitly set. Re-run the review after meaningful instrumentation or collector changes.
Keep safe diagnostic evidence of the test configuration and results. Do not introduce real credentials into a telemetry test merely to make the example convincing.
A practical tracing review
Consider a test endpoint that calls a controlled downstream service and fails once with a synthetic error. Inspect the span chain, timing, status, and attributes. The trace should explain the dependency failure without exposing the test token or entire submitted record.
Repeat the exercise through every approved exporter and under the deployed sampling policy. Check whether a processor removes the relevant field from stored and exported evidence, not only from one visible dashboard. Document any paths intentionally excluded from collection.
Our journal retention guide explains a complementary logging concern. Logs and traces can correlate useful events, but each path needs its own collection, privacy, and retention contract. Keep the review tied to the actual operational question rather than treating maximum recorded detail as maximum observability.
Review telemetry at release boundaries
Instrumentation changes can introduce new attributes without changing the visible business interface. Treat the emitted data contract as part of a release, especially when a dependency enables additional automatic spans or an exporter changes its output format.
- Compare representative records before and after an instrumentation update. Check names, attributes, events, errors, and resource metadata rather than reviewing only the span’s most obvious fields.
- Verify that safe correlation identifiers remain useful without exposing credentials. A trace should support diagnosis through approved references, not require copying complete confidential request bodies.
- Review sampling and retention together. Rare failures may need an approved evidence strategy, but recording more data does not remove the obligation to minimize sensitive material.
- Confirm collector configuration is applied to every intended path. A direct application exporter can bypass protections that work correctly on another route through the central pipeline.
- Test access to stored traces and exported files under representative roles. A protected ingestion endpoint does not prove that dashboards and support downloads have suitable permissions.
Keep a rollback and incident route for unintended collection. If a secret is emitted, follow the credential provider’s response process; deleting the visible trace does not reliably invalidate copies or the credential’s remaining authority.
Frequently asked questions
Does sampling prevent sensitive information from leaking?
No. It changes collection volume. Retained records still need safe fields, access controls, and approved handling.
Is a collector redaction rule enough?
Not by itself. Coverage depends on emitted data, processor behavior, and all destinations. Test the full pipeline and minimize data at the source.
What should I record first?
Stable operation names, safe outcome categories, timing, and correlation context needed for the task. Add other fields only with a clear purpose and data policy.