Better Synthetic Replacements Preserve the Context Data Needs
Redaction protects identity by destroying information. Limina's new Contextual Replacement capability takes a different approach, replacing sensitive identifiers with contextually accurate synthetic values so your data stays useful for research, collaboration and AI training.

There's a quiet tradeoff in de-identification that some teams feel long before they can name it. Sentiment analysis, topic modeling and conversation analytics work perfectly well on redacted data, because the identifiers themselves carry no meaning those tasks need. But when the transformed data has to stay readable and coherent, for research partners, human reviewers or AI training, protecting identity has often meant sacrificing meaning.
Redact aggressively and the dataset fills with black boxes. Substitute randomly and a regional bank becomes an airline, a mother and daughter stop sharing a surname, and a referring physician turns into a name that appears nowhere else in the record. The data is protected, technically. It's also, for many purposes, ruined.
This tradeoff has real consequences. Research collaborations stall because the shared dataset is missing key information for the receiving institution’s project. AI training projects run on text where placeholder markers have replaced the natural language patterns models learn from. The privacy work succeeded, but the data stopped working.
With release 4.5, Limina is introducing Contextual Replacement, a customer-ready capability designed to shrink that tradeoff. It replaces sensitive identifiers with synthetic values built to fit their surrounding context and preserve selected relationships, so protected data stays useful to the people and models that depend on it.
Contextual Replacement ships in release 4.5, targeted for September 14, 2026, and is included for all customers on that release.
When redaction and random substitution fall short
Redaction is still a great tool for a lot of privacy work. It's fast, it's proven and for tasks like sentiment analysis, topic modeling and conversation analytics it does exactly what's needed, because those tasks don't depend on the identifiers at all. What redaction can't do is preserve the demographic and relational detail that research depends on, or supply the natural language patterns an AI model needs to learn from. For those specific workflows, the information being gone is the problem.
Generic substitution goes a step further by swapping identifiers for placeholder values. This keeps the text flowing, but the replacements carry no awareness of context. The substitute for a hospital might be a hardware store. Two mentions of the same person might become two different people. A family in the source data dissolves into unrelated strangers. The text reads, but it no longer holds the original meaning or value.
Contextual replacement is the third approach, and it's the one Limina has built. Instead of removing information or substituting it blindly, the system generates replacement values designed to fit the context around them. A hospital is replaced by a plausible hospital. A referring physician remains a physician. Related family names can remain consistent after replacement. An organization such as a bank can be replaced with another contextually appropriate financial institution rather than an unrelated company. Where a research team intentionally retains a known population characteristic, replacement is designed to avoid unnecessarily erasing that signal.
This kind of synthetic replacement also carries a privacy property that redaction can't offer, one researchers describe as hiding in plain sight. In a redacted document, anything that slips past detection is immediately recognizable as real. In a document filled with realistic synthetic values, a residual real value is indistinguishable from its synthetic neighbors, so it no longer announces itself.
What this looks like in practice
Every identifier has been replaced. The patient's name, the hospital, the city, the date, the phone number and the referring physician are all synthetic. But the record still reads like a record. The hospital is still recognizably a hospital. The physician is still a physician. The phone number has valid structure. The date is plausible and can even be shifted by the same amount as other dates in the dataset. A researcher, a reviewer or a training pipeline can work with this text the way they'd work with the original, without the original's disclosure risk.
How the hybrid system works
Contextual Replacement isn't a single model doing everything. It's a hybrid system that combines two generation approaches and falls back automatically when neither applies, so replacement coverage stays complete across a request.
Any detected entity type outside the supported lists falls back automatically to Limina's existing synthetic generation system within the same request. Nothing is skipped, and customers don't have to choose between two competing products. The existing replacement system remains available and remains the default; Contextual Replacement is an option teams enable when their workflow calls for stronger contextual utility.
What this unlocks for research and AI training
The teams we expect to benefit first are the ones whose best data is currently their least usable data.
Research collaborations are the clearest case. A dataset that has to move between teams, institutions or external partners needs to survive the trip with its meaning intact. Contextual Replacement lets organizations replace original identifiers before data moves into an approved research or partner environment, while keeping the dataset coherent enough for the receiving side to work with. The privacy team gets a defensible, auditable transformation; the research team gets data that still makes sense.
AI training is the second. Models learn from patterns in language, and training on heavily redacted or randomly substituted text means learning from distorted patterns. Contextually accurate replacements keep the linguistic structure of the source data intact, producing higher-quality training material without feeding real identifiers into the pipeline.
Both use cases sit squarely in the workflows healthcare and pharma and life sciences organizations run every day, and both build on the same data de-identification platform Limina customers already deploy.
Getting started
If your organization is on release 4.5 with the GPU Synthetic Enhanced image, enabling Contextual Replacement is a configuration change away. If you're not yet a Limina customer and you have a research, collaboration or AI-training workflow stuck behind a privacy review, this is a good moment to talk.
Request a demo and bring a workflow. We'd rather show you what contextually accurate replacement does with your kind of data than tell you about it.
Test It on Your Own Data.
Free API key, no credit card. Enough calls to see what your current tool is missing.


