Identities
Every person and organization across the corpus, each with one consistent token.
PERSON_0412 · ORG_0091
Identity registry for AI training data
Every person in your corpus shows up a dozen ways: a name in a contract, a nickname in Slack, an email in the CRM. A7 Labs resolves them into one identity before de-identification, so the context survives all the way to the model.
The registry
After de-identification, each name is its own PERSON token, and his email, phone and SSN are typed tokens linked to no one.One canonical identity. Names resolve to it; typed identifiers stay typed and link to it. 12 identifiers, 8,258 mentions across 149 documents.
How training data is made today
Email, Slack and Teams, contracts, CRM records, decks: terabytes of unstructured data from one enterprise, bought to become training data.
Tommy in a Slack thread, Thomas Erickson in email, Mr. Erickson in an external email, his phone number in a deck, his SSN in a CRM document. De-identification turns the names into three different PERSON tokens and leaves his EMAIL, PHONE and SSN tokens linked to no one.
The record that says all six are Thomas has no owner: not the seller, not the data lab, not the de-identification vendor. Every link between Thomas and his decisions, actions and outcomes is lost.
Right after acquisition and before de-identification, we build the registry record: one identity, every alias and identifier confirmed, one token. PERSON_0412 is Thomas everywhere he appears.
Not just Thomas. Every person, organization and deal in the corpus gets one identity and keeps its relationships. More signal per token, less hallucination.
Where we sit
Everyone is racing to own dataset preparation end to end. Nobody owns the identity registry, and each player assumes someone else does.
Empty.
Nothing passes through here today. The corpus goes straight from the data lab to de-identification, so nothing links Tommy to Thomas.One identity per person, before anyone tokenizes anything.
Who builds the identity registry?
“The data lab will sort out who's who.”
“The de-identification vendor handles identities.”
“The data lab gives that to us.”
Everyone does a piece of it. Nobody does all of it. Every mention becomes a stranger, and the links are lost.
“We build it.” The only company working 100% on the identity registry for unstructured data.
What the registry resolves
Every person and organization across the corpus, each with one consistent token.
PERSON_0412 · ORG_0091
Names, nicknames, titles and handles, confirmed and mapped back to the right person.
"Tommy" → PERSON_0412
Emails, phone numbers, SSNs, employee IDs, roles and employers stay attached to one identity through de-identification.
role: "CFO" · treatment: Tokenize
Who signed, approved, reported to and worked with whom, preserved across documents and systems.
signed → DEAL_0210
About
We built the AI privacy product at a leading data privacy infrastructure company from the first line of code and ran it for more than three years. Our customers were the data providers that supply the frontier labs, AI-native healthcare companies, global retailers and the labs themselves.
That is where we saw the gap. Every one of them acknowledges it. None of them owns it.
The only company working 100% on the identity registry for unstructured data. We waited for the right problem. This is it.
Work with us
We are currently working with a small number of selected enterprise partners and data labs. Tell us about your data and we'll run a sample through the registry.