The five shapes of synthetic identity in India
Synthetic identity is not one document trick. Five operating patterns explain why valid records can still resolve to the wrong economic actor.
The category is broader than document fraud
A synthetic application can contain authentic records. Aadhaar authentication, for example, answers a defined authentication question about submitted identity data. The Unique Identification Authority of India describes the process as verification against information held in the Central Identities Data Repository. A successful response does not decide who controls a loan application, who will receive the funds or who will operate the resulting account.
The RBI KYC Master Direction sets duties for customer identification, ongoing due diligence and record management. Those duties are necessary. They still leave a separate risk question: do the verified attributes resolve to one independent person whose economic behaviour matches the application?
That question changes the unit of analysis. A document control inspects a credential. A personhood control compares credentials with account history, device evidence, relationship reuse and the timing of related applications. The five shapes below are operational categories for that comparison. They are threat models, not claims about prevalence.
1. Rented identity
A rented identity uses records that belong to a real person while another party controls the application. The named person may have handed over credentials, may be under pressure or may understand only part of the transaction. The document set can remain internally consistent because it came from one genuine record holder.
The break appears outside the document. The application may arrive from a device used for unrelated identities. A beneficiary account may recur across borrowers. Contact details can change immediately after approval. None of these observations proves rental alone, but together they test control rather than document validity.
The right response is not an automatic decline based on one shared attribute. It is a ranked reason code and a step-up path that asks the institution to confirm control through evidence independent of the submitted document bundle.
2. Hollow real identity
A hollow identity belongs to a real person and contains authentic credentials, but it has little sustained economic activity around it. The pattern matters because a genuine first-time borrower can also have a thin file. Thinness is not fraud, and absence of bureau history cannot carry a decline by itself.
The distinction is coherence. A genuine thin-file applicant can still show stable control of a mobile number, a plausible address history and a device that is not shared across a recent application ring. A hollow identity may show very little evidence while also connecting to infrastructure used elsewhere. The missing footprint and the repeated infrastructure must be evaluated together.
This is why a useful model needs an insufficient-evidence state. When the footprint is absent but the available signals agree, review is safer than decline. When absence is joined by contradictions or ring links, the review queue should move accordingly.
3. Recombinant identity
A recombinant identity joins valid attributes that do not belong to one person. A name may align with one credential while contact control, address history or a payment destination aligns with another. Each field can pass its own check because each underlying value exists.
The control failure is caused by independent verification paths. If every vendor returns only a field-level pass, the institution receives several positive answers without a test of whether the answers describe the same actor. Cross-field coherence must therefore be a first-class output, not an analyst note added after loss.
A reason code should identify the contradiction precisely. Generic labels such as identity mismatch are too weak for a regulated decision. The reviewer needs to know which attributes disagree, when each source was observed and whether the difference can be explained by an ordinary life event.
4. Seasoned ghost
A seasoned ghost is built to acquire history. Small facilities or low-risk accounts create a record that can later support a larger application. The earlier clean outcome is part of the construction, so an application-time score can look stronger at the second encounter.
The useful signal is change through time. The original device may disappear. Contact details can rotate. New applications may connect the identity to a shared beneficiary or submission cluster. A system that preserves the first decision as a baseline can rank those changes without treating all account updates as suspicious.
The RBI KYC Master Direction already defines periodic KYC duties by risk category. Scheduled personhood re-scoring can inform those existing processes, but it does not replace the institution's legal review or customer communication duties.
5. Farmed cluster
A farmed cluster is visible at graph level. One operator or coordinated group reuses devices, addresses, face vectors, beneficiary accounts or submission windows across many applications. A single record can look ordinary because the repeated infrastructure sits in other records.
The Reserve Bank Innovation Hub describes MuleHunter.AI as a system that analyses account and transaction behaviour to identify suspected mule accounts. The Ministry of Finance reports that the system is operating across 26 banks. That work concerns activity after onboarding, but the graph lesson also applies before account opening: relationships change what an isolated application means.
Common attributes need careful weighting. A pin code or family address can connect many legitimate applicants. A rare device fingerprint joined to one payout destination inside a short submission window is different evidence. The graph should preserve typed relationships so an analyst can see why a ring formed.
The shape is a hypothesis, not a verdict
An archetype helps an investigator decide what evidence to seek. It should not become a label attached to a person without review. The same shared address can describe a fraud ring, a student residence, a workplace or a family home. The same thin footprint can describe a fabricated record or a genuine applicant entering formal credit for the first time.
Evidence therefore needs provenance and time. A device link should carry the first and last observation, the applications it connects and the rule that produced the fingerprint. An address relationship should preserve how the value was normalised. A face-vector match should carry a quality state so poor capture is not treated as a confident identity link. The review record should also separate a missing observation from a negative result. A signal that was never collected cannot support the same inference as one that was collected and found to be clean.
The institution also needs a counterfactual review. Analysts should inspect clean cases that the model ranks highly and confirmed cases that it ranks poorly. Those samples reveal branch workflows, data-quality gaps and population differences that aggregate performance can hide.
A taxonomy earns its place only when it changes action. If every shape produces the same score threshold and the same decline path, the categories are decoration. Each one should point to a distinct evidence request, review owner or monitoring window.
A taxonomy should change the control
These five shapes should not collapse into one fraud score. They imply different evidence, review paths and time horizons. Rented and recombinant applications call for control checks. Hollow identities call for careful treatment of missing evidence. Seasoned ghosts require comparison over time. Farmed clusters require graph review.
The opinion behind this taxonomy is simple: document validity and human control are separate decisions. Institutions should retain the first, add the second and keep the reason codes visible to the person accountable for the outcome.
Preserve field-level results
Keep the original KYC, bureau and authentication responses. A personhood layer should not hide or rewrite them.
Test coherence separately
Return contradictions and missing evidence as named reason codes with source times.
Carry graph context
Show which shared infrastructure changed an application's risk and how strong each relationship is.
Match action to evidence
Use approve, step-up, review and decline paths that reflect evidence quality rather than one universal threshold.
