A fuzzy name duplicate is a vendor entered twice under name variants, such as a typo or a punctuation difference, rather than under a shared tax identifier. Because identifier-based checks compare PAN and GSTIN fields, not vendor names, a diagnostic that validates only those fields misses this pattern. Across two vendor master diagnostics, this check flagged 54 and 48 records respectively, most with no identifier overlap at all.
One vendor record carried a company name. A second record for the same entity carried the same name with one letter added, entered by someone re-keying the name instead of searching for the existing record. Neither record shared a GSTIN with the other. Neither shared a PAN. A duplicate PAN and GSTIN check, run correctly, would have cleared both records without a flag.
When we ran VCR-012 (fuzzy name match) across two vendor master diagnostics, a packaging manufacturer returned 54 flagged records and an auto dealership group returned 48 across 10 clusters.
How Fuzzy Name Duplicates Differ from PAN and GSTIN Duplicates
A duplicate GSTIN check compares one identifier field. A duplicate PAN plus GSTIN check compares two identifier fields together, see Duplicate PAN and GSTIN in Your Vendor Master: What the Check Finds. Both approaches assume the identifiers themselves were entered correctly, or entered at all. A fuzzy name duplicate does not make that assumption. It compares the vendor name field directly, using similarity scoring rather than exact match, to catch the case where the PAN field is blank on one record, the GSTIN field is blank on the other, or both identifiers are different because the second entry came from a different intake process.
Fuzzy vendor name matching in an Indian ERP vendor master catches the same vendor appearing under multiple names, not multiple identifiers. The added-letter pair above is the clean version of this pattern, same entity, one character different in the name, no identifier overlap to catch it. Vendor creation at the branch or plant level, without a check against a shared master, produces this pattern regularly. A typo during manual entry, a punctuation difference between “Pvt Ltd” and “Private Limited,” or a name entered in a different transliteration all pass the same way past a check that only compares tax fields.
What the Diagnostic Data Shows
| Client | Vendors flagged | Notes |
|---|---|---|
| A packaging manufacturer | 54 | 18 of the 54 pairs resolve to the same PAN under different GSTINs, a multi-state registration pattern. The remainder have fully distinct PAN and GSTIN. |
| An auto dealership group | 48 (10 clusters) | One pair shared an identical GSTIN despite the name difference; this pair was also independently flagged by the same client’s duplicate GSTIN check |
The auto dealership group’s shared-GSTIN pair did not need the name check alone. The client’s own duplicate GSTIN check had already flagged it. Where a fuzzy name duplicate shares an identifier, an identifier-based check catches it independently. The packaging manufacturer’s data shows the more common case. Of 54 flagged pairs, 18 share a PAN across two distinct, validly registered GSTINs. The remaining 36 share no identifier at all, records an identifier check has no field to compare against either way.
Why This Check Runs Alongside Identifier Checks, Not Instead of Them
A name duplicate with distinct GSTINs does not create the GST reconciliation exposure documented for identifier-based duplicates. GSTR-2B ITC matching runs strictly at the GSTIN level, so two genuinely distinct GSTINs reconcile independently. A non-filing or discrepancy against one has no bearing on ITC eligibility under the other, regardless of whether the two vendor records share a name or a parent PAN.
The exposure that does carry over is narrower and specific to TDS. Sections 194Q and 206C(1H) track cumulative transaction value per PAN, not per GSTIN or per vendor code, as typically interpreted. Of the packaging manufacturer’s 54 flagged pairs, 18 resolve to the same underlying PAN under different, state-specific GSTINs. An ERP that calculates the ₹50 lakh threshold per vendor code rather than per PAN will not aggregate those 18 pairs’ payment history together. Each record can sit under the threshold individually while the combined value to that PAN has crossed it, producing a genuine under-deduction under Section 194Q or under-collection under Section 206C(1H) that neither vendor record shows in isolation. The remaining 36 pairs, with no shared PAN or GSTIN, carry no GST or TDS exposure at all, only the payment and reconciliation risk described below.
Every flagged pair carries the same payment risk, regardless of identifier overlap. Two vendor IDs for the same entity mean two separate payment histories inside the ERP. An invoice matched and paid against one record does not appear in the other record’s aging or spend history. AP staff working from the vendor code, not the name, have no structural signal that a second, older or overdue balance exists under a near-identical name one field over.
A vendor master diagnostic that runs identifier-based and name-based checks together does not let a duplicate slip past just because it cleared the one field a single check happened to compare. Learn what a Diagnostic finds.
Key observations
- Fuzzy name duplicates use similarity scoring on the vendor name field, catching entity duplication that identifier-only checks structurally cannot detect when identifiers differ or are incomplete.
- Across two vendor master diagnostics, VCR-012 flagged 54 records at a packaging manufacturer and 48 at an auto dealership group.
- Of the packaging manufacturer’s 54 flagged pairs, 18 share a PAN under different, state-specific GSTINs, a pattern that risks 194Q/206C(1H) threshold under-aggregation, not GST reconciliation exposure, since GSTR-2B matching is GSTIN-level and unaffected by a shared PAN.
- One pair at the auto dealership group shared an identical GSTIN. That client’s duplicate GSTIN check had already caught it, independent of the name check.
- Every flagged pair carries the same underlying AP risk regardless of identifier overlap. Two vendor IDs mean two separate payment histories, with no structural signal to staff working from the vendor code that a second balance exists under a near-identical name.