IND imports ▲ 4.2%USA coffee 0901 ▲ 11.8%VNM exports ▲ 6.1%BRA 0901.11 ▲ 9.4%DEU machinery ▲ 2.7%Last refresh: 2026-08-01

Company names in trade data

The same business appears under a dozen spellings across declarations. Entity resolution is the unglamorous work that decides whether an analysis is right.

A declaration records the name as typed. Over thousands of filings by different agents in different years, one company becomes many strings, and any count you make of ‘suppliers’ or ‘buyers’ is really a count of strings unless someone has done the work.

Where the variance comes from

CauseExample pattern
Legal suffix variantsPvt Ltd, Private Limited, PVT. LTD., P Ltd
AbbreviationFull name against initials against a trading style
TransliterationNon-Latin scripts romanised differently by different filers
Branch and unit namesPlant or division filing separately from the parent
Agent filingsThe clearing agent's name appearing instead of the principal
Typos and spacingSimple keying errors, at scale

What it does to an analysis

Unresolved names inflate supplier counts, deflate individual volumes, and make a concentrated market look fragmented. A conclusion that a category has forty suppliers when it has twelve is not a small error — it changes the strategy that follows from it.

Check the variants folded into a name

Before you rely on a company profile, look at which spellings were merged into it. Over-merging is as damaging as under-merging, and both are invisible in the output.

Where identifiers help

Markets that publish a stable entity identifier alongside the name are far more tractable. India’s Importer Exporter Code is one such identifier, which is why Indian company-level aggregation is more reliable than free-text matching elsewhere.

More on that in the IEC explained, and on what a resolved profile assembles in company trade profiles.

How resolution is actually done

Entity resolution is a pipeline rather than a single trick. Names are normalised — case, punctuation, whitespace and legal suffixes standardised. Candidate pairs are generated using blocking keys so that every name is not compared against every other. Candidates are scored on string similarity plus corroborating evidence: shared address, shared identifier, shared counterparties, overlapping product mix, same clearance point. Above a threshold they merge; below it they do not; in the band between, a human decides or nothing happens.

Every one of those steps embeds a judgement, and the judgements are invisible in the output. That is the real problem: a resolved profile looks identical whether it was built carefully or carelessly, and the only way to tell is to look at the variants that were folded into it.

ErrorWhat it doesHow to spot it
Under-mergingOne company appears as severalSuspiciously many small counterparties with similar names
Over-mergingSeveral companies appear as oneA single entity with implausible product breadth
Agent contaminationA clearing agent counted as the traderA handful of names dominating a whole market
Branch fragmentationUnits counted separately from the parentThe same address under different names
Transliteration driftOne name in several romanisationsNear-identical strings with consistent substitutions

What this means for the numbers you quote

Any statement of the form ‘there are N suppliers of this product in this market’ is a statement about entity resolution as much as about the market. Unresolved data inflates N and deflates the volume attributed to each, which makes a concentrated market look fragmented and contestable when it is neither. Strategies built on that misreading are internally coherent and directed at the wrong opportunity.

Ask to see the variants

Before relying on a company profile, look at which spellings were merged into it. A provider that cannot show you is asking you to trust a judgement you cannot inspect.

Checking any of this against the record

Everything above is a framework, and a framework is only worth what it survives contact with. The useful discipline is to test each assumption against what consignments actually did, because customs data is one of the few commercial sources where the underlying event — goods crossing a border — physically happened and was documented under legal obligation at the time.

Two failure modes account for most wrong conclusions drawn from trade data, and both are easy to avoid once named. The first is reading the incomplete tail of a series as a decline — authorities publish on a lag and revise afterwards, so the last one or two periods will fill in after you look. The second is reading a value movement as a demand movement, when declared value can move because volume moved, because unit price moved, or because the product mix inside a tariff line changed.

What the record cannot answer

Customs data covers goods that crossed a border. It does not cover services, domestic trade, margin, contract terms or intent. Treat it as a dated, quantified observation to corroborate — not as a conclusion that arrives finished.

Turning company names in trade data into a repeatable process

The difference between teams that get value out of trade data and teams that ran one interesting project is almost never analytical sophistication. It is whether the work became a routine. A saved query reviewed weekly, a short written note against each counterparty you assessed, and a standing habit of checking the period stamp before quoting a figure will out-perform an elaborate one-off study within a quarter, because markets move and a study does not.

The second habit worth building is writing down not just what you concluded but why and when. Records get revised, prices move, and counterparties change behaviour. Six months later nobody remembers whether a supplier was rejected on volume, on price band or on timing, and without that note the assessment simply gets repeated from scratch. A one-line rationale is what converts a list into institutional knowledge, and it costs seconds at the point where the thinking has already been done.

Finally, be explicit with colleagues about the confidence attached to any figure you circulate. A declared value from a complete period, controlled for origin and unit, is strong evidence. The same figure pulled from an incomplete recent period, averaged across a whole chapter, is barely evidence at all — and the two look identical once they are in a slide. Saying which one you have is what keeps trade data credible inside an organisation over time.

Frequently asked questions

Why does one company have several spellings?

Because the declaration records the name as typed, by different agents, over years. Legal suffix variants, abbreviations, transliteration, branch filings and keying errors all produce distinct strings for one business.

Is over-merging worse than under-merging?

Both distort, in opposite directions, and both are invisible in the output. Over-merging is arguably more dangerous because it creates a profile that looks impressively complete and is partly fictional.

Do identifiers solve the problem?

Where a jurisdiction publishes a stable entity identifier alongside the name, aggregation becomes deterministic rather than probabilistic. That is a material improvement in reliability, and it is why Indian company-level analysis is more dependable than most.

How does this affect market concentration analysis?

Directly. Unresolved names inflate counterparty counts and deflate individual volumes, which systematically makes concentrated markets look fragmented.

How current is the trade data behind this?

Markets refresh on their customs authority's own release cycle — monthly for most, 45 to 60 days for a few. The most recent one or two periods are always still filling in, so exclude them when you are reading a trend rather than treating the gap as a decline.

Can I check this against my own product?

Yes. Give us the HS code or a product description and the market you care about, and we will return a sample of live customs records filed against it.

Keep reading

Related guides

The next questions this one usually raises are covered in The Importer Exporter Code, explained, Competitor analysis with trade data and How often trade data is updated. Each picks up where this article stops, and together they cover the sequence a consignment actually goes through — classification and duty before anything moves, documentation and payment while it moves, and verification of the counterparty before any of it is committed to. Reading them in that order is usually more useful than reading them by topic.