Join our FREE personalized newsletter for news, trends, and insights that matter to everyone in America

Newsletter
New

Why Global Migration Data Should Recognise Subgroups Like The Teochew

Card image cap

2026.08.04 01:50
David Tran, a Teochew-Vietnamese refugee who settled in California, created the world-famous Sriracha sauce, turning a regional flavour profile into a global household staple. Photo: Griffin Hammond

In countries like France and Norway, the government census strictly avoids asking about race and ethnicity. Governed by a desire for national unity, these states operate on a firm premise: categorising citizens by skin colour risks creating divisions and reinforcing prejudice.

Across the Atlantic and in the Commonwealth, you find a different philosophy. The United States, United Kingdom and Canada measure race and ethnicity. In these societies, data collection isn’t seen as a wedge to divide people, but as an essential tool to spot discrimination, design better social safety nets and address health disparities.

Yet both models are falling short. As global migration accelerates and people move fluidly between countries, our basic data buckets are failing to keep up. By lumping vast, highly diverse populations into monolithic labels like “Asian”, “Chinese” or “Indian”, standard census forms miss the rich subgroup dynamics that drive business, culture and community life.

Take the Teochew people, to whom I belong, a subgroup originating from the coastal Chaoshan region of Guangdong province. On a typical census form in North America or Europe, a Teochew individual is recorded simply as “Chinese” or “Asian”. Yet many Teochews in the West are also diasporic migrants from Southeast Asia self-identifying as Vietnamese, Cambodian, Thai or Singaporean.

A single checkbox erases a complex identity, a deep history of migration and an economic network that spans the globe.

Look at the box-office success of Dear You, a film rooted in the Teochew language and heritage. It touched a raw nerve across the global diaspora by portraying the history of , the family letters and remittances sent home by migrant workers to southern China.

In technological innovation, Moonshot AI founder Yang Zhilin was born in Shantou in the Chaoshan area. His rise reflects a long-standing tradition of Teochew entrepreneurs carving out new frontiers.

In global business, Li Ka-shing built one of Asia’s most powerful empires, drawing strength from deep-rooted Chaoshan commercial networks. In consumer goods, David Tran, a Teochew-Vietnamese refugee who settled in California, created the world-famous Sriracha sauce, turning a regional flavour profile into a global household staple.

When statistical systems reduce figures as varied as Li, Tran and Yang to a single generic label, they miss the cultural engine and connections that power their journeys.

This isn’t unique to Chinese communities. Consider the Indian diaspora, where broad national labels conceal vast internal realities. A Gujarati family running a business in London, a Malayali nurse working in the Gulf and a Telugu tech worker in Silicon Valley live in completely different linguistic worlds, rely on different community networks and face different economic realities.

In our age of big data and artificial intelligence, relying on crude demographic categories is a major missed opportunity. For businesses, treating a massive, diverse group as a single monolithic market leads to tone-deaf advertising and wasted budgets. A marketing push designed for the “Asian consumer” is likely to miss the cultural cues, food preferences and language nuances that set a Teochew family apart from a Cantonese one.

Teochew iron-rod puppets take centre stage at the Hong Kong Chiu Chow Festival on June 6, part of Teochew’s rich intangible cultural heritage. Photo: Dickson Lee

For public policy, the consequences are serious. When governments measure average income or health using broad racial categories, wealthier subgroups can hide deep poverty or health struggles in less fortunate communities. If a minority group looks successful “on average”, its struggling sub-communities risk being overlooked when governments distribute health resources, language support or social assistance.

This case has been made time and again in Asian American studies, notably through data disaggregation efforts to foreground ethnicities such as Chinese and Indian – although I am arguing to take it one step further by foregrounding subgroups within ethnicities. The answer isn’t to stop collecting population data or force people into rigid boxes. Instead, data science needs to catch up with how modern migration works.

Survey designers, market researchers and policy analysts need to collect more voluntary, flexible subgroup data, accounting for the language spoken at home, regional roots and migration history.

There’s also the need to incorporate the rich findings and theorising from decades of intricate qualitative research about ethnic economies: the economic relationships and networks that emerge from cultural and linguistic ties. Once a thriving area of academic study, it has made little impact on economic and social policymaking.

In an interconnected world, identity is layered, local and fluid. Until our data systems learn to see past broad labels and recognise subgroups like the Teochew, our picture of global migration – and our ability to build better societies – will remain incomplete.