DTWdailytechwire
Tech Intelligence, Wired Daily
Policy

Communities Build Data Collectives to Challenge Big Tech's Scraping Economy

Grassroots cooperatives in Pakistan, Africa, and India are negotiating directly with AI companies, ensuring local voices control how their linguistic and cultural data are used.

PN
Priya Nair
Staff Writer · Singapore
Jul 24, 2026
9 min read
Communities Build Data Collectives to Challenge Big Tech's Scraping Economy
Communities Build Data Collectives to Challenge Big Tech's Scraping EconomyCredit: iStock

The Power Shift in Data Ownership

When Meesum Alam first documented Dawoodi, a language spoken by roughly 300 people, he wasn't thinking about AI companies. He was grappling with what he calls a "linguistic identity crisis," unable to speak Balochi, the language of his Baloch ancestors in Pakistan. That personal disconnect drove him to record endangered languages before they vanished entirely.

Today, those recordings have become negotiating chips. Alam's voice data sets covering 39 Pakistani languages, totaling around 700 hours, sit on Mozilla Data Collective, where Meta and other firms must now ask permission to use them. The communities that generated this data dictate the terms: research or non-commercial use only. No blanket scraping. No silent exploitation.

This represents a fundamental recalibration in the relationship between communities and the companies that profit from their data. At DailyTechWire, we've tracked the rising tension between frontier AI labs and the populations whose languages, music, and cultural artifacts fuel training runs. Data collectives are emerging as a structural counterweight, one that treats information not as a commons to be harvested but as a resource requiring consent, compensation, and accountability.

Why Cooperatives Are Gaining Traction Now

The generative AI boom has made the value of data impossible to ignore. American and Chinese frontier models dominate the market, built on datasets scraped from the public web. Companies including OpenAI, Meta, Google, and Anthropic have ingested nearly everything accessible online, prompting lawsuits over copyrighted material and accusations of fair use abuse.

For communities with smaller or linguistically unique datasets, the traditional model offers no upside. Their languages don't appear in GPT or Gemini. Their cultural knowledge is absent from Claude or Qwen. Some governments have responded by building sovereign large language models, but that path requires capital and technical capacity most communities lack.

Data collectives offer an alternative architecture. They pool resources, share benefits, and impose governance structures that communities control. Raffi Krikorian, chief technology officer at Mozilla Foundation, describes the shift as communities "sitting on unique information, not generally available on the internet," and deciding to "turn the tables on what governance looks like for that data."

Mozilla Foundation launched Mozilla Data Collective in 2025 to provide a platform for exactly this kind of community-generated dataset. The timing was deliberate. The United Nations had designated 2025 as the International Year of Cooperatives, framing them as solutions to global challenges. That spotlight has accelerated interest in cooperative models for data stewardship.

From Fishers to Farmers to Language Activists

Data collectives span a wide range of use cases. Kerala Food Platform in southern India enables around 2,500 farmers to trace and market rice, fish, fruits, and vegetables. PescaData, based in Mexico, helps small-scale fishers across Latin America and the Caribbean manage catch records. The Native BioData Consortium stewards genetic and environmental data from Indigenous populations, ensuring researchers can't extract samples without consent.

Increasingly, these collectives are forming around AI-specific use cases. Low-resource language communities recognize that their voice data is valuable for training speech recognition tools and small models. For a language spoken by 1 million people, no frontier lab will prioritize development. But a locally trained chatbot that communicates in that language delivers real utility, and a data collective makes that possible without surrendering control.

Astha Kapoor, co-founder of Aapti Institute, a tech research firm in India, argues that collectives go beyond consent and compensation. "Collective action around data gives communities the opportunity to direct data towards issues they may care about," she notes. Communities can negotiate terms at every stage of the AI lifecycle and establish mechanisms for accountability if those terms are breached.

This is not a hypothetical framework. Alam's datasets have been used by Meta to build speech recognition tools for languages that were "never part of the digital world," as he puts it. For the communities involved, interacting with AI in their own language represents a first step into the digital economy. They chose research or non-commercial licensing not because they distrust technology, but because they distrust the extractive practices of large firms.

"They don't trust the big tech companies because they know they can take the data and monetize it," Alam explains. "They want a fair deal for the entire community, and they can only get that through a data collective."

African Languages Enter the Negotiation

In Africa, where extractive practices have deep historical roots, the Nwulite Obodo Open Data License launched in 2024 to enable creators and communities to share datasets without forfeiting the right to benefit from them. Around 70 African datasets now sit under NOODL within Mozilla Data Collective, covering speech data in more than 20 languages, plus music, lullabies, and poetry.

Emmanuel Ngue Um, a regional researcher at the Institute of African Digital Humanities, has published roughly 40 datasets on the platform. Most of these languages fall outside mainstream linguistic frameworks, making them invisible to model developers. Hosting them in a data collective increases both visibility and accessibility while preserving community control.

Crucially, users must clarify the purpose of their access request. This creates opportunities for cooperation between those with technical capacity to support language work and the communities themselves. It also shifts the default from extraction to partnership.

Earlier in 2026, Mozilla Data Collective began offering three community-generated datasets for paid commercial licensing. The feature will expand to all users, with the stated goal of ensuring communities are "valued, recognized and supported." This marks a pragmatic evolution. Collectives that rely solely on volunteer stewardship face sustainability challenges. Monetization, if structured correctly, can fund ongoing work without undermining the cooperative's mission.

Kapoor cautions that sustainability remains a core challenge. "Building data cooperatives solely to steward data is not feasible because sustainability becomes an issue, and the only viable pathway becomes monetization of the data the cooperative is meant to safeguard, which is problematic," she observes. The tension between financial viability and mission integrity will define the next phase of collective growth.

Alternative Governance Models

Data collectives are not the only framework in use. Data trusts appoint a trustee to manage data on behalf of a group. Data unions aggregate individual members' data to negotiate collectively with buyers. Data commons, such as Wikimedia and OpenStreetMap, operate under different governance structures with varying degrees of openness. Data donation schemes, like the Personal Genome Project, rely on individuals contributing information for public benefit.

Each model has trade-offs. Trusts centralize decision-making, which can streamline negotiations but reduce transparency. Unions prioritize individual member interests, which can fragment collective bargaining power. Commons maximize accessibility but offer limited control over downstream use. Donation schemes depend on altruism, which may not scale in commercial contexts.

Data collectives strike a balance. They centralize enough authority to negotiate effectively while maintaining community governance over terms of use. They generate economic value without surrendering ownership. They scale through federation, allowing smaller collectives to coordinate on shared interests.

The Emotional Dimension

For Alam, the impact extends beyond economics. He receives texts and voice notes from people who have communicated with a chatbot in their native language for the first time. "It's a very emotional thing for them when they can do that," he says.

This emotional dimension matters. Language is not just a communication tool; it's a carrier of identity, history, and belonging. When a language disappears from the digital world, it accelerates its disappearance from the physical world. Children grow up unable to speak with their grandparents. Cultural knowledge encoded in syntax and vocabulary vanishes. Communities lose a piece of themselves.

Data collectives offer a path to digital inclusion that doesn't require assimilation into dominant linguistic or cultural frameworks. They allow communities to participate in the AI economy on their own terms, preserving autonomy while gaining access to tools that were previously out of reach.

What This Means for the AI Industry

The rise of data collectives signals a shift in the power dynamics of AI development. Frontier labs have operated under the assumption that publicly accessible data is fair game for training. That assumption is now being challenged by communities asserting ownership and demanding negotiation.

This will complicate data sourcing. Instead of scraping the open web, companies will need to engage with collectives, negotiate licensing terms, and build partnerships. This adds friction, but it also creates opportunities. Companies that invest in these relationships gain access to high-quality, domain-specific datasets that competitors lack. They build trust with communities that may become early adopters of their products.

It also redistributes value. If collectives successfully monetize their datasets, revenue flows back to the communities that generated the data rather than concentrating in the hands of model developers. This won't eliminate inequality, but it shifts the baseline.

Krikorian frames the shift bluntly: "The big companies have built themselves up on the backs of all these people creating data, who think it's time to set their own terms now." That sentiment is driving the formation of new collectives and the politicization of data governance.

Scaling Challenges Ahead

Despite the promise, collectives face real obstacles. Governance structures can become unwieldy as membership grows. Deciding who speaks for the community, how decisions are made, and how disputes are resolved requires institutional capacity that many grassroots organizations lack.

Monetization introduces its own risks. If a collective becomes too dependent on licensing revenue, it may prioritize commercial partnerships over community interests. If it charges too much, it prices out researchers and nonprofits that could deliver real value. If it charges too little, it fails to cover operational costs.

There is also the question of scale. A collective representing 300 speakers of Dawoodi has leverage because the data is rare. A collective representing speakers of a more common language may struggle to compete with freely available datasets. The model works best for communities with unique, hard-to-access data, which limits its applicability.

Alam is now expanding his work to low-resource languages in India and Bangladesh. He sees data collectives as the only viable path for many communities to participate in the AI economy. Without them, these languages remain invisible to model developers. With them, communities gain a seat at the table.

The broader question is whether collectives can scale beyond niche use cases. If they remain limited to low-resource languages and Indigenous datasets, they will serve an important but narrow function. If they expand to encompass larger communities and more common datasets, they could reshape the data economy.

A New Baseline for Data Governance

Data collectives are not a panacea. They won't solve the structural inequalities that shape AI development. They won't prevent large companies from continuing to scrape publicly accessible data. They won't eliminate the need for regulation or legal frameworks that protect creators.

But they do establish a new baseline. They demonstrate that communities can organize, assert control, and force companies to negotiate. They show that data governance can be participatory rather than extractive. They prove that economic value can flow back to the people who generate the data.

At DailyTechWire, we've watched the AI industry grapple with questions of consent, compensation, and accountability since the launch of ChatGPT. Data collectives represent one of the more promising answers, precisely because they don't rely on corporate goodwill or regulatory intervention. They build power from the ground up.

Whether that power translates into lasting change will depend on the choices communities make in the coming years. Monetization strategies, governance structures, and partnerships will all shape the trajectory. But the fundamental shift is already underway. Communities are no longer passive sources of data. They are active participants in the economy their data enables.

For Alam and the communities he works with, the goal is clear. "They want a fair deal for the entire community," he says. Data collectives give them the tools to negotiate one.

Read next
Policy

AI Researchers Challenge Washington's Intellectual Property Claims on Model Distillation

Arjun S. Mehta · 5 min
Policy

Why Distillation Alone Can't Explain Kimi K3's Leap

Arjun S. Mehta · 6 min
Policy

Pentagon Moves to Block Humanoid Robots From Chinese Suppliers

Arjun S. Mehta · 6 min
Spot something wrong? Email corrections@dailytechwire.com. We log every correction publicly.