Gates Foundation launches coalition to build more representative language data sets for AI
LinguisticsInformation SystemsPublic Policy
THE AI ANGLE
Processing underrepresented languages and local dialectsThe Gates Foundation formed a coalition with tech companies and philanthropies to build AI data sets for underrepresented languages. This group plans to coordinate speech and text collection over five years to reach three billion people. The work aims to replace unrepresentative web-scraped training data with community-sourced local dialects.
THE TEACHING ANGLE
Students can examine whether large coalitions can establish sound governance while securing explicit consent from local communities during field data collection.Read the original at techxplore.com Generate teaching or study materials
More in Linguistics
- Stop relying on AI chatbots for customer care, UK banks and energy firms toldThe Guardian — Technology · September 22, 2026
- Meta’s AI agent has been blocked from using Amazon.comTechCrunch · September 22, 2026
- Treasury chief says AI bosses, not their bots, will carry the can for criminal actsThe Register · September 22, 2026
- Privacy group slams EU for changing the data rules to cater to AIThe Register · September 22, 2026
- AI and Financial RegulationKnowledge at Wharton · September 22, 2026