Abstract & Details
Description
Award ID: 2621972
America's ability to lead in artificial intelligence (AI) depends not only on algorithms and computing hardware, but on the quality and openness of the data used to train AI systems. Today, the public infrastructure for open AI training data is falling behind the pace of the field, creating barriers that favor well-funded institutions over small research teams, startups, and universities, while foreign competitors build proprietary, non-auditable data pipelines outside American oversight. This project establishes a public-philanthropic consortium to modernize the foundational data layer of AI development. The consortium reduces the cost and complexity of accessing large-scale web data, making it more affordable for a wider range of organizations to participate in cutting-edge AI research. It also expands the linguistic reach of publicly available training data from 160 to over 300 languages, ensuring that American-developed AI systems are competitive globally and can serve the needs of allied nations. By maintaining this critical infrastructure as an open, transparent, and auditable public good under American stewardship, the project protects against actors who seek to control how the world's AI systems are trained through closed, unaccountable data pipelines. This project addresses a set of challenges now facing openly available, large-scale AI training datasets: rising access costs, uneven linguistic coverage, and weak provenance across the data supply chain. It uses the Common Crawl (CC) web corpus (one of the most widely used such datasets) as the vehicle for developing solutions, transitioning its pipeline toward a componentizable, modular architecture whose methods generalize to other open corpora. The technical work proceeds along two thrusts. First, the project lowers data access costs through cross-cloud optimization and strengthened web crawling protocols, with an estimated reduction in technical overhead of 30 to 50 percent for small research teams. Second, the project expands the linguistic scope of the datasets from 160 to over 300 languages by integrating real-time natural language processing (NLP) methods for proactive language detection. This approach pairs human-curated seed lists with low-resource language labeling models to actively guide crawlers toward underrepresented languages, shifting from post-hoc corpus analysis to proactive, intelligent data collection. To execute these activities, the project convenes a stakeholder workshop to prioritize technical requirements, then funds five to eight expert teams through a managed grant lifecycle. Research outputs include independent, reusable modules for deduplication and parsing, a standardized framework for metadata and relationship mapping, and provenance guarantees across the AI data supply chain. All results are released as open, reusable components (modules, standards, and provenance tooling applicable across open training corpora), ensuring lasting public benefit from the investments made under this award. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
NSF Program Director: Jemin George
America's ability to lead in artificial intelligence (AI) depends not only on algorithms and computing hardware, but on the quality and openness of the data used to train AI systems. Today, the public infrastructure for open AI training data is falling behind the pace of the field, creating barriers that favor well-funded institutions over small research teams, startups, and universities, while foreign competitors build proprietary, non-auditable data pipelines outside American oversight. This project establishes a public-philanthropic consortium to modernize the foundational data layer of AI development. The consortium reduces the cost and complexity of accessing large-scale web data, making it more affordable for a wider range of organizations to participate in cutting-edge AI research. It also expands the linguistic reach of publicly available training data from 160 to over 300 languages, ensuring that American-developed AI systems are competitive globally and can serve the needs of allied nations. By maintaining this critical infrastructure as an open, transparent, and auditable public good under American stewardship, the project protects against actors who seek to control how the world's AI systems are trained through closed, unaccountable data pipelines. This project addresses a set of challenges now facing openly available, large-scale AI training datasets: rising access costs, uneven linguistic coverage, and weak provenance across the data supply chain. It uses the Common Crawl (CC) web corpus (one of the most widely used such datasets) as the vehicle for developing solutions, transitioning its pipeline toward a componentizable, modular architecture whose methods generalize to other open corpora. The technical work proceeds along two thrusts. First, the project lowers data access costs through cross-cloud optimization and strengthened web crawling protocols, with an estimated reduction in technical overhead of 30 to 50 percent for small research teams. Second, the project expands the linguistic scope of the datasets from 160 to over 300 languages by integrating real-time natural language processing (NLP) methods for proactive language detection. This approach pairs human-curated seed lists with low-resource language labeling models to actively guide crawlers toward underrepresented languages, shifting from post-hoc corpus analysis to proactive, intelligent data collection. To execute these activities, the project convenes a stakeholder workshop to prioritize technical requirements, then funds five to eight expert teams through a managed grant lifecycle. Research outputs include independent, reusable modules for deduplication and parsing, a standardized framework for metadata and relationship mapping, and provenance guarantees across the AI data supply chain. All results are released as open, reusable components (modules, standards, and provenance tooling applicable across open training corpora), ensuring lasting public benefit from the investments made under this award. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
NSF Program Director: Jemin George
| Status | Not started |
|---|---|
| Effective start/end date | 10/01/26 → 09/30/28 |
Funding
- (ET) Emerging Technology Pilots: $6,175,403.00
Active Fiscal Year
- FY2028
- FY2027
Start Fiscal Year
- FY2027
TIP Programs
- (ET) Emerging Technology Pilots
Key Technology Areas
- Artificial Intelligence
- (confidence score: 100%)
- Advanced Computing and Semiconductors
- (confidence score: 85%)
Technology Foci
- Advanced Computing and Semiconductors (Broad)
- (confidence score: 100%)
- Machine Learning Training Data
- (confidence score: 100%)
- Artificial Intelligence (excluding ML)
- (confidence score: 100%)
Congressional District at Award
- District n. 00 of District of Columbia
Current Congressional District
- District n. 98 of District of Columbia
United States
- District of Columbia
Core Based Statistical Area (CBSA)
- Washington-Arlington-Alexandria, DC-VA-MD-WV
County
- County: District of Columbia, DC
Fingerprint
Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint. Learn more about Elsevier's Fingerprint Engine here: https://beta.elsevier.com/products/elsevier-fingerprint-engine