Skip to main navigation Skip to search Skip to main content

SBIR Phase I: Automatic Data Series Extraction from a Text Corpus

Project: Research

Abstract & Details

Description

Award ID: 2110123

The broader impact of this Small Business Innovation Research (SBIR) Phase I project will be improved market efficiency and greater competition in financial services and adjacent industries. Currently, these sectors are dominated by large firms that have the resources to create and exploit asymmetric information advantages over smaller firms. A key reason for this asymmetry is that company information the basis for building high-quality, detailed financial models can be prohibitively expensive to surface, not because it is unavailable, but because it is reported in non-standardized ways and therefore difficult to extract and make actionable. To do so systematically and industry-wide requires thousands of man-hours per year of manually sifting through millions of documents, a cost that only the largest firms in the world can bear. Automating the extraction of such information and making it both widely available and easily accessible helps to level the playing field for small firms while simultaneously improving the speed and quality of decision-making at a lower total cost for large firms. This Small Business Innovation Research (SBIR) Phase I project aims to develop a machine learning platform for automatically extracting data from a collection of financial documents. The platform will take advantage of recent advances in natural language processing model architectures, but nonetheless faces the challenges of (a) achieving and maintaining a high level of accuracy, even as document text volume, and thus semantic variation, grows; and (b) generating sufficient labeled training data in a cost-effective way. This project addresses these dual challenges with a novel framework for continuous model training based on recent meta-learning techniques. Such an approach to supervised learning can substantially accelerate model improvement and simultaneously drive down training costs. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.

NSF Program Director: Peter Atherton
StatusClosed
Effective start/end date08/01/2105/31/23

Funding

  • SBIR Phase I: $256,000.00

Active Fiscal Year

  • FY2023
  • FY2022

Start Fiscal Year

  • FY2021

TIP Programs

  • SBIR Phase I

Small Business

  • Yes

Key Technology Areas

  • Artificial Intelligence
  • (confidence score: 100%)

Technology Foci

  • Machine Learning Training Data
  • (confidence score: 100%)
  • Machine Learning (ML)
  • (confidence score: 100%)
  • Artificial Intelligence (excluding ML)
  • (confidence score: 100%)

Congressional District at Award

  • District n. 16 of California

Current Congressional District

  • District n. 16 of California

United States

  • California

Core Based Statistical Area (CBSA)

  • San Jose-Sunnyvale-Santa Clara, CA

County

  • County: Santa Clara, CA

Fingerprint

Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint. Learn more about Elsevier's Fingerprint Engine here: https://beta.elsevier.com/products/elsevier-fingerprint-engine