Skip to main navigation Skip to search Skip to main content

I-Corps: Translation potential of verifiable reward-signal generators for agentic reinforcement learning (RL) environments

Project: Research

Abstract & Details

Description

Award ID: 2628056

This I-Corps project is based on the development of a software platform that generates feedback for training and evaluating artificial intelligence (AI) systems. As AI is increasingly trained to act on its own through trial and error, the central bottleneck is evaluating a system for correctness. Currently, this depends on costly human review or on hand-built checks that automated systems learn to exploit rather than satisfy, slowing progress across an industry now investing tens of billions of dollars annually. This technology is designed to generate precise, machine-checkable feedback signals at very low cost, allowing developers to train and measure AI against verifiable standards rather than approximations. This applies to most uses of AI including software development, computer chip design, cybersecurity, and scientific computing. This may make machine behavior easier to measure and support a safer, more transparent, and more competitive technology economy. This I-Corps project utilizes experiential learning coupled with first-hand investigation of the industry ecosystem to assess the translation potential of a platform for reinforcement learning from verifiable rewards (RLVR) that generates correctness feedback at near-zero marginal cost. The core technology combines formal methods, including temporal logic specification, automated synthesis, and satisfiability checking, with physical compute substrates, principally field-programmable gate arrays (FPGAs), to evaluate machine behavior against mathematically defined correctness criteria. Unlike conventional reward design, which relies on heuristic scoring or human labeling that agents can game, this approach derives feedback from formal guarantees, producing signals that are precise, reproducible, and resistant to exploitation. Certain reward dimensions, including real contention throughput, hardware side-channel behavior, and device-level process variation, are physically irreducible and cannot be faithfully simulated, motivating execution on physical FPGA fabric. Demonstrated proof-of-concept instances include benchmark environments for temporal-logic reasoning and high-throughput satisfiability solving, where verifiable rewards are produced automatically and at scale. Developers training autonomous agents, evaluating hardware design tools, and stress-testing system security may benefit from feedback that is cheaper and faster than current practice. This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.

NSF Program Director: Ruth Shuman
StatusActive
Effective start/end date08/01/2607/31/27

Funding

  • I-Corps Teams: $50,000.00

Active Fiscal Year

  • FY2027
  • FY2026

Start Fiscal Year

  • FY2026

TIP Programs

  • I-Corps Teams

Key Technology Areas

  • Artificial Intelligence
  • (confidence score: 100%)
  • Advanced Computing and Semiconductors
  • (confidence score: 100%)

Technology Foci

  • Semiconductors
  • (confidence score: 99%)
  • Machine Learning Training Data
  • (confidence score: 99%)
  • Advanced Computer Software
  • (confidence score: 95%)
  • Machine Learning (ML)
  • (confidence score: 97%)
  • Artificial Intelligence (excluding ML)
  • (confidence score: 90%)
  • Autonomy
  • (confidence score: 99%)

Congressional District at Award

  • District n. 13 of New York

Current Congressional District

  • District n. 13 of New York

United States

  • New York

Core Based Statistical Area (CBSA)

  • New York-Newark-Jersey City, NY-NJ

County

  • County: New York, NY

Fingerprint

Explore the research topics touched on by this project. These labels are generated based on the underlying awards/grants. Together they form a unique fingerprint. Learn more about Elsevier's Fingerprint Engine here: https://beta.elsevier.com/products/elsevier-fingerprint-engine