The confirmed problem
DARPA is seeking a black-box test environment that uses economic frameworks—such as auctions and markets—to characterize latent AI-agent preferences. The environment should measure how agents adapt, collaborate or deceive as information and incentives change, relying on queries and outputs rather than model internals.
Working market sandbox
Configurable market structures, payouts, utility and efficiency measures, plus a dynamic public news feed.
Comparable behavior
Behavioral classifiers, reference agents from at least ten LLMs and a proof of concept above 90% allocative efficiency.
Proposed contribution
I can deliver a bounded, non-confidential evaluation work package drawing on my public product experience with GERO and ATLAS Oracle:
- a sealed-bid auction and posted-price market;
- controlled dynamic news shocks;
- three to five reference agents for the initial partner demo;
- allocative-efficiency, utility and distribution-shift measurements;
- behavioral flags and replayable, versioned run evidence;
- a Phase I-aligned architecture, test plan, milestones and risk register.
Proprietary algorithms, private prompts, routing and internal scoring logic are neither disclosed nor transferred.
Fixed-scope starting point
Partner demonstrator
USD 4,900 fixed. A small executable specimen with declared acceptance tests and a reproducibility package.
Evaluation work package
USD 15,000 fixed. Architecture, milestones, quantitative test plan, implementation specimen and delivery risks.
These are subcontract work-package prices, not statements about the DARPA award ceiling.
Timing and official sources
The topic was published on 2 September 2026 and opens on 23 September. The detailed DARPA topic page lists 21 October as the close date, while the opportunity summary lists 23 October. The live DSIP solicitation must therefore be treated as controlling.
Eligible U.S. primes preparing a response are invited to contact me with their proposed role split and compliance approach.
Discuss the work package