OpenRSI Index opens long-horizon benchmark for model-development agents
OpenRSI Foundation has opened Preview v0.1, an Apache-2.0 benchmark stack that keeps a research agent in a persistent work container and evaluates each submission in a fresh judge with private tests. The catalogue ranges from smaller public tasks to model-training runs measured in thousands of H100-hours, while the preview remains tied to moving main rather than an immutable software release.