Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A benchmark dataset for evaluating whether a Large Language Model correctly allocates verification when orchestrating a fallible biology foundation model. The dataset, created by jang1563, includes a substrate table from the GEARS/Norman experiment. It was last updated on June 17, 2026.
License is unknown; users should verify licensing terms before use.