Loading...
Loading...
Available on 1 platform
Sign in to view source links and access this dataset
A dataset designed to evaluate the diagnostic capabilities of Large Language Models in the domain of oral disease. It was created by Lines and last updated on April 28, 2024. The dataset includes evaluation data for models such as GPT-3.5, GPT-4, Palm2, and Llama2-70B.
License is unknown, which may restrict usage.