The second batch of “First Proof” problems is meant to evaluate AI’s usefulness for research-level math. The best model got ...
A new benchmark pitting AI against previously unseen maths problems shows systems still fall short of top human expertise.
Active Learning Network for Accountability and Performance in Humanitarian Action (ALNAP)’s Humanitarian Evaluation, Learning and Performance (HELP) Library offers a variety of resources on evaluation ...