Benchmark
2 articles
Advertisement
Articles2
Given Only A Bibliography, Frontier Models Recover The Idea 3 To 15 Per Cent Of The Time
A benchmark called Reconstruction asks models to recover a paper's central idea from its reference list alone....
MA
22 Aug
Claude Sonnet Leads Real-World Web Agent Benchmark — But Only Completes 1 in 3 Tasks
A new benchmark called ClawBench tests AI agents on 153 real tasks across 144 live production websites — booki...
RE
1 May
Advertisement