OpenAI built a benchmark to test if LLMs actually understand genomics
8 min read
AI Benchmarking
OpenAI released GeneBench-Pro, a new evaluation framework designed to test AI models on complex genomics and biological research tasks. It moves beyond standard text benchmarks, forcing models to prove their utility on real-world scientific datasets....