News

China Team Unveils SeedBench, First Benchmark for Evaluating LLMs in Seed Science

First Benchmark for Evaluating LLMs in Seed Science Launches to Standardize AI Breeding SeedBench offers a standardized way to test whether large language models can actually help crop breeding research.

Seeds are the chips of agriculture. Yet gaps remain between China’s seed industry and global leaders, with some high-end germplasm still relying on imports. The practical challenges of breeding innovation are clear: long R&D cycles, scattered professional data, complex interdisciplinary demands, and a shortage of specialized talent.

Large language models (LLMs) open new possibilities. By learning from massive datasets, they could break down disciplinary barriers and drive the digital transformation of breeding. But scarce professional data and the lack of standardized evaluation systems are constraining how LLMs can be deployed in intelligent breeding.

SeedBench data is designed and validated by breeding experts. The team works with domain specialists to simulate real breeding scenarios, implementing a rigorous two-stage validation process.

Why SeedBench Is Needed

The global seed industry is transitioning from experience-based to intelligent breeding. FAO data shows global crop yield increases have exceeded 50% over the past two decades, with technological progress as the core driver. Advances in genomic sequencing mean a single trait may be regulated by hundreds of gene loci—traditional manual analysis can no longer keep pace, making data-driven AI integration essential.

But LLM applications in breeding face specific hurdles:

Data scarcity

Breeding-related data represents a small fraction of internet content, limiting model training. Some field records remain on paper, and substantial tacit knowledge has yet to be digitized.

Evaluation gap

Medicine, law, and finance have established benchmarks (MedBench, LawBench, FinBench). Breeding lacks comprehensive, full-process evaluation standards, leaving LLM optimization without clear direction.

Interdisciplinary complexity

Breeding spans genetics, molecular biology, environmental science, and more. LLMs must understand complex gene-trait relationships and generate field-applicable recommendations.

From Evaluation to Application

Fengdeng (SeedLLM), the first seed industry large language model evaluated through SeedBench, is now open for application at https://seedscientist.cn/, providing intelligent assistance for breeding research.