A new AI tool trained on 1.4 million publications aims to transform how scientists navigate the exploding volume of rice research data.

Rice feeds nearly half the world. Yet keeping pace with the science behind it has become a bottleneck for researchers. The problem isn’t a lack of data—it’s too much of it. Since the rice genome was sequenced in 2002, high-throughput technologies have flooded the field with transcriptomic, proteomic, and genomic datasets. Meanwhile, the scientific literature has grown exponentially. Extracting meaningful insights from this ocean of information remains labor-intensive and time-consuming, slowing the pace of discovery.
General-purpose large language models (LLMs) offer a partial fix, but they fall short on domain-specific tasks. Without standardized benchmarks for rice biology, evaluating their performance is guesswork. And crucially, these models struggle to synthesize the multimodal data—text, sequences, expression profiles—that rice research demands.
A research team from Yazhouwan National Laboratory, Shanghai AI Laboratory, and China Agricultural University has unveiled SeedLLM-Rice (SeedLLM), a 7-billion-parameter LLM purpose-built for rice biology.
In head-to-head evaluations against other AI models, SeedLLM achieved win rates of 57% to 88% on rice-specific tasks. The RBKG integration proved critical: when augmented with structured biological knowledge, SeedLLM significantly outperformed other AI models on advanced omics tasks.
The team has released both SeedLLM and the RBKG through an interactive web portal at https://seedscientist.cn/, freely available to researchers worldwide.