News

“Fengdeng” Awarded Hainan’s First Data Intellectual Property Registration Certificate

A structured agricultural knowledge database developed by the Fengdeng team becomes the first data asset to receive official IP protection in the province

On December 8, 2024, Hainan Province issued its first data intellectual property registration certificate to the Fengdeng research team. The registered asset: a "multi-omics knowledge fusion graph for intelligent breeding of staple crops."

During the development of the Fengdeng seed industry large language model, the team systematically collected and processed massive volumes of publicly available academic literature on staple crops including rice, corn, and soybean from around the world. After deep cleaning, precise annotation, and multi-dimensional fusion, the team completed this knowledge graph in April 2024. It serves as the underlying data infrastructure for the Fengdeng model and has become an independent data asset in its own right.

The core value of this dataset lies in transforming fragmented breeding knowledge scattered across millions of papers into a structured, computable relationship network—how genes link to traits, how traits respond to environmental factors, which publications provide supporting evidence. Information that previously required researchers to spend countless hours manually compiling is now organized in graph form.

Where is it being used?

The data currently supports three main applications:

Scientific agent knowledge retrieval

It provides real-time, precise evidence chain support for specialized agents such as Fengdeng · Gene Scientist. When researchers pose questions to the AI, the system delivers answers alongside traceable literature and data sources, improving both efficiency and reliability.

Breeding decision support

For breeding enterprises and research teams, it offers three-dimensional queries across gene–trait–environment relationships. Researchers can rapidly locate candidate genes and review supporting literature, reducing time spent on information screening.

LLM output calibration

It supplies traceable, structured professional knowledge to ground the outputs of general or specialized large language models, helping identify and correct factual deviations in generated content and mitigating hallucination risks.

Why Data Ownership Matters

Data intellectual property registration remains in exploratory stages in China. Hainan’s first certificate marks the first time structured data accumulated through scientific research has been formally recognized as a protectable intellectual property object at the institutional level.

For intelligent breeding, this carries several practical implications:

First, clarifying data asset ownership. The collection, cleaning, and structuring of breeding data involves substantial research investment, yet clear property rights were previously lacking. The registration system provides legal confirmation for such efforts.

Second, enabling data circulation. Clear ownership is a prerequisite for data transactions, licensing, and collaborative development. Data sharing between breeding enterprises and research institutions requires this institutional foundation.

Third, driving industry standardization. When data can be registered, priced, and transferred, the industry gains stronger incentives to establish unified data formats and quality standards, reducing interoperability costs across systems.

Background: From model to data asset

Fengdeng is China’s first large language model for the seed industry, launched in April 2024 by Yazhouwan National Laboratory, Shanghai AI Laboratory, and China Agricultural University. The model was trained on approximately 1.4 million rice-related academic publications.

The knowledge graph data registered this time represents an intermediate product from that development process. The team extracted, refined, and formally established rights over this data product, making it an independently tradable asset.

Industry significance

As artificial intelligence and breeding research deepen their integration, the core value of scientific data is becoming increasingly prominent. Hainan’s practice offers a reference case for data protection and value conversion in the intelligent breeding industry nationwide.

From a longer-term perspective, the standardization, propertization, and industrialization of breeding data relate directly to building independent innovation capacity in seed science and ensuring food security. The establishment and refinement of data intellectual property systems are foundational infrastructure for this progression.