Welcome to Smart Agriculture 中文

Smart Agriculture ›› 2026, Vol. 8 ›› Issue (3): 67-84.doi: 10.12133/j.smartag.SA202602016

• Special Issue--Digital Technologies Reshaping Agriculture and Agricultural Economics • Previous Articles     Next Articles

Artificial Intelligence Empowering Modern Agricultural Biological Breeding

XU Qiyu1, ZHENG Bo1,2(), ZHONG Shangwei1,3   

  1. 1. Peking University Institute of Advanced Agricultural Sciences, Weifang 261325, China
    2. Weifang science and technology innovation promotion center, Weifang 261071, China
    3. College of Life Sciences, Peking University, Beijing 100871, China
  • Received:2026-02-09 Online:2026-05-30
  • Foundation items:National Natural Science Foundation of China(32300248); Natural Science Foundation of Shandong Province(ZR2022QC010)
  • About author:

    XU Qiyu, E-mail:

  • corresponding author:
    ZHENG Bo, E-mail:

Abstract:

[Significance] The escalating complexity of genotype-phenotype-environment interactions and the explosive growth of multi-omics big data have necessitated a paradigm shift in crop breeding from empirical selection to intelligent design (Breeding 5.0). The aim of this paper is to systematically explore the underlying logic of the AI-driven crop breeding paradigm shift, comprehensively outline its generational evolution, analyze its core technical implementations in phenomics, genomics, multi-omics integration, and molecular design, and dissect how large agricultural foundation models reconstruct the entire seed industry workflow. [Progress] The historical evolution of crop breeding was first traced from 1.0 empirical domestication to 5.0 smart Breeding characterized by the deep integration of biotechnology (BT) and information technology (IT). In germplasm resource evaluation, deep learning algorithms enabled high-dimensional pattern recognition and unsupervised feature compression to unlock rare alleles from unannotated sequences, large language models (LLMs) like PlantConnectome and wheat germplasm information extraction (WGIE) leverage retrieval-augmented generation (RAG) to automatically construct structural knowledge graphs from unstructured historical literature, achieving predictive and dynamic germplasm evaluations. In high-throughput phenotyping, industrial platforms capture 3D point cloud and multi-spectral data at 0.1 mm resolution, while convolutional neural networks couple with the integrated genomic-enviromic prediction (iGEP) framework to build full-lifecycle digital twin crop models in virtual space. Regarding genomic prediction, the limitations of linear paradigms were dissected and cutting-edge deep learning architectures were highlighted: SoyDNGP applied a 3D-CNN to map chromosomal topology for complex soybean traits; DPCformer employed self-attention mechanisms to dynamically calculate environmental weights under multi-adversarial constraints; and HyenaDNA utilized long-convolution filters to bypass the O(N2) computational complexity limitation of standard Transformers, reducing it to O(N·log N) for chromosome-scale modeling. For multi-omics integration, intermediate fusion strategies were elucidated for their superior capacity to capture cross-layer biological compensatory pathways. In molecular design breeding, foundational plant language models were highlighted, such as AgroNT for zero-shot expression prediction and OpenCRISPR-1, the world's first de novo AI-generated genome editor built to capture underlying physico-chemical syntax. Furthermore, the emergence logic of major domestic and international agricultural foundation models was analyzed through the mathematical lenses of scaling laws, parameter-efficient fine-tuning, and multi-task evaluation benchmarks. Finally, empirical effectiveness was comprehensively evaluated through multinational success stories, including Bayer's Climate FieldView, IRRI's night-temperature thermal models for "Green Super Rice", and China's state-led breeding platforms for stress-resistant maize and high-yield soybean. [Conclusions and Prospects] Key structural challenges that smart breeding faces are thoroughly dissected: data silos and standardization dilemmas, extreme computational resource asymmetry and high training costs, the lack of biological causal logic in deep learning "black boxes", and the structural scarcity of interdisciplinary BT-IT talent. To overcome these bottlenecks, future research and policy efforts should focus on four pillars: 1) Establishing standardized open data ecosystems following FAIR (Findable, Accessible, Interoperable, Reusable) principles and promoting federated learning under a national AI data copyright integration platform to safeguard digital borders; 2) Exploring cloud-to-edge lightweight model deployment pathways through parameter-efficient fine-tuning, quantization, and knowledge distillation to empower real-time field-side decision-making and realize compute equity; 3) Deepening the mechanistic integration of AI and synthetic biology by absorbing breakthroughs from multi-scale modeling frameworks into the virtual "design-build-test-learn" cycle to break natural evolutionary thresholds; and 4) improving regulatory approval frameworks for AI-designed crops and modernizing agricultural education to cultivate a new generation of digital agronomists.

Key words: smart breeding, artificial intelligence, breeding model, multi-omics fusion, high-throughput phenotyping, genomic selection, intelligent design breeding

CLC Number: