AgriSkill-Agent — Code & Showcase

Repository: github.com/Wispertise/AgriSkill-Agent

System Overview

AgriSkill-Agent accepts either a leaf image accompanied by a natural-language request or a text-only agricultural question. GPT-4o serves as the central agent and activates task-specific skills for phenotype analysis, disease diagnosis, management support, and report generation. Image consultations invoke specialist visual models, whereas text-only questions proceed directly to knowledge-supported reasoning.

Procedural Skills

Each skill defines a role, procedure, permitted tools, domain rules, and output schema. This makes tool invocation reproducible: visual requests activate phenotype analysis and deterministic measurement; diagnosis requests combine observations with classifier candidates and literature evidence; management requests query literature and structured pesticide-use records; report generation integrates the stage outputs.

User request
  ├── Phenotype Analysis Skill
  │     ├── constrained image description
  │     ├── disease classification
  │     └── leaf / lesion segmentation
  ├── Disease Diagnosis Skill
  │     └── literature retrieval
  ├── Disease Management Skill
  │     ├── hybrid RAG
  │     └── pesticide-use knowledge graph
  └── Report Generation Skill
        └── structured, traceable output
Visual Tools

EfficientNet-V2-S was selected as the deployed classifier after reaching 97.91% accuracy and 94.17% macro-F1 across 50 validation categories. SegFormer was selected for both binary segmentation tools, reaching 97.68% leaf mIoU and 92.98% lesion mIoU.

Representative leaf-disease images from ten crops
Figure 4: Representative leaf-disease images from the ten crops included in the benchmark.
End-to-End Results
MethodDiagnosis Acc. (%)CompletenessManagement GroundingOverall Quality
GPT-4o (direct)42.947.940.363.9
GPT-5.449.353.343.369.6
Claude Sonnet 550.049.042.667.4
Gemini 3.1 Pro Preview76.475.157.380.2
AgriSkill-Agent92.993.178.789.9
Ablation comparison
Figure 5: Paired report-quality differences on the same 40 image cases.
Hybrid retrieval benchmark
Figure 6: Retrieval results. The complete pipeline reached 80.0% Hit@1, 96.7% Hit@3, and 98.3% Hit@5.
Code & Artifacts

Code: github.com/Wispertise/AgriSkill-Agent

Code, benchmark manifests, prompts, and non-restricted experimental artifacts are available through the repository or from the corresponding authors.