The deployment of robust computer vision models in precision agriculture is frequently limited by the scarcityand high cost of acquiring large, accurately annotated datasets in unstructured orchard environments. This studyevaluates photorealistic synthetic imagery as a scalable alternative and complement to real data for apple detection. We first performed a systematic benchmarking campaign on publicly available apple-detection datasets,identifying the specific characteristics that drive model performance. Guided by these insights, we developed afully procedural generation pipeline (Blender + Unreal Engine 5) to create three synthetic datasets with progressively refined scene composition and annotation geometry. Our evaluation demonstrates that models trainedon high-quality synthetic data achieve cross-domain generalization comparable to models trained on establishedreal-world datasets. Notably, we found that the domain gap between synthetic and real imagery could be substantially bridged by a mixed-data strategy. Introducing a small fraction of real imagery into a predominantlysynthetic training set corrected low-precision behaviours and yielded marked improvements in detection performance, with diminishing returns observed at higher ratios of real data. For the apple-detection task and rendering pipeline studied here, these results confirm that synthetic datasets provide a cost-effective, scalable trainingbackbone for agricultural object detection, reducing manual data collection and labelling while maintaining highdetection accuracy.

Designing effective synthetic datasets for fruit detection in agriculture

Hueller, Jhonny;Vecchio, Massimo
;
Antonelli, Fabio
2026-01-01

Abstract

The deployment of robust computer vision models in precision agriculture is frequently limited by the scarcityand high cost of acquiring large, accurately annotated datasets in unstructured orchard environments. This studyevaluates photorealistic synthetic imagery as a scalable alternative and complement to real data for apple detection. We first performed a systematic benchmarking campaign on publicly available apple-detection datasets,identifying the specific characteristics that drive model performance. Guided by these insights, we developed afully procedural generation pipeline (Blender + Unreal Engine 5) to create three synthetic datasets with progressively refined scene composition and annotation geometry. Our evaluation demonstrates that models trainedon high-quality synthetic data achieve cross-domain generalization comparable to models trained on establishedreal-world datasets. Notably, we found that the domain gap between synthetic and real imagery could be substantially bridged by a mixed-data strategy. Introducing a small fraction of real imagery into a predominantlysynthetic training set corrected low-precision behaviours and yielded marked improvements in detection performance, with diminishing returns observed at higher ratios of real data. For the apple-detection task and rendering pipeline studied here, these results confirm that synthetic datasets provide a cost-effective, scalable trainingbackbone for agricultural object detection, reducing manual data collection and labelling while maintaining highdetection accuracy.
File in questo prodotto:
Non ci sono file associati a questo prodotto.

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/11582/374390
 Attenzione

Attenzione! I dati visualizzati non sono stati sottoposti a validazione da parte dell'ateneo

Citazioni
  • ???jsp.display-item.citation.pmc??? ND
social impact