arrow
Return

PhytoOracle: Scalable, modular phenomics data processing pipelines

delete2023-03-06
delete4
delete
OA
AI
E
Emmanuel Gonzalez
A
Ariyan Zarei
N
Nathanial Hendler
T
Travis Simmons
A
Arman Zarei
J
Jeffrey Demieville
B
Bruno Rozzi
S
Sebastian Calleja
H
Holly Ellingson
M
Michele Cosi
S
Sean Davey
D
Dean Lavelle
M
María José Truco
T
Tyson L. Swetnam
N
Nirav Merchant
R
Richard W. Michelmore
E
Eric Lyons
D
Duke Pauli *
DOI:10.3389/fpls.2023.1112973delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
As phenomics data volume and dimensionality increase due to advancements in sensor technology, there is an urgent need to develop and implement scalable data processing pipelines. Current phenomics data processing pipelines lack modularity, extensibility, and processing distribution across sensor modalities and phenotyping platforms. To address these challenges, we developed PhytoOracle (PO), a suite of modular, scalable pipelines for processing large volumes of field phenomics RGB, thermal, PSII chlorophyll fluorescence 2D images, and 3D point clouds. PhytoOracle aims to (i) improve data processing efficiency; (ii) provide an extensible, reproducible computing framework; and (iii) enable data fusion of multi-modal phenomics data. PhytoOracle integrates open-source distributed computing frameworks for parallel processing on high-performance computing, cloud, and local computing environments. Each pipeline component is available as a standalone container, providing transferability, extensibility, and reproducibility. The PO pipeline extracts and associates individual plant traits across sensor modalities and collection time points, representing a unique multi-system approach to addressing the genotype-phenotype gap. To date, PO supports lettuce and sorghum phenotypic trait extraction, with a goal of widening the range of supported species in the future. At the maximum number of cores tested in this study (1,024 cores), PO processing times were: 235 minutes for 9,270 RGB images (140.7 GB), 235 minutes for 9,270 thermal images (5.4 GB), and 13 minutes for 39,678 PSII images (86.2 GB). These processing times represent end-to-end processing, from raw data to fully processed numerical phenotypic trait data. Repeatability values of 0.39-0.95 (bounding area), 0.81-0.95 (axis-aligned bounding volume), 0.79-0.94 (oriented bounding volume), 0.83-0.95 (plant height), and 0.81-0.95 (number of points) were observed in Field Scanalyzer data. We also show the ability of PO to process drone data with a repeatability of 0.55-0.95 (bounding area).
Keywords:
phenomics
morphological phenotyping
physiological phenotyping
distributed computing
high performance computing
image analysis
point cloud analysis
data management
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

Frontiers in Plant Science cover
Frontiers in Plant Science
IF:
4.8
Papers:
3.4W
Citations:
14.7W

Organization

S
Sharif University of Technology
Scholars:
1.1W
Papers: 1.1W
Citations: 9.5K
U
University of Arizona
Scholars:
3.6W
Papers: 3.2W
Citations: 980
University of California System cover
University of California System
Scholars:
37.5W
Papers: 33.7W
Citations: 6.6K
researcher View more organizations