【目的】杉木是我国最重要的经济树种之一,在生态防护、木材供给等方面具有不可替代的价值。杉木苗期三维(3D)表型参数是反映其生长状态的重要根据,从图像、三维点云中提取单株信息是执行后续监测分析的第一步。新兴的立体视觉与人工智能(artificial intelligence, AI)技术给植物表型分析领域提供了高效、自动化的数据分析工具,但常规的模型训练依赖大批量精细的人工标注进行全监督学习,且三维成像设备价格昂贵、点云处理效率低,制约了其在林木信息感知方面的大规模运用。为了突破常规AI模型训练依赖大批量精细的人工标注、三维成像设备价格昂贵、点云处理效率低等局限性,提出融合零标注/少标注学习和双目视觉的杉木图像-点云分割算法。【方法】本研究提出一种融合零标注/少标注学习和双目视觉的杉木图像-点云分割方法,依据杉木季节性色差选用差异化提示文本,辅助SAM3视觉语言模型生成高质量二维(2D)实例掩码;将SAM3输出作为伪标签,训练基于YOLO网络的轻量化杉木图像分割模型;运用双目相机提供的2D-3D坐标映射,将二维掩码投影至三维空间实现单株杉木点云的快速提取并计算株高和冠幅。【结果】野外种植的杉木苗叶色随季节变化明显,采集的图像含复杂背景,SAM3在针对性的提示词辅助下可输出更高质量的伪标签,与人工真实标注的平均交并比最高可达0.819。零标注模式下,利用完整数据集的原始伪标签训练YOLOv8n模型,在测试集上的平均精度(mAP50)为0.851;仅处理地面平台拍摄植株的子数据集时,mAP50可达0.926。少标注模式下,人工增补少量伪标签中缺失个体再训练模型,在处理完整数据集、子数据集的mAP50分别提升至0.925、0.995,同步实现模型低复杂度、高精度。通过简洁高效的二维到三维的投影映射分割取得单株点云,株高和冠幅计算的平均百分比偏差分别为5.78%、11.70%。【结论】本研究提出的杉木图像点云分割方法简化点云分割计算流程,降低智能模型训练对人工精细标注的依赖,可为后续优株筛选、胁迫监测和移栽管理中的表型数据获取提供方法参考。
【Objective】Cunninghamia lanceolata (Chinese fir) is one of the important economic tree species in China, possessing irreplaceable ecological, economic, cultural, and medicinal value. The 3D phenotypic parameters of Chinese fir seedlings serve as a crucial basis for evaluating their growth status. Precisely extracting individual plant information from multimodal images is a prerequisite for subsequent single-tree phenotyping. Although existing stereo vision and artificial intelligence (AI) algorithms have been widely applied to various plant phenotyping tasks, conventional fully supervised training and modeling methods heavily rely on large-scale, fine-grained manual annotations. Furthermore, the relatively high cost of 3D imaging equipment and the time-consuming nature of point cloud analysis restrict their application and development in forestry engineering. Emerging stereo vision and artificial intelligence (AI) technologies provide efficient and automated data analysis tools for plant phenotyping. To overcome the limitations including the heavy reliance on large-scale, manually annotated datasets for AI training, the high cost of 3D imaging equipment, and the low efficiency of point cloud processing, this study proposed a joint image-point cloud segmentation method for Chinese fir by integrating few- or zero-shot deep learning and binocular vision technologies.【Method】This study proposed a joint image-point cloud segmentation method for Chinese fir by integrating few-/zero-shot deep learning and binocular vision technologies. A differentiated prompting mechanism, tailored to the seasonal color variations of Chinese fir, was constructed to assist the vision-language model, SAM3, in generating high-quality 2D instance masks. A pseudo-label-based learning process was designed, which utilized the SAM3 outputs as pseudo-labels for training a lightweight YOLO instance segmentation network. Utilizing the 2D-3D coordinate mapping, the 2D masks were projected into 3D space to rapidly extract point clouds of single plant and calculate its height and crown width.【Result】The leaf color of field-grown Chinese fir seedlings changed significantly with seasons. The captured plant images had complex backgrounds. With proper text prompts, SAM3 could generate higher-quality pseudo-labels, achieving a maximum mean Intersection over Union (mIoU) of 0.819 compared to ground-truth annotations. Based on zero-shot mode, the YOLOv8n segmentation model trained with these pseudo-labels of the full dataset achieved an mAP50 of 0.851 on the test set. When processing the sub-set including only plants captured from ground-level perspective, the mAP50 was 0.926. Based on few-shot mode, some missing objects are corrected manually in the pseudo-labels. Training using the modified pseudo-labels, the mAP50 values concerning the full- and sub-set were improved to 0.925, 0.995, respectively. The YOLOv8n-based plant segmentation model had a weight size of only 6.46 MB, which successfully balanced low model complexity with high precision. Through concise and efficient 2D-to-3D projection mapping for segmentation, point cloud of single plant was obtained. It achieved average relative errors of 5.78% for tree height calculation and 11.70% for crown width calculation, respectively.【Conclusion】The proposed image-point cloud segmentation method for Chinese fir simplifies the 3D segmentation workflow and reduces the reliance on manual annotations for intelligent model training. It provides methodological references for obtaining plant phenotypic traits in downstream applications such as superior plants screening, stress monitoring, and transplantation management.