Forests and plantations are critical for carbon sequestration, but accurately monitoring their growth has traditionally required expensive lidar surveys or labor-intensive field measurements. A new artificial intelligence model, described in a study published October 20, 2025, in the Journal of Remote Sensing (DOI: 10.34133/remotesensing.0880), demonstrates that sub-meter precision in estimating tree heights can be achieved using only standard RGB satellite imagery.
Developed by researchers from Beijing Forestry University, Manchester Metropolitan University, and Tsinghua University, the model combines a large vision foundation model with self-supervised learning to produce high-resolution canopy height maps. The approach addresses the longstanding challenge of balancing cost, precision, and scalability in forest monitoring, offering a practical tool for managing plantations and tracking carbon sequestration under programs such as China's Certified Emission Reduction initiative.
The researchers designed a canopy height estimation network with three components: a feature extractor based on the DINOv2 large vision foundation model, a self-supervised feature enhancement unit that preserves fine spatial details, and a lightweight convolutional height estimator. When tested against airborne lidar measurements, the model achieved a mean absolute error of just 0.09 meters and an R² of 0.78, outperforming traditional convolutional neural networks and transformer-based methods. It also enabled over 90% accuracy in single-tree detection and showed strong correlations with measured above-ground biomass.
Field tests were conducted in the Fangshan District of Beijing, an area with fragmented plantations of Populus tomentosa, Pinus tabulaeformis, and Ginkgo biloba. Using one-meter-resolution Google Earth imagery and lidar-derived references, the AI model produced canopy height maps that closely matched ground truth data. It significantly outperformed global canopy height model products, capturing subtle variations in tree crown structure that existing models often missed. The generated maps supported individual-tree segmentation and plantation-level biomass estimation with R² values exceeding 0.9 for key species. When applied to a geographically distinct forest in Saihanba, the network maintained robust accuracy, confirming its cross-regional adaptability.
“Our model demonstrates that large vision foundation models can fundamentally transform forestry monitoring,” said Dr. Xin Zhang, corresponding author at Manchester Metropolitan University. “By combining global image pretraining with local self-supervised enhancement, we achieved lidar-level precision using ordinary RGB imagery. This approach drastically reduces costs and expands access to accurate forest data for carbon accounting and environmental management.”
The team used an end-to-end deep-learning framework implementing pre-trained LVFM features with self-supervised enhancement. High-resolution Google Earth imagery from 2013 to 2020 served as input, while UAV-based lidar data provided reference for training and validation. The model was implemented in PyTorch using the fastai framework on an NVIDIA RTX A6000 GPU. Comparative experiments with conventional networks (U-Net and DPT) and global CHM datasets confirmed superior accuracy and efficiency.
The ability to reconstruct annual growth trends from archived satellite imagery provides a scalable solution for long-term carbon sink monitoring and precision forestry management. This innovation bridges the gap between expensive lidar surveys and low-resolution optical methods, enabling detailed forest assessment with minimal data requirements. Future research will extend the method to natural and mixed forests, integrate automated species classification, and support real-time carbon monitoring platforms. As the world advances toward net-zero goals, such intelligent, scalable mapping tools could play a central role in achieving sustainable forestry and climate-change mitigation.


