<p>This data uses 2023 Sentinel-2 optical satellite data as the data source, and a sampling random forest regression (RFR) method is used to build a remote sensing estimation model of suspended matter concentration in Xinjiang lakes. The data is Abers projection in the CGCS2000 coordinate system with an accuracy of 10 meters. The numerical coefficient is 0.001, that is, the data pixel value is multiplied by 0.001 to obtain the actual suspended solids concentration in the water body (mg/L). </p>
| data size | 2.2 GiB |
|---|---|
| data format | grid |
| Coordinate system | CGCS2000 |
| Projection | Albers projection |
Use 2023 Sentinel 2 optical satellite data as the data source.
Use Sentinel 2A/BMSI with better radiation performance. Sentinel2A/BMSI has 13 spectral bands with spatial resolutions of 10m, 20m and 60m respectively. The revisit period is 10 days and 5 days after the double satellite network. It has demonstrated good performance in inland water bodies. Obtained L1C data from ESA's Copernicus Data Open Center. Atmospheric correction is carried out using the SEN2COR algorithm provided by ESA. However, SEN2COR is a terrestrial atmosphere correction algorithm that needs to be aimed at further removing the effects of sky light, solar flares and residual aerosol scattering: R_rs(λ)=(R(λ)-min (R(865),R(2202)))/π Where, Rrs is the remote sensing reflectance of water body, and R is the surface reflectance obtained by SEN2COR correction. A sampling random forest regression (RFR) method was used to build a water quality parameter model. Random forest is an integrated learning method whose basic unit is a decision tree. During the training process, a large number of independent decision trees are built to form a "random forest", and finally the results of these decision trees are synthesized to improve the accuracy of the model (such as output average). The "randomness" of the RFR algorithm is mainly reflected in two aspects: when building each decision tree, the Bagging method (i.e., the self-service sampling method, where samples are put back after each sampling) is randomly sampled from the original training data set to generate a training data subset; When nodes are split, all feature variables are not used to participate in the comparison, but a subset of the feature variables is randomly selected to participate in node splitting. Two "randomness" make the RFR algorithm less prone to overfitting and has good tolerance for outliers and noise. The most important hyperparameters of the RFR algorithm during the training process are: (1) The number of decision trees (n_estimators). The larger the n_estimators, the better the model results, but the longer the calculation time is; when a certain amount of data is reached, the model tends to be stable. (2) The maximum number of features (max_features) when the node is split. The decision tree finds the best split feature from randomly selected max_features when the node is split. (3) The maximum depth of the decision tree (max_depth). If it is not set (i.e. None), the decision tree will grow to the maximum extent until the segmentation termination condition is met. RFR is implemented through the Pythonscikit-learn software package. By adjusting the optimal input variables of the input algorithm, the optimal hyperparameters of each algorithm are obtained through the grid search method. The algorithm based on RFR-has high accuracy: the uncertainty is 24.29%, the deviation (β) is 6.77%, the slope is 0.90, and the root mean square logarithmic error is 0.219.
A sampling random forest regression (RFR) method was used to build a water quality parameter model. Random forest is an integrated learning method whose basic unit is a decision tree. During the training process, a large number of independent decision trees are built to form a "random forest", and finally the results of these decision trees are synthesized to improve the accuracy of the model (such as output average). The optimal input variables for each water quality parameter algorithm are determined by adjusting the inputs, and the optimal hyperparameters of each algorithm are obtained through a grid search method. A large number of field surveys and satellite-earth synchronous data were used to conduct model research. The results showed that the RFR suspended matter concentration algorithm had high accuracy, with an uncertainty of 24.29%, a deviation of 6.77%, a slope of 0.90, and a root-mean-square logarithmic error of 0.219.
| # | number | name | type |
| 1 | 2021xjkk1400 | 2021xjkk1400 | National Science and technology support program |
This work is licensed under a
Creative
Commons Attribution 4.0 International License.
| # | title | file size |
|---|---|---|
| 1 | tetx.txt | 0 Bytes |
| 2 | 2021xjkk1400-24-2023121324 |
©Copyright 2021-. Xinjiang Institute of Ecology and Geography, CAS
No. 818 Beijing South Road, Urumqi, Xinjiang, China, 830011
