Geometric evaluation in crack segmentation: an analysis of the complementary relationship between region-based (IoU) and boundary-based (BF1) metrics

  • Affiliations:

    1 University of Transport and Communications, Hanoi, Vietnam
    2 Hanoi University of Mining and Geology, Hanoi, Vietnam

  • *Corresponding:
    This email address is being protected from spambots. You need JavaScript enabled to view it.
  • Received: 10th-Mar-2026
  • Revised: 2nd-Sept-2026
  • Accepted: 12th-Sept-2026
  • Online: 1st-Oct-2026
Pages: 113 - 125
Views: 25
Downloads: 1
★★★★★Rating: , Total rating: 0
  • ★
  • ★
  • ★
  • ★
  • ★
Yours rating

Abstract:

Crack segmentation from images plays a crucial role in infrastructure inspection and maintenance, where the reliability of predicted masks directly affects the assessment of structural deterioration. Although the Intersection over Union (IoU) metric is widely adopted for segmentation evaluation, it primarily measures region overlap and may not adequately capture geometric boundary accuracy for thin, elongated structures such as cracks. This study investigates the complementary relationship between the region-based IoU metric and the boundary-based Boundary F1-score (BF1), with the aim of clarifying their respective strengths and limitations in crack segmentation assessment. Experiments were conducted using a U-Net model with a ResNet-34 encoder on two datasets with distinct labeling characteristics, including a multi-source ground-image dataset and a consistently annotated UAV dataset. In addition, a controlled boundary-deviation experiment was designed to directly evaluate the response of both metrics to predefined geometric distortions. The results indicate that IoU tends to increase withcrack width and mainly reflects area overlap, whereas BF1 is more sensitive to boundary localization errors and annotation quality. For wide cracks, IoU achieved similar values across both datasets (0.706), while BF1 clearly discriminated data quality, increasing from 0.273 on the multi-source dataset to 0.875 on the UAV dataset. The controlled experiments further demonstrated that BF1 decreases onceboundary deviations exceed the tolerance threshold θ, whereas IoU degrades more gradually with increasing geometric distortion. These findings demonstrate that IoU and BF1 characterize different geometric aspects of segmentation quality and therefore should not be used interchangeably. Based on this evidence, the study recommends a standardized evaluation framework that reports both region-based and boundary-based metrics to provide a more reliable assessment of thin linear structure segmentation.

How to Cite
Vu, P.Ngoc and Pham, D.Trung 2026. Geometric evaluation in crack segmentation: an analysis of the complementary relationship between region-based (IoU) and boundary-based (BF1) metrics (in Vietnamese). Journal of Mining and Earth Sciences. 67, 5 (Oct, 2026), 113-125. DOI:https://doi.org/10.46326/JMES.2026.67(5).09.
References

Al-Huda, Z., Peng, B., Algburi, R. N. A., Al-Antari, M., Al-Jarazi, R., and Zhai, D. (2023). A hybrid deep learning pavement crack semantic segmentation. Eng. Appl. Artif. Intell., 122, 106142. https://doi.org/10.1016/j.engappai. 2023.106142.

Bộ Khoa học và Công nghệ. (2024). Bảo dưỡng thường xuyên đường bộ - Yêu cầu kỹ thuật (TCVN 14182:2024). Bộ Khoa học và Công nghệ. https://tieuchuan.vsqi.gov.vn/tieu chuan/view?sohieu=TCVN+14182%3A2024.

Chen, Z., Shamsabadi, E. A., Jiang, S., Shen, L., and Dias-Da-Costa, D. (2024). Vision Mamba-based autonomous crack segmentation on concrete, asphalt, and masonry surfaces. ArXiv, abs/2406.16518. https://doi.org/10.48550/ arxiv.2406.16518.

Cheng, B., Girshick, R., Dollár, P., Berg, A. C., and Kirillov, A. (2021). Boundary IoU: Improving Object-Centric Image Segmentation Evaluation (arXiv:2103.16562). arXiv. https://doi.org/ 10.48550/arXiv.2103.16562.

Csurka, G., Larlus, D., and Perronnin, F. (2013). What is a good evaluation measure for semantic segmentation? Proceedings of the British Machine Vision Conference 2013, 32.1-32.11. https://doi.org/10.5244/C.27.32.

Du, Z., Yuan, J., Xiao, F., and Hettiarachchi, C. (2021). Application of image technology on pavement distress detection: A review. Measurement, 184, 109900. https://doi.org/10.1016/ j.measurement.2021.109900.

Dung, C. V, Anh, L. D. (2019). Autonomous concrete crack detection using deep fully convolutional neural network. Automation in Construction, 99, 52-58. https://doi.org/10. 1016/j.autcon.2018.11.028.

Fan, Z., Li, C., Chen, Y., Mascio, P. D., Chen, X., Zhu, G., and Loprencipe, G. (2020). Ensemble of Deep Convolutional Neural Networks for Automatic Pavement Crack Detection and Measurement. Coatings, 10(2), 152. https://doi.org/10.3390/ coatings10020152.

Fei, Y., Wang, K. C. P., Zhang, A., Chen, C., Li, J. Q., Liu, Y., Yang, G., and Li, B. (2020). Pixel-Level Cracking Detection on 3D Asphalt Pavement Images Through Deep-Learning- Based CrackNet-V. IEEE Transactions on Intelligent Transportation Systems, 21(1), 273-284. https://doi.org/10.1109/TITS.2019.2891167.

Goo, J. M., Milidonis, X., Artusi, A., Boehm, J., and Ciliberto, C. (2025). Hybrid-Segmentor: Hybrid approach for automated fine-grained crack segmentation in civil infrastructure. Automation in Construction, 170, 105960. https://doi.org/10.1016/j.autcon.2024.105960.

He, K., Zhang, X., Ren, S., and Sun, J. (2016). Deep Residual Learning for Image Recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770-778. https://doi.org/10.1109/CVPR.2016.90.

Kang, D., and Cha, Y. (2022). Efficient attention-based deep encoder and decoder for automatic crack segmentation. Structural Health Monitoring, 21(5), 2190-2205. https://doi. org/10.1177/14759217211053776.

Kheradmandi, N., and Mehranfar, V. (2022). A critical review and comparative study on image segmentation-based techniques for pavement crack detection. Construction and Building Materials, 321, 126162. https://doi. org/10.1016/j.conbuildmat.2021.126162.

Korcynski, N. (2026). A Boundary-Metric Evaluation Protocol for Whiteboard Stroke Segmentation Under Extreme Imbalance (arXiv:2603.00163). arXiv. https://doi.org/ 10.48550/arXiv.2603.00163.

Martin, D. R., Fowlkes, C. C., and Malik, J. (2004). Learning to detect natural image boundaries using local brightness, color, and texture cues. IEEE Transactions on Pattern Analysis and Machine Intelligence, 26(5), 530-549. https:// doi.org/10.1109/TPAMI.2004.1273918.

Middha, L. (2020). Crack Segmentation Dataset [Data set]. https://www.kaggle.com/datasets /lakshaymiddha/crack-segmentation-dataset.

Miller,J.S and Bellinger,W.Y. (2014). Distress Identification Manual for the Long-Term Pavement Performance Program. (5th rev.ed.; Report No. FHWA-HRT-13-092.)Federal Highway Administration. https://www.fhwa. dot.gov/publications/research/infrastructure/pavements/ltpp/13092/13092.pdf.

Rahman, M. A., and Wang, Y. (2016). Optimizing Intersection-Over-Union in Deep Neural Networks for Image Segmentation. In G. Bebis, R. Boyle, B. Parvin, D. Koracin, F. Porikli, S. Skaff, A. Entezari, J. Min, D. Iwai, A. Sadagic, C. Scheidegger, and T. Isenberg (Eds.), Advances in Visual Computing (pp. 234-244). Springer International Publishing. https://doi.org/ 10.1007/978-3-319-50835-1_22.

Ronneberger, O., Fischer, P., and Brox, T. (2015). U-Net: Convolutional Networks for Biomedical Image Segmentation. In N. Navab, J. Hornegger, W. M. Wells, and A. F. Frangi (Eds.), Medical Image Computing and Computer-Assisted Intervention - MICCAI 2015 (pp. 234-241). Springer International Publishing. https://doi. org/10.1007/978-3-319-24574-4_28.

Shit, S., Paetzold, J. C., Sekuboyina, A., Ezhov, I., Unger, A., Zhylka, A., Pluim, J. P. W., Bauer, U., and Menze, B. H. (2021). clDice-A Novel Topology-Preserving Loss Function for Tubular Structure Segmentation. 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 16560-16569. https:// doi.org/10.1109/CVPR46437.2021.01629.

Sudre, C. H., Li, W., Vercauteren, T., Ourselin, S., and Cardoso, M. J. (2017). Generalised Dice overlap as a deep learning loss function for highly unbalanced segmentations. In Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Support (Vol. 10553, pp. 240-248). https://doi.org/ 10.1007/978-3-319-67558-9_28 (Original work published In: DLMIA Workshop, MICCAI.).

Wang, Z., and Pang, J. (2025). Fast Measuring Pavement Crack Width by Cascading Principal Component Analysis (arXiv:2511.02144). arXiv. https://doi.org/10.48550/arXiv.2511. 02144.

Yu, G., Dong, J., Wang, Y., and Zhou, X. (2023). RUC-Net: A Residual-Unet-Based Convolutional Neural Network for Pixel-Level Pavement Crack Segmentation. Sensors, 23(1), 53. https://doi.org/10.3390/s23010053

Ziya. (2025). UAV-Based Crack Detection Dataset [Data set]. https://www.kaggle.com/ datasets/ziya07/uav-based-crack-detection-dataset.

Other articles