Vision Transformer-Based Recognition of Riau Malay Architectural Features with Cross-Regional Comparison for Digital Heritage Documentation

Authors

  • Heri Pramono Architecture, Star YKPN, Gagak Rimang No 1, Yogyakarta, Indonesia
  • Sri Winiarti Doctoral Program of Informatics, Universitas Ahmad Dahlan, Ringroad Selatan, Tamanan Bantul, Yogyakarta, Indonesia
  • Abdul Fadlil Doctoral Program of Informatics, Universitas Ahmad Dahlan, Ringroad Selatan, Tamanan Bantul, Yogyakarta, Indonesia
  • Sunardi Doctoral Program of Informatics, Universitas Ahmad Dahlan, Ringroad Selatan, Tamanan Bantul, Yogyakarta, Indonesia

DOI:

https://doi.org/10.24002/jarina.v5i2.14591

Keywords:

Architectural Feature Recognition, Digital Heritage Documentation, Pontianak Malay Architecture, Riau Malay Architecture, Vision Transformer

Abstract

Traditional Riau Malay architecture requires systematic digital documentation for heritage preservation. This study evaluates a Vision Transformer (ViT-B/16) model initialised with ImageNet-1K pretrained weights for recognising Riau Malay architectural features, using Pontianak Malay architecture for cross-regional comparison. The dataset was constructed from 24 architectural videos covering roof shapes, building structures, ornaments, windows, staircases, and full-building views. Using automated spatiotemporal segmentation at five frames per second, 13,230 frames were extracted, resized to 224×224 pixels, normalised, augmented, and divided into 16×16-pixel patches. Evaluation on a balanced, held-out test set of 32 clips yielded an overall accuracy of 84.38%, macro precision of 84.51%, macro recall of 84.38%, and macro F1-score of 84.36%. Distinctive elements, such as roofs, windows, staircases, and full buildings, achieved higher recognition performance when clearly visible. Conversely, partially visible structures and detailed ornaments exhibited variable performance due to lighting, viewpoint, and visual complexity. Given the single hold-out split and the limited number of source videos, these findings are preliminary; high feature-specific accuracies should not imply perfect recognition or generalizability. Nonetheless, the results demonstrate ViT-B/16’s strong potential to support the digital recognition of Malay architectural heritage. Future work should incorporate grouped five-fold cross-validation, independent building-level testing, CNN baseline comparisons, ROC–AUC analysis, and attention map visualisations.

References

[1] S. S. Workflow, R. Salah, N. Géczy, and K. A. Károlyfi, "Architectural Heritage Digitization : A Classification-Driven," Buildings, vol. 16, no. 21, pp. 1-24, 2026, https://doi.org/10.3390/buildings16010021

[2] F. Colace, R. Gaeta, A. Lorusso, M. Pellegrino, and D. Santaniello, "New AI challenges for cultural heritage protection : A general overview," J. Cult. Herit., vol. 75, pp. 168-193, 2025, https://doi.org/10.1016/j.culher.2025.07.019

[3] X. Li and F. Chiabrando, "Machine Learning and Deep Learning for Cultural Heritage Conservation : A Bibliometric and Task-Oriented Review," Remote Sens., vol. 18, no. 268, 2026, https://doi.org/10.3390/rs18040628

[4] S. Alsheikh Mahmoud, H. Bin Hashim, M. F. Shamsudin, and H. Alsheikh Mahmoud, "Effective Preservation of Traditional Malay Houses: A Review of Current Practices and Challenges," Sustain. , vol. 16, no. 11, p. 4773, 2024, https://doi.org/10.3390/su16114773

[5] K. Siountri and C. N. Anagnostopoulos, "The Classification of Cultural Heritage Buildings in Athens Using Deep Learning Techniques," Heritage, vol. 6, no. 4, pp. 3673-3705, 2023, https://doi.org/10.3390/heritage6040195

[6] A. Sasithradevi, B. Chanthini, T. Subbulakshmi, and P. Prakash, "MonuNet : a high performance deep learning network for Kolkata heritage image classification," Herit. Sci., pp. 1-14, 2024, https://doi.org/10.1186/s40494-024-01340-z

[7] A. Luiz and C. Ottoni, "ImageOP : The Image Dataset with Religious Buildings in the World Heritage Town of Ouro Preto for Deep Learning Classification," Heritage, vol. 7, no. 11, pp. 6499-6525, 2024, https://doi.org/10.3390/heritage7110302

[8] Z. Li et al., "A deep learning-based method for identifying traditional villages across cultural types," Herit. Sci., vol. 13, no. 671, pp. 1-15, 2025, https://doi.org/10.1038/s40494-025-02187-8

[9] T. Luo, X. Sun, W. Zhao, W. Li, L. Yin, and D. Xie, "Ethnic Architectural Heritage Identification Using Low-Altitude UAV Remote Sensing and Improved Deep Learning Algorithms," Buildings, vol. 15, no. 15, 2025, https://doi.org/10.3390/buildings15010015

[10] S. Wang, J. Zhang, A. N. Tun, and K. Sein, "Research on Identification, Evaluation, and Digitization of Historical Buildings Based on Deep Learning Algorithms: A Case Study of Quanzhou World Cultural Heritage Site," Buildings, vol. 15, no. 11, pp. 1-18, 2025, https://doi.org/10.3390/buildings15111843

[11] P. Han, S. Hu, and R. Xu, "Formal Feature Identification of Vernacular Architecture Based on Deep Learning-A Case Study of Jiangsu Province, China," Sustain., vol. 17, no. 4, pp. 1-30, 2025, https://doi.org/10.3390/su17041760

[12] J. Zhou, L. Xie, P. Fricker, and K. Liu, "ConvNeXt-L-Based Recognition of Decorative Patterns in Historical Architecture : A Case Study of Macau," Buildings, vol. 15, no. 20, pp. 1-27, 2025, https://doi.org/10.3390/buildings15203705

[13] S. Alsheikh Mahmoud, H. Bin Hashim, M. F. Shamsudin, and H. Alsheikh Mahmoud, "Effective Preservation of Traditional Malay Houses: A Review of Current Practices and Challenges," Sustain. , vol. 16, no. 11, pp. 1-19, 2024, https://doi.org/10.3390/su16114773

[14] Y. Wang, Y. Deng, Y. Zheng, P. Chattopadhyay, and L. Wang, "Vision Transformers for Image Classification: A Comparative Survey," Technologies, vol. 13, no. 1, pp. 1-32, 2025, https://doi.org/10.3390/technologies13010032

[15] J. Montrezol, H. S. Oliveira, and H. P. Oliveira, "Decoding Vision Transformer Variations for Image Classification: A Guide to Performance and Usability," Machine Learning with Applications, vol. 23, p. 100844, Mar. 2026, https://doi.org/10.1016/j.mlwa.2026.100844

[16] W. Xu et al., "ViT-HVE : a vision transformer-based framework for recognition and weighted evaluation of cultural heritage values," Herit. Sci., vol. 13, no. 571, 2025, https://doi.org/10.1038/s40494-025-02143-6

[17] T. Fan and H. Wang, "Multimodal cultural heritage image recognition based on quantum and classical multimodal fusion network," Herit. Sci., vol. 14, p. 160, 2026, doi: https://doi.org/10.1038/s40494-026-02419-5

[18] S. Khan, M. Naseer, M. Hayat, and S. W. Zamir, "Transformers in Vision : A Survey," ACM Comput. Surv., pp. 1-30, 2021, doi: https://doi.org/10.1145/3505244

[19] T. Luo, X. Sun, W. Zhao, W. Li, L. Yin, and D. Xie, "Ethnic Architectural Heritage Identification Using Low-Altitude UAV Remote Sensing and Improved Deep Learning Algorithms," Buildings, vol. 15, no. 1, 2025, https://doi.org/10.3390/buildings15010015

[20] F. N. U. Neha and A. Bansal, "Understanding the architecture of vision transformer and its variants: A review," 1st Int. Conf. Innov. Eng. Sci. Technol. Res. ICIESTR 2024 - Proc., no. May, 2024, https://doi.org/10.1109/ICIESTR60916.2024.10798341

[21] A. Gil, Y. Arayici, B. Kumar, and R. Laing, "Machine and Deep Learning Implementations for Heritage Building Information Modelling : A Critical Review of Theoretical and Applied," ACM J. Comput. Cult. Herit., vol. 17, no. 3, 2024, https://doi.org/10.1145/3649442

[22] A. Reshetnikov et al., "DEArt : Building and evaluating a dataset for object detection and pose classification for European art," J. Cult. Herit., vol. 75, pp. 258-266, 2025, https://doi.org/10.1016/j.culher.2025.07.022

[23] A. M. Tasir, M. Nor, and A. Khalid, "OPEN A novel deep neural model for efficient and scalable historical place image classification," Sci. Rep., vol. 15, no. 42745, pp. 1-29, 2025, https://doi.org/10.1038/s41598-025-26897-y

[24] A. M. Shehata and N. S. Alaboud, "Identification of Features of Architectural Heritage Using Deep Learning Techniques," Civ. Eng. Archit., vol. 13, no. 4, pp. 3280-3298, 2025, https://doi.org/10.13189/cea.2025.130431

[25] D. W. Alexey Dosovitskiy, Lucas Beyer∗, Alexander Kolesnikov∗, M. M. Xiaohua Zhai∗, Thomas Unterthiner, Mostafa Dehghani, and N. H. Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale," in Published as a conference paper at ICLR 2021, 2021.

[26] A. Dosovitskiy et al., "An image is worth 16 x 16 words :," Int. Conf. Learn. Represent., pp. 1-21, 2021, https://doi.org/10.48550/arXiv.2010.11929.

[27] A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lučić, and C. Schmid, "ViViT: A Video Vision Transformer," Proc. IEEE Int. Conf. Comput. Vis., pp. 6816-6826, 2021, https://doi.org/10.1109/ICCV48922.2021.00676

[28] T. Zhang, W. Xu, B. Luo, and G. Wang, "Depth-Wise Convolutions in Vision Transformers for Efficient Training on Small Datasets," Neurocomputing, vol. 617, p. 128998, Nov. 2024, https://doi.org/10.1016/j.neucom.2024.128998

[29] E. Matsuyama, H. Watanabe, and N. Takahashi, "Performance Comparison of Vision Transformer- and CNN-Based Image Classification Using Cross Entropy : A Preliminary Application to Lung Cancer Discrimination from CT Images," J. Biomed. Sci. Eng., vol. 17, no. 9, pp. 157-170, 2024, https://doi.org/10.4236/jbise.2024.179012

Downloads

Published

2026-08-18

How to Cite

[1]
H. Pramono, S. Winiarti, A. Fadlil, and Sunardi, “Vision Transformer-Based Recognition of Riau Malay Architectural Features with Cross-Regional Comparison for Digital Heritage Documentation”, JARINA, vol. 5, no. 2, pp. 108–125, Aug. 2026.