Please use this identifier to cite or link to this item:
Title: Harnessing Large-Scale Herbarium Image Datasets Through Representation Learning
Authors: Walker, BE
Tucker, A
Nicolson, N
Keywords: deep learning;digitized herbarium specimens;natural history collections;machine learning;computer vision
Issue Date: 13-Jan-2022
Publisher: Frontiers Media SA
Citation: Walker, B.E., Tucker, A. and Nicolson, N. (2022) 'Harnessing Large-Scale Herbarium Image Datasets Through Representation Learning', Frontiers in Plant Science, 12, 806407, pp. 1-12. doi: 10.3389/fpls.2021.806407.
Abstract: Copyright © 2022 Walker, Tucker and Nicolson. The mobilization of large-scale datasets of specimen images and metadata through herbarium digitization provide a rich environment for the application and development of machine learning techniques. However, limited access to computational resources and uneven progress in digitization, especially for small herbaria, still present barriers to the wide adoption of these new technologies. Using deep learning to extract representations of herbarium specimens useful for a wide variety of applications, so-called “representation learning,” could help remove these barriers. Despite its recent popularity for camera trap and natural world images, representation learning is not yet as popular for herbarium specimen images. We investigated the potential of representation learning with specimen images by building three neural networks using a publicly available dataset of over 2 million specimen images spanning multiple continents and institutions. We compared the extracted representations and tested their performance in application tasks relevant to research carried out with herbarium specimens. We found a triplet network, a type of neural network that learns distances between images, produced representations that transferred the best across all applications investigated. Our results demonstrate that it is possible to learn representations of specimen images useful in different applications, and we identify some further steps that we believe are necessary for representation learning to harness the rich information held in the worlds’ herbaria.
Appears in Collections:Dept of Computer Science Research Papers

Files in This Item:
File Description SizeFormat 
FullText.pdf10.08 MBAdobe PDFView/Open

This item is licensed under a Creative Commons License Creative Commons