arrow
Return

Machine learning methods for isolating indigenous language catalog descriptions

delete2025-02-24
delete0
PRE
AI
Y
Yi Liu
C
Carrie Heitman
L
Leen‐Kiat Soh
P
Peter M. Whiteley
DOI:10.1007/s00146-025-02223-ydelete
deleteOriginal
deleteOriginal request for help
deleteShare
deleteSave
Abstract

Abstract

En 中文
Museum collection databases contain echoes of encounter between colonial collectors (broadly defined) and Indigenous people from around the world. The moment of acquisition-when an item passed out of a community and into the hands of the collector-often included multilingual acts of translation. An artist may have shared the Indigenous name of the object, or the terms associated with its origin and use. Late nineteenth and twemtieth century museum registrars would in turn transcribe this information from field logs into museum catalogs. Over time, these catalog entries were transformed into digital records within collections managements systems (e.g., EMu, PastPerfect, etc.). As a result of this 150-year process, today's museum collection databases are riddled with Indigenous words and descriptions, scattered across various metadata fields. They may include Native place-names, family names or vocabulary terms that, when translated, extend far beyond the categories ascribed by museum collection managers. These instances of Indigenous description may also serve as a crucial bridge for reconnecting source communities with items of particular interest to their cultural heritage and linguistic preservation efforts. Aiming to enhance the accessibility of Indigenous languages contained in the metadata of cultural heritage collections, this paper explores applications of machine learning methodologies to identify Indigenous terms present in museum catalogs. Specifically, we discuss methods that incorporate the Google Cloud Language Identification Service to detect A:shiwi (Pueblo of Zuni) language terms through a case study of metadata records from the two largest natural history museums in the USA. We utilize an elimination mechanism to exclude specific languages (e.g., English and Spanish) at the word and phrase levels to detect A:shiwi terms. Our approach outperforms conventional methods in terms of accuracy, recall, precision, and F1-scores. This method can be used to confront the Digital Heap of cultural heritage records across institutions to improve the discoverability of Indigenous languages in metadata descriptions and reconnect source communities with items of cultural patrimony.
Keywords:
Language identification
Indigenous language identification
Natural language processing
Museum informatics
Cultural heritage
Metadata

Journal

AI and Society cover
AI and Society
IF:
4.7
Papers:
354
Citations:
4.8K

Organization

U
Univ Nebraska Lincoln
Scholars:
361
Papers: 187
Citations: 31
A
Amer Museum Nat Hist
Scholars:
89
Papers: 77
Citations: 34