arrow
返回

Standardizing XBRL Tags with Natural Language Processing

delete2025-06-11
delete0
PRE
AI
R
Richard Wang *
DOI:10.1080/08874417.2025.2507710delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
This paper proposes a machine learning (ML) based approach to automatically standardize tens of thousands of custom XBRL tags by mapping them to the closest standard XBRL tags in the US Financial Reporting Taxonomy. This paper addresses a practical issue faced by many financial analysts, investors, regulators, and researchers when utilizing XBRL data for financial analysis and modeling. The algorithm utilizes three natural language processing (NLP) techniques: (1) Bag-Of-Words (TF-IDF), (2) Word2Vec Embedding, and (3) Huang et al.’s state-of-the-art FinBERT. The algorithm is implemented for all US custom tags between 2009 and 2022, and an evaluation of the results shows the algorithm is fast, efficient, and has a reasonable level of accuracy.
Keyword:
XBRL
natural language processing
financial reporting
FinBERT
standardization

期刊

Journal of Computer Information Systems 封面图
Journal of Computer Information Systems
IF:
4.2
论文数:
209
被引数:
3.1K

机构

S
st. john fisher university
学者数:
6
论文数: 4
被引数: 0
引用论文

引用论文

A survey of approaches to automatic schema matching
err2001-12-01
err0
errOAAI
errErhard Rahm; Philip A. Bernstein
err分享
err收藏
err分享
err收藏
Effects of the SEC's XBRL mandate on financial reporting comparability
err2015-12-01
err52
PREAI
errDhole, Sandip; Lobo, Gerald J.; Mishra, Sagarika; Pal, Ananda M.
err分享
err收藏
学者 查看更多内容