返回
Standardizing XBRL Tags with Natural Language Processing
DOI:10.1080/08874417.2025.2507710.png)
摘要
En 中文
This paper proposes a machine learning (ML) based approach to automatically standardize tens of thousands of custom XBRL tags by mapping them to the closest standard XBRL tags in the US Financial Reporting Taxonomy. This paper addresses a practical issue faced by many financial analysts, investors, regulators, and researchers when utilizing XBRL data for financial analysis and modeling. The algorithm utilizes three natural language processing (NLP) techniques: (1) Bag-Of-Words (TF-IDF), (2) Word2Vec Embedding, and (3) Huang et al.’s state-of-the-art FinBERT. The algorithm is implemented for all US custom tags between 2009 and 2022, and an evaluation of the results shows the algorithm is fast, efficient, and has a reasonable level of accuracy.
Keyword:
XBRL
natural language processing
financial reporting
FinBERT
standardization
期刊
IF:
4.2
论文数:
209
被引数:
3.1K
机构
引用论文
The determinants of eXtensible Business Reporting Language (XBRL) adoption: a cross-country study可扩展商业报告语言(XBRL)采纳的决定因素:一项跨国研究
Customization versus Standardization in Electronic Financial Reporting: Early Evidence from the SEC XBRL Mandate电子金融报告中的定制化与标准化:来自SEC XBRL强制规定的早期证据

