arrow
Return

HTMLDownloader: An open-source tool for dynamic web scraping and archiving using WebView2

delete2025-09-01
delete0
delete
OA
AI
T
Truong, Ba-Vinh
N
Nguyen, Loan T. T.
P
Pham, Phu
B
Bay Vo *
DOI:10.1016/j.softx.2025.102373delete
deleteOriginal
deleteShare
deleteSave
View PDF
Abstract

Abstract

En 中文
The increasing complexity and dynamism of modern websites present major challenges for traditional web scraping tools such as Scrapy, BeautifulSoup and wget, which often fail to capture dynamic content or offer accessible user interfaces. To address these limitations, we introduce HTMLDownloader which provides a graphical interface (GUI) that makes it accessible to non-technical users and enhances scraping reliability by integrating browser-based rendering. Experimental evaluations on 75,516 links from 35 diverse domains demonstrate a 98.4% success rate, significantly outperforming Selenium (81.3%), Scrapy (56.6%), BeautifulSoup (34.5%) and wget (45.1%). These results confirm HTMLDownloader's robustness and scalability, making it a powerful solution for dynamic content extraction and long-term web archiving. HTMLDownloader ships as enduser MSI/portable ZIP with documented workflows, enabling non-specialists to reproducibly archive JavaScriptheavy pages. A DOI-tagged release supports verification, reuse and citation (DOI: 10.5281/zenodo.16935169).
Keywords:
Web scraping
Dynamic content
Web archiving
Link extraction
HTML data collection
AI Summary

AI Summary

Key information extracted from the uploaded paper, including a brief overview, abstract, background, key highlights, visual analysis, and future outlook.

Journal

SoftwareX cover
SoftwareX
IF:
2.4
Papers:
325
Citations:
7.3K

Organization

No organization information available