arrow
返回

Same or Different? Diff-Vectors for Authorship Analysis

delete2023-09-06
delete3
delete
OA
AI
S
Silvia Corbara *
A
Alejandro Moreo
F
Fabrizio Sebastiani
DOI:10.1145/3609226delete
delete原文链接
delete原文求助
delete分享
delete收藏
摘要

摘要

En 中文
In this article, we investigate the effects on authorship identification tasks (including authorship verification, closed-set authorship attribution, and closed-set and open-set same-author verification) of a fundamental shift in how to conceive the vectorial representations of documents that are given as input to a supervised learner. In classic authorship analysis, a feature vector represents a document, the value of a feature represents (an increasing function of) the relative frequency of the feature in the document, and the class label represents the author of the document. We instead investigate the situation in which a feature vector represents an unordered pair of documents, the value of a feature represents the absolute difference in the relative frequencies (or increasing functions thereof) of the feature in the two documents, and the class label indicates whether the two documents are from the same author or not. This latter (learner-independent) type of representation has been occasionally used before, but has never been studied systematically. We argue that it is advantageous, and that, in some cases (e.g., authorship verification), it provides a much larger quantity of information to the training process than the standard representation. The experiments that we carry out on several publicly available datasets (among which one that we here make available for the first time) show that feature vectors representing pairs of documents (that we here call Diff-Vectors) bring about systematic improvements in the effectiveness of authorship identification tasks, and especially so when training data are scarce (as it is often the case in real-life authorship identification scenarios). Our experiments tackle same-author verification, authorship verification, and closed-set authorship attribution; while DVs are naturally geared for solving the 1st, we also provide two novel methods for solving the 2nd and 3rd that use a solver for the 1st as a building block. The code to reproduce our experiments is open-source and available online.(1)
Keyword:
Supervised learning
vector-based representations
authorship analysis

期刊

ACM Transactions on Knowledge Discovery from Data 封面图
ACM Transactions on Knowledge Discovery from Data
IF:
4.8
论文数:
1.3K
被引数:
4.4K

机构

S
scuola normale superiore di pisa
学者数:
2.8K
论文数: 3.1K
被引数: 2
C
consiglio nazionale delle ricerche (cnr)
学者数:
6.2W
论文数: 5.7W
被引数: 48
引用论文

引用论文

err
IF0
err
err0
PREAI
err
err分享
err收藏
err
IF0
err
err0
PREAI
err
err分享
err收藏
err分享
err收藏
Predicting Traffic Sign Retro-Reflectivity Degradation Using Deep Neural Networks
err2021-12-07
err0
errOAAI
errAbdolmaged Alkhulaifi; Arshad Jamal; Irfan Ahmad
err分享
err收藏
Zinc replacement in hepatic encephalopathy among the Egyptian patients埃及患者肝性脑病中的锌替代治疗
err2020-04-01
err0
errOAAI
errMohamed Al-Alfy; Amin Amin Hegazy; Kamel Soliman Hammad; Ahmed El-Sayed Shehata
err分享
err收藏
Peptide aldehyde inhibitors challenge the substrate specificity of the SARS-coronavirus main protease
err2011-11-01
err0
errOAAI
errLili Zhu; Shyla George; Marco F. Schmidt; Samer I. Al-Gharabli; Jörg Rademann; Rolf Hilgenfeld
err分享
err收藏
学者 查看更多内容