返回
The reusability prior: comparing deep learning models without training
DOI:10.1088/2632-2153/acc713.png)
摘要
En 中文
Various choices can affect the performance of deep learning models. We conjecture that differences in the number of contexts for model components during training are critical. We generalize this notion by defining the reusability prior as follows: model components are forced to function in diverse contexts not only due to the training data, augmentation, and regularization choices, but also due to the model design itself. We focus on the design aspect and introduce a graph-based methodology to estimate the number of contexts for each learnable parameter. This allows a comparison of models without requiring any training. We provide supporting evidence with experiments using cross-layer parameter sharing on CIFAR-10, CIFAR-100, and Imagenet-1K benchmarks. We give examples of models that share parameters outperforming baselines that have at least 60% more parameters. The graph-analysis-based quantities we introduced for the reusability prior align well with the results, including at least two important edge cases. We conclude that the reusability prior provides a viable research direction for model analysis based on a very simple idea: counting the number of contexts for model parameters.
Keyword:
entropy
deep learning
parameter efficiency
reusability
期刊
M
IF:
4.6
论文数:
1.1K
被引数:
3.4K
机构
引用论文
ACORT: A compact object relation transformer for parameter efficient image captioning
NEUROCOMPUTING
IF6.5
Gradient-based learning applied to document recognition基于梯度的学习在文档识别中的应用
PROCEEDINGS OF THE IEEE
IF25.9
Two-dimensional bricklayer arrangements of tolans using halogen bonding interactions使用卤素键相互作用的tolans的二维瓦工层布置
Dietary trans and saturated fatty acids effects on semen quality, hormonal levels and expression of genes related to steroid metabolism in mouse adipose tissue
Andrologia
IF0

