返回
Framework for kernel regularization with application to protein clustering
DOI:10.1073/pnas.0505411102.png)
摘要
En 中文
We develop and apply a previously undescribed framework that is designed to extract information in the form of a positive definite kernel matrix from possibly crude, noisy, incomplete, inconsistent dissimilarity information between pairs of objects, obtainable in a variety of contexts. Any positive definite kernel defines a consistent set of distances, and the fitted kernel provides a set of coordinates in Euclidean space that attempts to respect the information available while controlling for complexity of the kernel. The resulting set of coordinates is highly appropriate for visualization and as input to classification and clustering algorithms. The framework is formulated in terms of a class of optimization problems that can be solved efficiently by using modern convex cone programming software. The power of the method is illustrated in the context of protein clustering based on primary sequence data. An application to the globin family of proteins resulted in a readily visualizable 3D sequence space of globins, where several subfamilies and subgroupings consistent with the literature were easily identifiable.
Keyword:
classification
convex cone programming
dissimilarity information
trace penalty
sequence data
AI总结
对已上传原文的论文进行重点信息的提取,主要内容包括:简要概述、研究摘要、背景介绍、关键亮点、图文解析、展望与总结。
期刊
P
IF:
9.1
论文数:
10.8W
被引数:
73.5W
机构
暂无机构信息
引用论文
JASPAR:: an open-access database for eukaryotic transcription factor binding profiles
NUCLEIC ACIDS RESEARCH
IF13.1

