Inferring feature importance with uncertainties in high-dimensional data

Johnsen, Pål Vegard; Strümke, Inga; Riemer-Sørensen, Signe; DeWan, Andrew Thomas; Langaas, Mette

Computer Science > Machine Learning

arXiv:2109.00855 (cs)

[Submitted on 2 Sep 2021 (v1), last revised 20 Sep 2021 (this version, v3)]

Title:Inferring feature importance with uncertainties in high-dimensional data

Authors:Pål Vegard Johnsen, Inga Strümke, Signe Riemer-Sørensen, Andrew Thomas DeWan, Mette Langaas

View PDF

Abstract:Estimating feature importance is a significant aspect of explaining data-based models. Besides explaining the model itself, an equally relevant question is which features are important in the underlying data generating process. We present a Shapley value based framework for inferring the importance of individual features, including uncertainty in the estimator. We build upon the recently published feature importance measure of SAGE (Shapley additive global importance) and introduce sub-SAGE which can be estimated without resampling for tree-based models. We argue that the uncertainties can be estimated from bootstrapping and demonstrate the approach for tree ensemble methods. The framework is exemplified on synthetic data as well as high-dimensional genomics data.

Subjects:	Machine Learning (cs.LG); Methodology (stat.ME); Machine Learning (stat.ML)
Cite as:	arXiv:2109.00855 [cs.LG]
	(or arXiv:2109.00855v3 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2109.00855

Submission history

From: Inga Strümke [view email]
[v1] Thu, 2 Sep 2021 11:57:34 UTC (237 KB)
[v2] Mon, 6 Sep 2021 07:24:58 UTC (236 KB)
[v3] Mon, 20 Sep 2021 06:52:11 UTC (240 KB)

Full-text links:

Access Paper:

view license

Current browse context:

cs.LG

< prev | next >

new | recent | 2021-09

Change to browse by:

cs
stat
stat.ME
stat.ML

References & Citations

DBLP - CS Bibliography

listing | bibtex

Signe Riemer-Sørensen

export BibTeX citation

Computer Science > Machine Learning

Title:Inferring feature importance with uncertainties in high-dimensional data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Inferring feature importance with uncertainties in high-dimensional data

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators