Correction of overfitting bias in regression models

Massa, Emanuele; Jonker, Marianne; Roes, Kit; Coolen, Anthony

Statistics > Methodology

arXiv:2204.05827 (stat)

[Submitted on 12 Apr 2022 (v1), last revised 4 Sep 2023 (this version, v3)]

Title:Correction of overfitting bias in regression models

Authors:Emanuele Massa, Marianne Jonker, Kit Roes, Anthony Coolen

View PDF

Abstract:Regression analysis based on many covariates is becoming increasingly common. However, when the number of covariates $p$ is of the same order as the number of observations $n$, maximum likelihood regression becomes unreliable due to overfitting. This typically leads to systematic estimation biases and increased estimator variances. It is crucial for inference and prediction to quantify these effects correctly. Several methods have been proposed in literature to overcome overfitting bias or adjust estimates. The vast majority of these focus on the regression parameters. But failure to estimate correctly also the nuisance parameters may lead to significant errors in confidence statements and outcome prediction.
In this paper we present a jacknife method for deriving a compact set of non-linear equations which describe the statistical properties of the ML estimator in the regime where $p=O(n)$ and under the hypothesis of normally distributed covariates. These equations enable one to compute the overfitting bias of maximum likelihood (ML) estimators in parametric regression models as functions of $\zeta = p/n$. We then use these equations to compute shrinkage factors in order to remove the overfitting bias of maximum likelihood (ML) estimators. This new derivation offers various benefits over the replica approach in terms of increased transparency and reduced assumptions. To illustrate the theory we performed simulation studies for multiple regression models. In all cases we find excellent agreement between theory and simulations.

Comments:	6 figures, 38 pages including appendices
Subjects:	Methodology (stat.ME); Statistics Theory (math.ST); Data Analysis, Statistics and Probability (physics.data-an)
MSC classes:	62J99
Cite as:	arXiv:2204.05827 [stat.ME]
	(or arXiv:2204.05827v3 [stat.ME] for this version)
	https://doi.org/10.48550/arXiv.2204.05827

Submission history

From: Emanuele Massa [view email]
[v1] Tue, 12 Apr 2022 14:11:13 UTC (1,554 KB)
[v2] Mon, 3 Jul 2023 15:49:18 UTC (955 KB)
[v3] Mon, 4 Sep 2023 09:44:43 UTC (1,138 KB)

Statistics > Methodology

Title:Correction of overfitting bias in regression models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Methodology

Title:Correction of overfitting bias in regression models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators