arXiv Machine Learning By Xiaodong Li, Zhentao Li

Multitask Regression with Pairwise Fusion

Read the original on arXiv Machine Learning →

The paper investigates multitask regression where different predictors may have varying degrees of coefficient sharing across tasks. It introduces a framework that quantifies sharing by the number of active predictors and the total number of task-specific coefficient deviations, and proposes a pairwise penalty estimator that achieves matching upper and lower bounds in terms of these quantities. The method also handles scenarios where a large subset of tasks shares an identical coefficient vector, with explicit sample‑size conditions ensuring exact pooling of those tasks while allowing others to differ.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 10

Multi-Task Learning with Covariate-Overlap Regularization

The paper introduces COVER, a multi‑task learning framework that regularizes covariate overlap to mitigate the negative effects of sharing information across tasks with differing covariate distributions and response relationships. COVER blends a common component function, a shared neural representation, and low‑dimensional task‑specific coefficients, using taskwise second‑moment matrices to guide coefficient integration. The authors provide theoretical bias‑variance analysis, oracle inequalities, and neural‑network convergence rates, and demonstrate that COVER outperforms existing deep‑learning and statistical integration methods in simulations and a GTEx central‑nervous‑system study.

By Yang Sui, Qi Xu, Yang Bai, Annie Qu
arXiv Machine Learning
1d ago

Model Merging via Data-Free Covariance Estimation

The paper introduces a data‑free method for model merging that estimates per‑layer covariance matrices directly from difference matrices, eliminating the need for auxiliary data. This approach reduces computational costs while maintaining a principled interference‑minimization framework. Experiments on vision and language benchmarks with models from 86 M to 7 B parameters show that the method outperforms existing data‑free merging techniques.

By Marawan Gamal Abdel Hameed, Derek Tam, Pascal Jr Tikeng Notsawo, Colin Raffel, Guillaume Rabusseau