Deep Multitask Learning for Mixed-Type Outcomes with Shared Sparsity
About
Most existing multitask learning approaches are limited by their reliance on task-specific loss functions tailored to the scale and type of each outcome. When outcomes differ across tasks, these losses are generally not directly comparable, which makes it difficult to formulate a unified objective and may limit information sharing across tasks. We propose a multitask transformation framework in which task-specific responses may differ through unknown monotone transformations. Motivated by high-dimensional biological applications in which the predictor dimension may diverge with the sample size while only a common subset of predictors is informative, we consider shared sparsity across tasks. Under this framework, we estimate the target functions and identify important predictors by optimizing a smoothed rank-based criterion with a group-Lasso penalty, implemented through a multitask deep neural network with a shared first layer. We establish the nonasymptotic excess-risk bounds, and variable-selection consistency for the proposed estimator. Simulation studies show that the proposed method achieves competitive prediction and variable-selection performance compared with competing approaches. Analyses of gene-expression studies with continuous, binary, and mixed outcomes further illustrate that the proposed method improves prediction and identifies biologically meaningful shared predictors.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Classification | Simulation Single-type (Setting 3, p=100) | Accuracy71.1 | 6 | |
| Classification | Simulation Single-type Setting 4, p=100 | Accuracy59.9 | 6 | |
| Mixed-type Multitask Learning | Lung cancer | Accuracy96.8 | 6 | |
| Mixed-type outcome prediction | Simulation Setting 5 | Accuracy62.7 | 6 | |
| Mixed-type outcome prediction | Simulation Setting 6 | Accuracy62.6 | 6 | |
| Multitask regression | Mouse genetics | MAE0.827 | 6 | |
| Regression | Simulation Single-type Setting 1, p=100 | MAE2.046 | 6 | |
| Regression | Simulation Single-type Setting 2, p=100 | MAE2.779 | 6 | |
| Multitask classification | METABRIC | Accuracy77.9 | 6 |