Attention-based representations for multi-task computation
arXiv:2608. 04243v1 Announce Type: new Abstract: Multi-head attention layers produce vector representations that support multiple downstream tasks.
arXiv:2502. 01015v5 Announce Type: replace Abstract: Task arithmetic, representing downstream tasks through linear operations on task vectors, has emerged as a simple yet powerful paradigm for transferring knowledge across diverse settings.
arXiv:2608. 04243v1 Announce Type: new Abstract: Multi-head attention layers produce vector representations that support multiple downstream tasks.
arXiv:2608. 10837v1 Announce Type: cross Abstract: The strong performance of foundation models for tabular tasks comes at substantial inference costs.
arXiv:2606. 18627v1 Announce Type: new Abstract: Model merging has emerged as a training-free alternative to multi-task learning, aiming to combine multiple task-specific fine-tuned models into a single multi-task model.
arXiv:2505. 23696v2 Announce Type: replace Abstract: Solving systems of polynomial equations, particularly those with finitely many solutions, is a crucial challenge across many scientific fields.
The paper introduces a new multi‑task model‑merging framework that tackles task interference by projecting task vectors into a high‑dimensional sparse feature space using Sparse Autoencoders, enabling feature‑level disentanglement before fusion. It also proposes a lightweight Group‑Ranked Zeroth‑Order Optimizer to identify task‑critical layers for selective merging, reducing computational overhead. Experiments on Qwen2.5‑1.5B and Qwen2.5‑7B show consistent performance gains over several baselines across reasoning, code generation, instruction following, and general knowledge tasks, with a 2.78% improvement in a highly conflicting four‑task setting.
arXiv:2606. 28831v1 Announce Type: cross Abstract: Long-context LLM inference faces a fundamental conflict: head-adaptive compression algorithms (e.
arXiv:2608. 01528v1 Announce Type: new Abstract: Vector symbolic architectures (VSA) are widely used for reasoning in neuro-symbolic (NeSy) AI, yet high-dimensional codebooks often create severe memory bottlenecks that limit scalability and deployment.
arXiv:2605.23200v2 Announce Type: replace-cross Abstract: The linear growth of the Key-Value (KV) cache is a critical bottleneck in long-form LLM inference. Existing KV compression methods mitigate t...
arXiv:2512. 01461v2 Announce Type: replace Abstract: Model merging has emerged as a promising paradigm for enabling multi-task capabilities without additional training.
arXiv:2510. 01718v2 Announce Type: replace Abstract: Attention is a core operation in large language models (LLMs).
arXiv:2608. 15602v1 Announce Type: cross Abstract: While binary quantization theoretically promises extreme compression and acceleration for Large Language Models (LLMs), existing research often overlooks the necessity of specialized hardware kernels, thus failing to unleash the full acceleration potential due to persistent reliance on expensive floating-point arithmetic or runtime dequantization overheads.
The paper revisits the impact of pruning on large language models (LLMs) during test-time scaling (TTS). While prior work found that structured pruning degrades reasoning performance, this study shows that unstructured pruning—removing only specific redundant weights—can actually improve TTS performance on reasoning benchmarks for models s1.1-7B and Qwen3-8B, sometimes surpassing the full-weight models. The authors also examine how different layer-wise sparsity allocation strategies affect these outcomes.