Robust Regression of General ReLUs with Queries
arXiv:2606. 11130v1 Announce Type: new Abstract: We study the task of agnostically learning general (as opposed to homogeneous) ReLUs under the Gaussian distribution with respect to the squared loss.
The paper investigates how the ability to synthesize arbitrary queries (membership queries) changes the sample complexity of active learning compared to the traditional pool-based setting. It shows that some hypothesis classes that only achieve polynomial error decay with pool-based queries become exponentially learnable when synthesis is allowed, revealing a significant gap in learning difficulty. The authors propose sufficient conditions, provide examples, and suggest a conjectural framework to identify classes that benefit from synthesized queries.
arXiv:2606. 11130v1 Announce Type: new Abstract: We study the task of agnostically learning general (as opposed to homogeneous) ReLUs under the Gaussian distribution with respect to the squared loss.
arXiv:2311. 00260v2 Announce Type: replace-cross Abstract: In collaborative active learning, where multiple agents try to learn labels from a common hypothesis, we introduce an innovative framework for incentivized collaboration.
arXiv:2607. 02895v1 Announce Type: cross Abstract: Learning-based approaches for selectivity estimation in databases have gained significant traction in recent years.
arXiv:2606. 10125v1 Announce Type: cross Abstract: Few-shot example retrieval is the dominant paradigm for grounding large language models (LLMs) in domain-specific text-to-SQL systems.
arXiv:2607. 04869v1 Announce Type: new Abstract: Motivated by real-world scenarios where malicious entities tamper with existing networks, we define a model where an adversary seeks to hide a set of \emph{corrupted vertices} inside a graph $G^*$.
arXiv:2607. 17632v1 Announce Type: cross Abstract: Time series data play a pivotal role across numerous domains, including healthcare and manufacturing.
arXiv:2608. 03249v1 Announce Type: new Abstract: Cold-Start Active Learning (CSAL) aims to select a valuable subset from an unlabeled pool without any prior knowledge or human assistance.
arXiv:2307.02719v5 Announce Type: replace Abstract: Uncertainty sampling is a classical active-learning strategy, yet the statistical objective induced by its query rule is often implicit. We introdu...
arXiv:2512.12870v2 Announce Type: replace-cross Abstract: Active Learning (AL) is commonly used in applications where labeling data is expensive or time-consuming. In practice, however, labels are of...
arXiv:2606. 17805v1 Announce Type: new Abstract: Data acquisition is a major bottleneck for learning in real-time streams: analysts must decide on the fly which labels to purchase while respecting a rolling budget.
The paper extends the study of relatively smart learning, showing that ERM and any proper consistent learner are relatively smart for binary classification in the distribution‑free setting, achieving a quadratic sample‑complexity blowup. It further demonstrates that semi‑supervised relatively smart learning is possible with only a quadratic blowup in unlabeled data and no blowup in labeled data, though this requires a leave‑most‑out transductive approach and incurs intractability when only an agnostic ERM oracle is available. The results clarify the trade‑offs between sample efficiency, label efficiency, and computational tractability in relatively smart learning.
LIMIT (Less Is More for Instruction Tuning in Text-to-SQL) challenges the belief that large instruction corpora are necessary for effective Text-to-SQL models. The framework uses a four‑stage data‑centric process—difficulty‑aware filtering, chain‑of‑thought synthesis, LLM‑as‑judge quality scoring, and genetic algorithm optimization—to select a compact set of examples that still achieve full schema coverage. On the BIRD and Spider benchmarks, LIMIT’s 796 and 863 samples enable Qwen3‑8B to reach 69.1% and 88.9% execution accuracy, outperforming methods trained on twenty times more data and setting a new state‑of‑the‑art for open‑source approaches.