arXiv Machine Learning

Binary Classification from Coupled Pairwise Labels

The paper introduces SD-Pcomp learning, a binary classification framework that jointly utilizes Similarity/Dissimilarity (SD) labels and Pairwise Comparison (Pcomp) labels from instance pairs. It proposes an objective function that can be decomposed into either an SD estimator plus ordering information or a Pcomp estimator plus pair-type information, thereby integrating complementary relational cues. Experiments on eight datasets demonstrate that combining both label types improves classification accuracy and AUC compared to using either alone or a simple convex combination.

arXiv Machine Learning
Sep 18

Alliance Beats Isolation: Unifying Heterogeneous Allied Datasets Improves Classifier Performance

The paper introduces a method for combining heterogeneous, allied datasets—datasets that share the same class labels but have disjoint objects and largely distinct feature spaces—into a single unified feature space. By applying matrix completion to this merged space, the authors create a unified dataset that enables knowledge transfer between the original datasets. Experiments across multiple dataset pairs and classifiers show that models trained on the unified representation consistently outperform those trained separately on each dataset.

By Girish Keshav Palshikar
arXiv AI
Sep 21

The Impact of Semantic Pairs on Self-Supervised Representation Learning

The paper investigates the effect of using semantic positive pairs—different instances of the same class—in self‑supervised visual representation learning. By creating matched ImageNet‑1K subsets of augmented pairs and manually curated semantic pairs, the authors compare contrastive and non‑contrastive SSL methods under identical training conditions. Across transfer learning and object detection tasks, semantic‑pair pretraining consistently outperforms augmented‑pair pretraining, with contrastive methods like SimCLR showing the largest gains, indicating that semantic pairs foster additional invariances beyond standard augmentations.

By Mohammad Alkhalefi, Georgios Leontidis, Mingjun Zhong
arXiv Machine Learning
Sep 2

Convergence issues in Relational Concept Analysis based on AOC-posets

The paper examines convergence problems in Relational Concept Analysis (RCA) when applied to AOC-posets instead of full concept lattices. It explains why RCA’s iterative process may fail to converge in the AOC-poset setting, identifies conditions that can still guarantee convergence, and proposes a convergent variant that preserves the AOC-poset structure by never removing relational attributes. The study also discusses data transformations that can restore convergence.

By Xavier Dolques, Agn\`es Braud, Alain Gutierrez, Marianne Huchard, Florence Le Ber
arXiv Machine Learning
Jun 17

Rethinking Dataset Distillation for Classification: Do Distilled Sets Outperform Coresets?

arXiv:2606. 18209v1 Announce Type: new Abstract: Dataset distillation (DD) has emerged as a prominent approach in data centric machine learning, aiming to synthesize compact training sets for efficient training by compressing the information in large datasets into a small number of synthetic samples.

By Trisha Mittal, Akshay Mehra, Joshua Kimball