arXiv:2606. 24986v1 Announce Type: new Abstract: Automated cattle posture-classification systems frequently report near-perfect accuracy, yet their robustness under realistic deployment conditions remains largely unknown.
By Leutrim Uka, Severino Pinto, Gundula Hoffmann, Marina M. -C. H\"ohne
arXiv:2608. 09943v1 Announce Type: cross Abstract: Monitoring livestock behaviour under extensive conditions would provide valuable insights to assess animal adaption to environmental perturbations in agroecological systems (e.
By Lucile Riaboff (GenPhySE, INRAE), Ny Aina Andriamampandry (GenPhySE, GenPhySE), Jean-Fran\c{c}ois Bompa (GenPhySE, GenPhySE), Mathias Aletru (GenPhySE, GenPhySE), Christian Durand (UEF), S\'ebastien Douls (UEF), Ga\"etan Bonnafe (UEF), Morgane Costes-Thir\'e (GenPhySE, GenPhySE), Guillaume Delosi\`eres (GenPhySE, GenPhySE), Jean- Marc Mongrelet (GenPhySE, GenPhySE), Enzo Niro (GenPhySE, GenPhySE), N\'emuel Tadi (GenPhySE, GenPhySE), S\'everine Deretz (DEPT GA, UEF, INRAE), Sara Parisot (UEF), Margot Lamarque (UEF), Dominique Hazard (GenPhySE), Emilie Cobo (GenPhySE)
arXiv:2605. 16301v2 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries.
By Isabella Luong, Joyee Chen, Arturs Kanepajs, Jasmine Brazilek, Sankalpa Ghose, David Williams-King, Linh Le, Allen Lu
The paper investigates whether detailed articulated human pose provides more discriminative power than coarse spatial relationships for early violence detection. By fixing the downstream pipeline and comparing five interaction representations—including bounding‑box geometry, handcrafted pose analogues, enriched pose descriptors, and a learned joint encoder—the study finds that pose‑based representations do not outperform coarse geometry. When visual encoders are frozen and evaluated on larger datasets, person‑crop appearance and whole‑frame context outperform geometry, but cropping to interacting people offers no advantage over encoding the entire frame. The authors further demonstrate that pre‑onset frames contain source‑related artifacts (e.g., title cards, watermarks) that contribute significantly to discrimination, suggesting that benchmark performance may reflect these artifacts rather than true event evidence.
By Parishruthi Ganesh
The paper introduces ChartBias, a benchmark of 820 real-world charts covering six social attributes, designed to audit bias in vision‑language models (VLMs) that interpret charts. Across 12 VLMs, the study identifies three failure modes—narrative shift, group hallucination, and preference polarity—where models produce different or misleading narratives when the referenced social group changes. A multi‑agent mitigation framework is proposed, separating evidence extraction from group‑conditioned generation and using a counterfactual judge, which reduces narrative shift while maintaining chart‑grounded reasoning.
By Mizanur Rahman, Huan Wu, Arash Asgari, Enamul Hoque Prince, Laleh Seyyed-Kalantari
arXiv:2608. 11322v1 Announce Type: cross Abstract: Human-AI research often evaluates individual capabilities, combined performance, or final outputs, but these approaches do not preserve how one party's response becomes part of the conditions under which the other party's next contribution is formed.
By Mehmed Zahid \c{C}\"ogenli