arXiv:2607. 09039v1 Announce Type: new Abstract: The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability.
By Chaoran Cheng, Zhanghan Ni, Yanru Qu, Yuxin Chen, Ruihan Guo, Jiajun Fan, Ge Liu
Vision-language models (VLMs) such as CLIP enable zero-shot classification by comparing image features with text prompts in a shared embedding space. A fundamental property underlying this capability is the global comparability of logits across arbitrary candidate classes.
Automated evaluation is essential for scaling generative 3D systems, where exhaustive human review is costly and slow. However, the reliability of an automated judge depends on the entire evaluation pipeline, not only the underlying vision-language model (VLM), but also how assets are rendered, what visual evidence is provided, how the task is specified, and how human reference labels are constructed.
The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions. A good choice of action representation and loss function can help to address these concerns, but there are often trade offs.
Preparing for job interviews is important for securing desired positions, yet realistic practice remains difficult to access: real interviews are infrequent, expert mock coaching is costly, and self-practice offers neither adaptive dialogue nor structured assessment. Existing systems typically address only parts of this need through fixed question sequences, limited communication channels, or feedback with little supporting evidence.
Human-pet interaction estimation and generation remain underexplored due to the absence of a high-quality large-scale dataset. We present InterPet4D, the first multimodal dataset capturing natural interactions between humans and dogs.
We present a novel integrated architecture for robust online 3D Gaussian splatting, real-time VR exploration, and speech-driven Vision-Language-Model interaction. Unlike methods assuming clean depth or external poses, our system combines ORB-SLAM3-based pose estimation with online Gaussian reconstruction for noisy real-world data.
arXiv:2607. 08371v1 Announce Type: cross Abstract: Synthetic multi-speaker conversations are widely used to train conversational automatic speech recognition (ASR) systems, but it remains unclear which timing properties make simulated data most useful.
By M\'at\'e Gedeon, P\'eter Mihajlik
arXiv:2607. 08725v1 Announce Type: cross Abstract: Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable.
By Ayda Eghbalian, Kevin Desai
arXiv:2607. 08711v1 Announce Type: cross Abstract: Accurate 3D terrain maps are essential for emergency response when assessing wildfire hazards.
By Xiao Fu, Yue Hu, Meida Chen, Peter Anthony Beerel, Barath Raghavan
arXiv:2607. 08233v1 Announce Type: new Abstract: A central challenge in building intelligent systems is enabling agents to jointly perceive complex inputs, form hypotheses about hidden patterns, and design informative experiments to test them.
By Sophia Koehler, Antonia W\"ust, Inga Ibs, Wasu Top Piriyakulkij, Wolfgang Stammer, Constantin Rothkopf, Kevin Ellis, Kristian Kersting
arXiv:2607. 07758v1 Announce Type: new Abstract: Foundation models (FMs) have transformed machine learning from isolated task-specific model development toward general-purpose models pretrained on broad data and adapted to multiple downstream tasks.
By Syed Usama Imtiaz, Mitra Nasr Azadani, Nasrin Alamdari
arXiv:2607. 08409v1 Announce Type: cross Abstract: LLM-based ASR adapted to regulated domains such as banking is bottlenecked by privacy: real speech is costly and legally constrained to collect, making synthetic text-to-speech (TTS) an attractive substitute.
By Shashi Kumar, Yanis Labrak, Hasindri Watawana, Sergio Burdisso, Esa\'u Villatoro-Tello, Kadri Hacio\u{g}lu, Petr Motlicek, Andreas Stolcke
arXiv:2607. 08312v1 Announce Type: new Abstract: How should language interface with a world model's discrete symbol system?
By Jiayi Fang
arXiv:2607. 08202v1 Announce Type: new Abstract: Estimating original-space conditional expectations is central to value-driven recommender systems, including dwell time, GMV, and LTV forecasting.
By Mingyu Zhao, Zhaohan Li, Zhenxiong Miao, Xu Zhang, Dewei Leng, Yanan Niu, Kun Gai
arXiv:2607. 08073v1 Announce Type: new Abstract: Fetal electrocardiogram (fECG) and Doppler ultrasound provide complementary views of fetal cardiovascular function: fECG captures electrical activity while Doppler reflects mechanical hemodynamics shaped by factors such as placental resistance and vascular compliance.
By Tongli Su, Alireza Rafiei, Marly van Assen, Reza Sameni, Gari D. Clifford, Faezeh Marzbanrad, Nasim Katebi
arXiv:2607. 07756v1 Announce Type: new Abstract: Multimodal learning usually requires a dedicated encoder per modality.
By Ilia Koloiarov, Diego Coello de Portugal Mecke, Vijaya Krishna Yalavarthi, Tom Hanika, Lars Schmidt-Thieme
arXiv:2607. 07855v1 Announce Type: new Abstract: Hierarchical Implicit Q-Learning (HIQL), an offline goal-conditioned RL method, selects subgoals by value-function advantages alone.
By Erdemt Bao, Xing Lei, Jun Chen
arXiv:2607. 08024v1 Announce Type: cross Abstract: Long-horizon robot planning requires jointly reasoning over semantic task structure and geometric feasibility.
By Emily Jin, Joy Hsu, Yiqing Xu, Weiyu Liu, Nick Haber, Jiajun Wu
arXiv:2601. 00969v3 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models provide strong action priors for robotic manipulation, but their reactive behavior can fail under distribution shift and long-horizon task structure.
By Ke Ren, Ali Salamatian, Kieran Pattison, Cyrus Neary