arXiv:2609.37569v1 Announce Type: new
Abstract: Rendering accurate Chinese text remains challenging for text-to-image models. Existing OCR-based reinforcement-learning rewards compare decoded transcr...
By Yazhen Xie, Xingsong Ye, Zhineng Chen
arXiv:2609.37648v1 Announce Type: new
Abstract: Preoperative liver-tumor assessment requires segmentation, physical-space measurement, visual evidence, and resection planning from the same three-dime...
By Binghong Qian, Xuanhe Liu, Yifan Xing, Wenjie Deng, Jian Wu, Haochao Ying
arXiv:2609.37712v1 Announce Type: cross
Abstract: Optical Character Recognition (OCR) is evolving from plain-text transcription toward general visual intelligence, requiring models to recognize, loca...
By GuangJian Team, Kaili Huang, Yongshuo Zhang, Bingtao Fu, Changjiang Jiang, Chenfan Qu, Chenfeng Zhang, Fangming Cui, Gaoyang Zhang, Jiangwei Xie, Jianshu Li, Jing Huang, Jingwen Bai, Mingqi Fang, Tao Fang, Weihong Zhang, Wenbo Du, Xiongfei Bai, Xuekang Zhu, Yinan Xia, Zhenming Wang, Jian Liu, Jingjing Liu, Xiang Qi, Weiqiang Wang
arXiv:2605.16981v3 Announce Type: replace
Abstract: Streaming 3D reconstruction under a strict constant-memory budget hinges on how the recurrent state is updated as the stream evolves. We profile TT...
By Kejun Ren, Lei Jin, Tianxin Huang, Lianming Xu, Li Wang
arXiv:2609.36302v1 Announce Type: new
Abstract: While foundation models have revolutionized natural language processing and computer vision by leveraging universal vocabularies, Graph Machine Learnin...
By Ben Finkelshtein, Andr\'{e} Linhares, Petar Veli\v{c}kovi\'{c}, Bryan Perozzi, Mikhail Galkin
arXiv:2609.36134v1 Announce Type: cross
Abstract: Functional Kolmogorov-Arnold Networks (FunKAN) achieve state-of-the-art accuracy on MRI Gibbs artifact removal and anatomical segmentation, but their...
By Mohammad Sadegh Sirjani
arXiv:2609.36172v1 Announce Type: cross
Abstract: Figuring out how objects relate to each other, like whether they touch, overlap, stay completely separate or one sits inside another, matters a lot i...
By Saptak Das, Monidipa Das
arXiv:2609.36348v1 Announce Type: cross
Abstract: Generative and representation learning remain asymmetrically connected: semantic representations are used to improve diffusion generation, whereas th...
By Xiaoyu Wu, Yifei Wang, Chen Wei
arXiv:2609.37177v1 Announce Type: cross
Abstract: Persistent homology (PH) is a frequently used tool for extracting and preserving topological information from image data, particularly in image segme...
By Alexander H. Berger, Marco Fontana, Daniel Rueckert, Johannes C. Paetzold, Laurin Lux, Ulrich Bauer
arXiv:2609.37784v1 Announce Type: cross
Abstract: Satellite observations, precomputed embeddings, and map products describe the same evolving Earth, yet are stored as independent, petabyte-scale data...
By Arjun Rao, Sebastian Loeschcke, Anthony Fuller, Isaac Corley, Nico Lang, Evan Shelhamer
arXiv:2609.38112v1 Announce Type: cross
Abstract: Many applications of black-box predictive models require controlling task-relevant error rates, such as missed lesion pixels in segmentation or misse...
By Bruno Marcondes e Resende, Helton Graziadei, Thiago Rodrigo Ramos, Rafael Izbicki
arXiv:2609.38165v1 Announce Type: cross
Abstract: The landscape of satellite imagery time series datasets and boundary-pushing architectures for cropland segmentation has never been richer. However,...
By Joseph Metcalfe, Sara Sharifzadeh, Fabio Caraffini
The paper investigates the impact of low‑resolution data on machine learning models, especially when high‑resolution data is scarce. It provides a theoretical analysis using Kullback‑Leibler divergence to quantify how data resolution affects a datapoint’s influence and derives bounds on information loss from downsampling. Empirical results with a vision transformer and a convolutional neural network show that incorporating low‑resolution data consistently improves performance in such settings.
By Theresa Dahl Frehr, Niels Henrik Pontoppidan, Hiba Nassar, Tommy Sonne Alstr{\o}m
arXiv:2509.25549v3 Announce Type: replace-cross
Abstract: Choroidal nevi are common benign pigmented lesions in the eye, with a small risk of transforming into melanoma. Early detection is critical t...
By Mohammadmahdi Eshragh, Emad A. Mohammed, Behrouz Far, Ezekiel Weis, Carol L Shields, Sandor R Ferenczy, Trafford Crump
arXiv:2606.16996v2 Announce Type: replace-cross
Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabu...
By Tran Dinh Tien, Zhiqiang Shen
arXiv:2603.10861v2 Announce Type: replace
Abstract: SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication date...
By Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, Johan Sofalas
Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this canc...
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understandin...
The ByteTrack algorithm is a widely used and computationally efficient multi-object tracking architecture. Its core innovation lies in the combination of lenient bounding box associations with trackle...
There has been significant work on understanding the In-Context Learning capabilities of Large Language Models, especially on the induction circuit. For a few-shot classification task, the induction c...