ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.
By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
arXiv:2607. 09955v1 Announce Type: cross Abstract: Predictive modeling is a core component of modern financial services, where a wide range of tasks are traditionally addressed using separate models trained on manually engineered tabular features.
By Nikita Rusakov, Vladislav Meshkov, Konstantin Zorin, Gleb Zaripov, Alexander Uglov, Alexey Vasilev, Anton Klenitskiy
FINESSE is an agent‑based simulation framework that generates synthetic, structured datasets of multiple interdependent financial event streams, such as transactions, payments, account status changes, and policy interventions. Each stream has its own action space, schema, and variable types, and the streams are coupled through agents’ evolving latent states, allowing temporally rich interactions. The accompanying FINESSE‑Bench dataset supports four tasks—balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction—and baseline results are provided using various time‑series and event‑sequence methods.
By Tyler Farnan, Benjamin Eng, Adam Abate, Xirui Hou, Rizal Fathony, Nam H. Nguyen, Senthil Kumar
arXiv:2604. 00513v3 Announce Type: replace-cross Abstract: With the rapid growth of e-commerce, exploring general representations rather than task-specific ones has attracted increasing attention.
By Junxian Wu, Chenghan Fu, Zhanheng Nie, Daoze Zhang, Bowen Wan, Wanxian Guan, Chuan Yu, Jian Xu, Bo Zheng
Just Pass Twice (JPT) is a method that allows causal large language models to perform token classification for zero‑shot named entity recognition by concatenating the input with itself, giving each token full bidirectional context without architectural changes. The approach combines these representations with definition‑guided entity embeddings to enable flexible zero‑shot generalization. JPT achieves state‑of‑the‑art results, outperforming prior methods by an average of +7.9 F1 on CrossNER and MIT benchmarks and running over 20× faster than comparable generative approaches.
By Ahmed Ewais, Ahmed Hashish, Amr Ali
arXiv:2604.08649v2 Announce Type: replace-cross
Abstract: Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper present...
By Maxim Ostroukhov, Ruslan Mikhailov, Vladimir Iashin, Artem Sokolov, Andrei Akshonov, Vitaly Protasov, Andrey Goncharov, Dmitrii Beloborodov, Vince Mullin, Roman Yokunda Enzmann, Georgios Kolovos, Jason Renders, Pavel Nesterov, Anton Repushko
arXiv:2607. 20228v1 Announce Type: new Abstract: We propose a hybrid approach for user-centric modeling of transactional event sequences that combines contrastive representation learning (CoLES) with State Space Models (SSMs).
By Ivan Palagin
Merchant risk control at large payment platforms screens tens of millions of merchants daily, where false positives harm legitimate merchants and false negatives leave harmful activity undetected. The hardest cases require jointly understanding a merchant's textual profile and long behavioral sequence.
While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context limits. To address this, sequential approaches like...
Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image by editing a reference image with a natural-language instruction, without relying on domain-specific annotated triplets. Most existing ZS-CIR methods rely on textual inversion to translate the reference image into pseudo-text tokens and then compose them with the instruction via simple concatenation in the text space, which can be lossy and brittle for fine-grained semantics.
arXiv:2609.10441v1 Announce Type: cross
Abstract: While Large Language Models (LLMs) have demonstrated impressive capabilities, they often struggle with extremely long contexts due to fixed context l...
By Hongming Zhang, Zhaozhen Gu, Fengshuo Bai, Ming Hao, Qingyang Zhang, Yuanyuan Wang, Shiyang Tang, Yanna Wang, Bo Xu
The paper documents the migration of a live conversational recommendation system from a gradient‑boosted multiclass model to a pairwise‑binary deep recommender. It explains how reformulating the task, using negative sampling, noise injection, and attention pooling over transcript chunks enabled the new model to handle dynamic, multimodal data and long conversation context. The authors compare several architectures and loss functions, showing that the deep recommender matches or surpasses the CatBoost baseline, especially in later conversational stages.
By Sonia Sharma, Jeyendran Balakrishnan, Shreya Rajpal, Swapnil Parekh, Nagaraj Janardhana, Andrew Mattarella-Micke