arXiv Machine Learning

MINT: A Universal Zero-Shot Predictor for Transaction Data

arXiv:2608. 14198v1 Announce Type: new Abstract: Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization.

arXiv Machine Learning
Sep 24

ChronoSteer: Bridging Large Language Model and Time Series Foundation Model via Synthetic Cross-Modal Alignment Dataset

ChronoSteer is a decoupled agentic framework that bridges large language models and time series foundation models by learning cross‑modal alignment from synthetic paired supervision. It converts textual events into revision instructions that steer a frozen time‑series model, discretizes these instructions into a compact codebook to reduce semantic divergence, and then refines the predictions with a two‑stage training strategy. The authors also release a leakage‑controlled multimodal benchmark and report a 25.8% improvement in zero‑shot prediction accuracy over the unimodal backbone.

By Chengsen Wang, Qi Qi, Zhongwen Rao, Lujia Pan, Jingyu Wang
arXiv Machine Learning
Sep 14

FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences

FINESSE is an agent‑based simulation framework that generates synthetic, structured datasets of multiple interdependent financial event streams, such as transactions, payments, account status changes, and policy interventions. Each stream has its own action space, schema, and variable types, and the streams are coupled through agents’ evolving latent states, allowing temporally rich interactions. The accompanying FINESSE‑Bench dataset supports four tasks—balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction—and baseline results are provided using various time‑series and event‑sequence methods.

By Tyler Farnan, Benjamin Eng, Adam Abate, Xirui Hou, Rizal Fathony, Nam H. Nguyen, Senthil Kumar
arXiv Computation and Language
Aug 27

Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER

Just Pass Twice (JPT) is a method that allows causal large language models to perform token classification for zero‑shot named entity recognition by concatenating the input with itself, giving each token full bidirectional context without architectural changes. The approach combines these representations with definition‑guided entity embeddings to enable flexible zero‑shot generalization. JPT achieves state‑of‑the‑art results, outperforming prior methods by an average of +7.9 F1 on CrossNER and MIT benchmarks and running over 20× faster than comparable generative approaches.

By Ahmed Ewais, Ahmed Hashish, Amr Ali
arXiv Computation and Language
Aug 25

PRAGMA: Revolut Foundation Model

arXiv:2604.08649v2 Announce Type: replace-cross Abstract: Modern financial systems generate vast quantities of transactional and event-level data that encode rich economic signals. This paper present...

By Maxim Ostroukhov, Ruslan Mikhailov, Vladimir Iashin, Artem Sokolov, Andrei Akshonov, Vitaly Protasov, Andrey Goncharov, Dmitrii Beloborodov, Vince Mullin, Roman Yokunda Enzmann, Georgios Kolovos, Jason Renders, Pavel Nesterov, Anton Repushko
Hugging Face Trending Papers
Jul 2

FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval

Zero-shot composed image retrieval (ZS-CIR) aims to retrieve a target image by editing a reference image with a natural-language instruction, without relying on domain-specific annotated triplets. Most existing ZS-CIR methods rely on textual inversion to translate the reference image into pseudo-text tokens and then compose them with the instruction via simple concatenation in the text space, which can be lossy and brittle for fine-grained semantics.

arXiv AI
Aug 26

From Gradient-Boosted Trees to Deep Recommenders: Practical Lessons from Migrating a Production Customer Support Recommender

The paper documents the migration of a live conversational recommendation system from a gradient‑boosted multiclass model to a pairwise‑binary deep recommender. It explains how reformulating the task, using negative sampling, noise injection, and attention pooling over transcript chunks enabled the new model to handle dynamic, multimodal data and long conversation context. The authors compare several architectures and loss functions, showing that the deep recommender matches or surpasses the CatBoost baseline, especially in later conversational stages.

By Sonia Sharma, Jeyendran Balakrishnan, Shreya Rajpal, Swapnil Parekh, Nagaraj Janardhana, Andrew Mattarella-Micke