Sentiment analysis with frozen pre-trained language model (PLM) backbones has become a common paradigm, yet the practical benefit of explicit domain adaptation remains unclear, particularly when backbones encode varying degrees of target-domain knowledge. We present a preliminary case study evaluating a controlled family of frozen embedding backbones (Qwen3-Embedding 0.
arXiv:2104. 08928v4 Announce Type: replace-cross Abstract: Unstructured text provides decision-makers with a rich data source in many domains, ranging from product reviews in retail to nursing notes in healthcare.
By Kan Xu, Xuanyi Zhao, Hamsa Bastani, Osbert Bastani
arXiv:2608.30609v1 Announce Type: cross
Abstract: Large language models are increasingly capable in general, but their utility can remain modest in niche or understudied areas. One approach to addres...
By Lukas Borggren, Jenny Kunz, Marco Kuhlmann
arXiv:2507.09601v3 Announce Type: replace-cross
Abstract: Financial text embeddings must distinguish changes in event status, perspective, and obligations even when passages share similar wording. NM...
By Hanwool Lee, Sara Yu, Yewon Hwang, Jonghyun Choi, Heejae Ahn, Sungbum Jung, Youngjae Yu
The paper investigates cross‑lingual transfer for sequential sentence classification (SSC) in research papers, focusing on 13 non‑English languages. Experiments show that linguistic proximity does not reliably predict transfer success, whereas structural similarity in rhetorical organization—particularly label distribution similarity—correlates positively with performance. The authors introduce three generative‑model methods that exploit structural cues, achieving parity with strong encoder baselines on‑domain and outperforming them when transferring to unseen languages.
By Kazuhiro Yamauchi, Marie Katsurai
The paper introduces Distilled Rapid Embedding Transfer (DRET), a parameter‑efficient method that injects biomedical domain knowledge from large specialized models into a smaller general‑purpose model without retraining on the original specialized corpora. DRET evolves through iterative strategies—tokenizer‑merge (DRET 1.x), hybrid embedding averaging (DRET 2.0), priority‑based embedding transfer (DRET 3.x), and further refinements (DRET 4.x)—and demonstrates that a 66‑million‑parameter DistilBERT can achieve competitive or superior performance on token‑level PICO classification compared to much larger models, while remaining lightweight. The authors validate the embedding‑level transfer with cosine similarity, semantic‑shift, and t‑SNE analyses, highlighting DRET’s potential for scalable, resource‑efficient biomedical text mining.
By Girish Sundaram, Daniel Berleant
arXiv:2605.28190v2 Announce Type: replace
Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embeddin...
By Manuel Frank, Haithem Afli
arXiv:2608.00042v2 Announce Type: replace-cross
Abstract: Domain adaptation of small language models (SLMs) has emerged as a practical strategy for deploying capable NLP systems in resource-constrain...
By Ramesh B. Paramkusham
arXiv:2606. 10392v1 Announce Type: new Abstract: Financial named-entity recognition (NER) is essential for translating unstructured financial reports and news into structured knowledge graphs.
By Wu Yuerong, Mingni Luo
Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deployment in resource-constrained or privacy-sensitive settings. We investigate how far small language models (SLMs) can close this gap across general-domain and literary text.
arXiv:2608. 09834v1 Announce Type: cross Abstract: Financial sentiment analysis converts unstructured financial news into quantitative signals that can support market analysis and decision-making.
By Fan Zhang, Jiaming Li
The article surveys fake review detection research, focusing on how pre‑trained language models (PLMs) and large language models (LLMs) influence both the generation of deceptive reviews and their detection. It reviews 211 studies from 2018 to early 2026, categorizing methods by evidence source—such as review text, sentiment, rating behavior, temporal metadata, user‑product graphs, multimodal content, external knowledge, and LLM‑generated signals—and by fusion level. The survey traces the evolution from traditional machine learning to PLM‑based and LLM‑based approaches, evaluates performance on Amazon, Yelp, and OpSpam benchmarks, and highlights open challenges including adversarial generation, cross‑domain transfer, uncertainty‑aware fusion, robustness to missing sources, interpretability, and trustworthy evaluation of AI‑generated deceptive content.
By Fanji Yang (Guizhou University of Finance and Economics), Huiyao Chen (Harbin Institute of Technology), Xi Yu (Guizhou University of Finance and Economics), Meishan Zhang (Harbin Institute of Technology), Xiaohong Xiao (Guizhou University of Commerce), Mingsen Deng (Guizhou University of Finance and Economics)