Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder
arXiv:2608. 19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science.
Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.
arXiv:2608. 19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science.
arXiv:2608. 19200v1 Announce Type: cross Abstract: Text summarization refers to the task of condensing a document into a shorter version while preserving its key information.
ContractScrub is a new benchmark that evaluates large language models on the task of contract scrubbing—reviewing legal agreements for errors and inconsistencies. The benchmark includes contracts crafted by experienced lawyers covering diverse error types such as misuse of defined terms, incorrect references, and inconsistent language. Initial tests show that even leading models perform poorly, with only one achieving a 0.75 macro‑average recall, highlighting the gap between general LLM performance and real‑world legal tasks.
Question Answering over Temporal Knowledge Graphs (TKGQA) requires reasoning over time-sensitive facts, yet existing embedding-based methods struggle with multi-step queries due to single-pass reasoning pipelines. We propose SABET-QA, a framework that iteratively refines reasoning states across multiple hops via a bidirectional entity-temporal scoring mechanism and a slot-aware contextualization module that aligns question semantics with temporal KG embeddings.
LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers. The deployment question...
Multimodal LLMs can see a document, but they often can't read it reliably. Small text, tables, visual cues, and topological elements still trip them up under direct visual inference, even when the pag...
NE‑BERT is a multilingual encoder trained on about 8.3 million sentences from nine Northeast Indian languages plus Hindi and English. Using weighted sampling and a custom SentencePiece tokenizer, it achieves significantly lower perplexity than IndicBERT‑V2, MuRIL, and mBERT, and improves tokenization fertility. The model also addresses vocabulary fragmentation in extremely low‑resource languages through aggressive upsampling, and its effectiveness is validated on part‑of‑speech tagging for three of the languages.
SuTRA (Structurally-Unified Tokenization with Root Awareness) is a morphology-aware tokenization algorithm designed to address the problem of Morphological Shattering in morphologically rich Indic languages. It preserves the indivisibility of aksharas—complex orthographic syllables—by penalizing merges that cross morphological boundaries, thereby reducing over-fragmentation of words. The authors also release a new morphological segmentation dataset for Hindi, Marathi, and Gujarati, and demonstrate that SuTRA improves morphological alignment by up to 14.7% and semantic recoverability by 34% over BPE, leading to an average machine translation gain of +8.08 chrF2.
EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.
DeepWeaver is a framework designed to improve open‑ended question answering by weaving noisy retrieved evidence into comprehensive, well‑cited answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and supporting evidence, and uses subordinate TBCs to refine and expand the evidence before final generation. Evaluations on LoQA and DeepResearch Bench show that DeepWeaver enhances content sufficiency, citation grounding, and detail preservation across multiple LLMs.
The paper discusses how recent progress in deep learning has improved Natural Language Processing and Vision‑Language Models, but also introduced new vulnerabilities, particularly backdoor attacks that threaten security. It explores methods for analyzing, detecting, and designing such attacks in both NLP and VLMs. Additionally, it proposes efficient multimodal representation techniques aimed at clinical and medical imaging applications.
The scoping review examines how artificial intelligence (AI) is applied across all stages of medication management in rural healthcare settings, from prescribing to post-administration monitoring. It identifies four main themes: the types of AI used, the medication phases impacted, the effectiveness in reducing errors, and rural-specific challenges such as infrastructure and alert fatigue. Studies show machine‑learning surveillance can cut prescribing and transcription errors by 34% to 80%, yet barriers like governance gaps, funding limits, and clinician resistance remain.
The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.
The study investigates how post‑training quantization (PTQ) affects proactive interference (PI) in large language models. Using bitsandbytes, the authors compare FP16, INT8, and INT4/NF4 precision across three instruction‑tuned models and find that INT4 quantization markedly degrades accuracy under high interference, with INT8 also incurring a smaller penalty in two of the three models. The degradation is linked to increased same‑key intrusion errors and originates in the quantized transformer backbone rather than the output layer.
DentAgent is an evidence‑centric multi‑agent framework designed for multimodal dental reasoning. It coordinates five specialized agents that process domain knowledge, radiographs, intraoral photographs, and 3D dental data, converting observations into structured evidence records. The Evidence Blackboard tracks coverage, gaps, and conflicts before generating responses, and the system outperforms senior specialists by 17.3 percentage points on multi‑label diagnosis across four benchmarks.
MITRE‑SAGE is a multi‑agent retrieval‑augmented generation framework that combines semantic and structural cybersecurity knowledge to enhance large language model question‑answering. It decomposes tasks into query interpretation, evidence retrieval, and answer synthesis, supporting vulnerability assessment, threat profiling, and relationship extraction. The authors also introduce MITRE‑QA, a benchmark of 3,000 question‑answer pairs, and show that MITRE‑SAGE outperforms standalone LLMs and conventional RAG methods, with a lightweight configuration achieving top performance on most tasks.
The paper introduces Fulcrum, a topology‑aware differential privacy scheme for hierarchical federated learning that allocates noise based on the size and exposure of regional aggregation groups. By deriving a closed‑form exposure dispersion metric from region structure and weights, the method optimally balances privacy and utility, achieving up to 14.84% accuracy gains on image tasks and 12.16% on text tasks at ε = 0.99 compared to uniform noise allocation. The approach ensures each participant receives noise commensurate with its actual exposure, eliminating unnecessary privacy overhead.
The paper presents the CDSP (context-conditional deliberation signal pipeline), which transforms investment committee meeting transcripts into structured predictive features. CDSP segments transcripts into topical chunks, assigns asset‑class context labels via a large language model, maps financial keywords to a taxonomy, and adds sentiment polarity and mention frequency features. Using these engineered features on 48 monthly meetings, the best model—combining sentence embeddings with CDSP features—achieves 73% accuracy and a 0.73 F1 score, outperforming a simple stock‑choice baseline, though the improvement is not statistically significant.
The paper proposes a systematic framework for creating a "Map of Datasets in Engineering Design and Systems Engineering" (EDSE) to address the fragmented and inaccessible nature of existing datasets. It introduces a multi‑dimensional taxonomy that classifies datasets by domain, lifecycle stage, data type, and format, and presents an interactive discovery tool built on a knowledge graph data model. The authors analyze the current data landscape, identify underrepresented areas such as early‑stage design and system architecture, and suggest strategies for curation and sustainability to build a dynamic, community‑driven resource.
The paper presents a method for extracting key information from OCR‑digitized clinical reports, addressing challenges posed by heterogeneous documents and noisy OCR output. It introduces an open key space that is iteratively mined, normalized, clustered, and verified to build a canonical key inventory, and defines key coverage as a metric for inventory completeness. Experiments on reports from over 20 hospitals using a 0.2B BERT model show that performance improves steadily with key coverage, achieving high F1 scores when the top 90 keys are covered and outperforming a fine‑tuned Qwen3‑0.6B baseline.