arXiv:2607. 28126v2 Announce Type: replace Abstract: Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles.
By Bingchen Liu, Yuanyuan Fang, Lei Liu, Guangyuan Dong, Xing Fu, Yuanyuan Gao, Shuyue Wei, Xin Li, Xiangtian Meng
Long-horizon steel-equipment inspection requires reasoning over heterogeneous records accumulated across repeated inspection cycles. Existing retrieval-augmented generation systems treat historical logs as a static corpus and retain records without estimating their diagnostic value, failing to report early risk.
PinSieve is a production system that selectively serves vision‑language models (VLMs) for enterprise content‑quality triage, operating only on the grey‑zone cases that lightweight models cannot resolve. The deployed VLM Serving Agent filters 2.05× more non‑actionable items, improves review productivity by 25.7%, cuts operating costs by 16.2%, and delivers signals the same day instead of the next. A governed memory flywheel with selective feedback, audit sampling, and a bounded proposal‑verifier loop further reduces false‑negative rates from 17.73% to 13.29% over six months, while a reasoning review agent audits teacher‑generated rationales for keep/repair/drop decisions.
whyItMatters":"The system demonstrates how selective VLM serving and governed feedback loops can substantially improve efficiency, cost, and accuracy in enterprise AI content‑quality pipelines."
By Chuqing Gao, Yuanfang Song, Jonathan Zhang, Yifan Wu, Vishwakarma Singh, Qinglong Zeng, Andrey Gusev
arXiv:2608. 02786v1 Announce Type: new Abstract: AI systems can fail silently.
By Priyanka Bajaj (Independent Researcher)
arXiv:2608. 19303v1 Announce Type: new Abstract: When a tool call times out, the agent sees the failure and can route around it.
By Sugam Panthi, Rabab Abdelfattah
arXiv:2608. 05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers.
By Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
The paper introduces CORA (Counterfactual, Observable Redundancy Audit), a protocol for auditing website redundancy by measuring repetition load, normal-use tax, and failure-domain recovery reserve. Each audit run records screenshots, stable element identities, and task traces, while a versioned vision‑language model generates annotations that are validated and released only if they meet calibrated criteria. Experiments on a transparent mechanistic testbed show that CORA’s factorized representation separates reserve from normal-use tax and predicts perturbed success more accurately than scalar-load baselines, but it withholds automated scores when instruments fail to meet release requirements, indicating that CORA is an auditable candidate procedure for the studied benchmark rather than a universal standard.
By Ge Kong, Yongtong Cao
SafeRestore introduces a framework for certifying when an industrial image restoration should be automatically returned to a detector or require human review. It ranks five restoration candidates using action‑specific fitted scores, selects a threshold gate on tuning data, and evaluates the gate on a separate certification sample with two one‑sided exact binomial bounds—one for evidence‑loss incidents and one for excess‑activation incidents. In a retrospective study of 4,591 Carinthia‑S images, the protocol demonstrates auditable risk‑coverage behavior, with varying pass rates across different policies and morphologies.
By Shaoliang Yang, Jun Wang
The paper introduces candidate‑fate accounting, an audit framework for transparent sensor diagnostic pipeline search that records every candidate, including invalid, pruned, or skipped ones, and assigns a terminal fate to each. It enhances traceability by hashing repeated observations, flagging illegal candidates, and documenting budget rationales. Experiments on three bearing‑diagnostic datasets demonstrate that the framework uncovers 30–41 omitted candidates and verifies complete accounting while preserving competitive performance.
By Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang
arXiv:2607. 12469v1 Announce Type: cross Abstract: Many agent-safety evaluation results are not yet load-bearing evidence: identical nominal outcomes (task success, attack success, monitor scores) may sit atop materially different evidence regimes.
By Oleg Solozobov
SIDScope is a diagnostic tool that evaluates Semantic‑ID interfaces used in generative recommendation systems. It normalizes item‑to‑code artifacts, verifies provenance, profiles mapping structure, and compares revisions while tracking path‑to‑item outcomes in generated traces. Using nine tokenizer exports from Amazon and Yelp data, SIDScope shows that interface health depends on multiple signals and reveals gaps in prefix alignment, trace accounting, and mapping refresh effects.
By Jiandong Ding, Huijie Qin, Tiandeng Wu, Yi Cao
arXiv:2609.08071v1 Announce Type: new
Abstract: Firms making inventory decisions have access to operational data, optimization tools, and large language models (LLMs). Typically, data characterize th...
By Fenghua Yang, Preet Baxi, Yi Zhang, Stefanus Jasin, Yanzhe Lei, Mo Liu, Parshan Pakiman