arXiv AI

Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics

Semantic Signal-Assisted Inspection and Recovery Allocation in Reverse Logistics proposes a decision-support framework that transforms return notes into a condition factor and a signal‑quality score. These metrics guide how deeply to inspect returned assets and how to allocate them for recovery, balancing labor constraints. In synthetic benchmarks across IT decommissioning, aircraft maintenance, and consumer‑electronics returns, the keyword‑based implementation outperforms a structured‑feature comparator by improving net recovery value and lowering inspection costs, with the greatest economic benefit observed in the aircraft scenario.

arXiv Machine Learning
Aug 26

PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage

PinSieve is a production system that selectively serves vision‑language models (VLMs) for enterprise content‑quality triage, operating only on the grey‑zone cases that lightweight models cannot resolve. The deployed VLM Serving Agent filters 2.05× more non‑actionable items, improves review productivity by 25.7%, cuts operating costs by 16.2%, and delivers signals the same day instead of the next. A governed memory flywheel with selective feedback, audit sampling, and a bounded proposal‑verifier loop further reduces false‑negative rates from 17.73% to 13.29% over six months, while a reasoning review agent audits teacher‑generated rationales for keep/repair/drop decisions. whyItMatters":"The system demonstrates how selective VLM serving and governed feedback loops can substantially improve efficiency, cost, and accuracy in enterprise AI content‑quality pipelines."

By Chuqing Gao, Yuanfang Song, Jonathan Zhang, Yifan Wu, Vishwakarma Singh, Qinglong Zeng, Andrey Gusev
arXiv AI
Aug 7

SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents

arXiv:2608. 05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers.

By Zhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
arXiv Computer Vision
Aug 25

From Subjective Judgments to Auditable Standards:Protocol-Guided AI Auditing of Website Redundancy

The paper introduces CORA (Counterfactual, Observable Redundancy Audit), a protocol for auditing website redundancy by measuring repetition load, normal-use tax, and failure-domain recovery reserve. Each audit run records screenshots, stable element identities, and task traces, while a versioned vision‑language model generates annotations that are validated and released only if they meet calibrated criteria. Experiments on a transparent mechanistic testbed show that CORA’s factorized representation separates reserve from normal-use tax and predicts perturbed success more accurately than scalar-load baselines, but it withholds automated scores when instruments fail to meet release requirements, indicating that CORA is an auditable candidate procedure for the studied benchmark rather than a universal standard.

By Ge Kong, Yongtong Cao
arXiv Computer Vision
Sep 4

SafeRestore: Detector-Relative Risk Certificates for Selective Industrial Image Restoration

SafeRestore introduces a framework for certifying when an industrial image restoration should be automatically returned to a detector or require human review. It ranks five restoration candidates using action‑specific fitted scores, selects a threshold gate on tuning data, and evaluates the gate on a separate certification sample with two one‑sided exact binomial bounds—one for evidence‑loss incidents and one for excess‑activation incidents. In a retrospective study of 4,591 Carinthia‑S images, the protocol demonstrates auditable risk‑coverage behavior, with varying pass rates across different policies and morphologies.

By Shaoliang Yang, Jun Wang
arXiv AI
Aug 20

Candidate-Fate Accounting for Transparent Sensor Diagnostic Pipeline Search

The paper introduces candidate‑fate accounting, an audit framework for transparent sensor diagnostic pipeline search that records every candidate, including invalid, pruned, or skipped ones, and assigns a terminal fate to each. It enhances traceability by hashing repeated observations, flagging illegal candidates, and documenting budget rationales. Experiments on three bearing‑diagnostic datasets demonstrate that the framework uncovers 30–41 omitted candidates and verifies complete accounting while preserving competitive performance.

By Haotao Xie, Yutian Chen, Yangqi Liu, Xiaoyu Jiang
arXiv AI
Aug 20

SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation

SIDScope is a diagnostic tool that evaluates Semantic‑ID interfaces used in generative recommendation systems. It normalizes item‑to‑code artifacts, verifies provenance, profiles mapping structure, and compares revisions while tracking path‑to‑item outcomes in generated traces. Using nine tokenizer exports from Amazon and Yelp data, SIDScope shows that interface health depends on multiple signals and reveals gaps in prefix alignment, trace accounting, and mapping refresh effects.

By Jiandong Ding, Huijie Qin, Tiandeng Wu, Yi Cao