ReCast: Contract-Preserving Protection for Fixed-Interface Multimodal Reasoning
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
ReCast is a plug‑in framework that protects private inputs for fixed‑interface multimodal reasoning by locally converting them into a shared textual evidence‑query record, rewriting entities and topics with a distilled model, and mapping numerical values through an invertible, role‑aware map. A reconstruction agent then generates the required media from this protected record, allowing a remote solver to return a program whose operands are restored locally before execution. On 4,000 held‑out ChartQA and NMSQA examples, ReCast achieves 75.10% accuracy, retaining 92.43% of the unprotected remote accuracy, and flags source‑content leakage in 7.95% of solver‑bound requests, outperforming all evaluated local baselines.
arXiv:2606. 24623v1 Announce Type: cross Abstract: Retrieval-Augmented Generation enhances large language models by incorporating external knowledge, but deploying it in sensitive scenarios risks privacy leakage via malicious prompts.
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation.
arXiv:2607. 28374v1 Announce Type: new Abstract: Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy.
arXiv:2610. 01871v1 Announce Type: cross Abstract: Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of Multimodal Large Language Models (MLLMs) into relevant, up-to-date, external knowledge.
arXiv:2606. 04067v1 Announce Type: cross Abstract: As LLMs become increasingly woven into everyday workflows, user queries sent to cloud hosted LLMs routinely mix task-essential content with task non-essential sensitive disclosures, yet type based PII redaction is context agnostic and may raise two issues: over disclosing untyped sensitive context and over removing answer bearing spans.