arXiv Computation and Language

BLANC: Discovering Patent White Space via Changes in Normalized Pointwise Mutual Information Between Multi-View Clusters

BLANC (Blank Landscape Analysis through NPMI Conditioning) is a three‑phase pipeline that uses multi‑view neural topic modeling across application/use, novelty, and inventive step, computes Normalized Pointwise Mutual Information (NPMI) to measure cross‑dimensional cluster association, and introduces a conditional detection step that flags combinations whose NPMI drops when the corpus is filtered by a keyword. The drop is quantified by a new metric, ΔNPMI, which identifies combinations that are established globally but unexplored locally. BLANC was evaluated on two USPTO corpora—machine learning/AI and glass compositions—by artificially depleting known technology combinations; it recovered 34.1% and 27.3% of the depleted pairs, respectively, while random removals rarely recovered the target, and it successfully identified a fluorine surface‑treatment × warpage‑suppression candidate in a proprietary float‑glass case.

arXiv Computation and Language
Sep 23

ABAI at COLIEE 2026 Task 1: Multi-Stage Retrieval with GraphRAG-Enhanced Meta-Learning, and a Post-Hoc Study of the Cross-Validation-to-Test Gap

The paper reports the ABAI submission to COLIEE 2026 Task 1, a case law retrieval challenge that suppresses cited passages, and details a four‑stage retrieval pipeline: multi‑view BM25 with reciprocal rank fusion, neural reranking, graph‑based features via a graph attention network, and a LightGBM meta‑learner over 34 features. The best run achieved an F1 score of 0.177 on the official test set, compared to a cross‑validated 0.311, and the authors attribute the gap to a recall ceiling, temporal distribution shift, and threshold miscalibration. A controlled post‑hoc study examined the impact of threshold transfer, decision quality across time, and query similarity, and identified specific remedies—such as BM25 length‑normalisation tuning, event‑triple views, and dense fusion—that improved recall, while other interventions had no effect.

By Minhan Cho, Soyoung Park, Daejin Choi, Jinyoung Han
arXiv Computation and Language
Sep 3

HyGRAIL: Cost-Aware and Evidence-Grounded Scientific Hypothesis Discovery over Knowledge Graphs

HyGRAIL is a framework for discovering scientific hypotheses in incomplete knowledge graphs by combining a graph neural network (GNN) triage with large language model (LLM) review. The GNN scores candidate hypotheses and routes only ambiguous cases to the LLM, which receives structured evidence from the graph converted into natural language. Experiments on MatKG show HyGRAIL achieves the highest F1 score, improves over baselines, and cuts LLM calls by over 54%.

By Yihang Sun, Zhihan Zhu, Zhiyuan Jiang, Jingyi Ge, Zixuan Li, Jiaxuan You