arXiv Machine Learning

hoBIT: A Profile-Aware Retrieval-Augmented Chatbot for University Academic Advising

The paper introduces proFILL, a method that transforms the existing rule-based advising chatbot hoBIT into a profile-aware retrieval-augmented generation system. Instead of needing a full student profile at the start, proFILL incrementally gathers only the attributes required for each query, using both the query intent and initially retrieved evidence to guide retrieval from a profile-aware index. Experiments and a human preference study demonstrate that proFILL outperforms various RAG baselines, is favored by users, and remains effective when deployed with open-weight models for cost-efficient on‑premise use.

arXiv Machine Learning
Sep 4

Comparing Retrieval Methods for Academic Advisor Discovery: A Six-Method Study of 768 CS Faculty Profiles Across 9 US Universities

The study evaluates six retrieval methods for ranking computer science faculty as potential academic advisors based on graduate applicants’ research interest statements. Using a new dataset of 768 faculty profiles from nine U.S. universities and 162 graded relevance judgments across five queries, the reranked hybrid approach achieved the highest mean NDCG@10 (0.477). Ablation experiments showed that faculty biographies alone outperform the full model, and adding arXiv abstracts actually decreased performance, leading to a late‑fusion design. All code, data, and labels are publicly released.

By Biraj Subedi
arXiv AI
Jun 2

NBQ: Next-Best-Question for Dynamic Profiling

arXiv:2606. 00809v1 Announce Type: new Abstract: Many real-world conversational settings for knowledge discovery, including podcasts, hiring screens, and marketplaces, require a purpose-driven understanding of a person.

By Yimin Shi, Clarice Wang, Haixun Wang, Xiaokui Xiao
arXiv Computation and Language
Sep 4

Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent

The paper introduces PersonaLink, a training‑free method that distills a user’s interaction history into a bounded three‑field persona and iteratively refines it by self‑evaluating a frozen 7B language model on held‑out labeled data. Each refinement rewrites the persona only if it does not regress on that slice, ensuring the persona remains bounded and query‑independent. On a 200‑user news categorization task (LaMP‑2), PersonaLink achieves 0.745–0.755 accuracy, statistically indistinguishable from BM25 retrieval’s 0.760–0.765 accuracy, demonstrating that distilled personas can match retrieval for classification but not for regression tasks.

By JaeHa Yoon, Minjun Park, Seoyeon Kim, Jiwoo Lee, Hyunwoo Choi, Dohyun Kang
arXiv AI
Sep 10

Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models

The paper introduces ENOUGH, a method for creating minimal sufficient user profiles for personalized language models. ENOUGH iteratively adds behavioral records or stops, evaluating profile prefixes with a counterfactual search that balances downstream gains, user specificity, and token costs. The resulting profiles are distilled into a lightweight controller that orders records and triggers the generator only once, achieving better effectiveness and efficiency than existing baselines across six tasks.

By Minghang Liu, Qiang Qiu, Yuanzhuo Wang, Huawei Shen, Xueqi Cheng
arXiv AI
Jul 28

PeopleSearchBench: A Multi-Dimensional Benchmark for Evaluating AI-Powered People Search Platforms

arXiv:2603. 27476v2 Announce Type: replace Abstract: AI-powered people search platforms are increasingly used in recruiting, sales prospecting, and professional networking, yet no widely accepted benchmark exists for evaluating their performance.

By Wei Wang, Tianyu Shi, Shuai Zhang, Boyang Xia, Zequn Xie, Chenyu Zeng, Qi Zhang, Lynn Ai, Yaqi Yu, Kaiming Zhang, Feiyue Tang, Lei Ding
arXiv AI
2d ago

Scientific Agents: Evaluating Profession-Specific System Prompts on Scientific Tasks

The study evaluates whether detailed, profession‑specific system prompts improve performance on scientific tasks. Using an open‑source corpus of 503 agent profiles and Gemini 3.8 Flash, the authors compared matched profiles to four control prompts across nine text‑based science benchmarks and a tool‑using bioinformatics benchmark. Results show no consistent accuracy gains; matched profiles actually increased token usage and cost, and in some cases reduced success rates, with only a minor advantage in one benchmark likely due to prompt length rather than domain expertise.

By Timothy Kassis