arXiv:2509. 14474v3 Announce Type: replace Abstract: The debate around Artificial General Intelligence (AGI) remains open due to two fundamentally different goals: replicating human-level performance versus replicating human-like cognitive processes.
By Meltem Subasioglu, Nevzat Subasioglu
Benchy is a semantic language and execution engine designed to standardize task-oriented AI benchmarks. Each benchmark is fully defined by a program, a scoring function, and a dataset (B=(P,S,D)), and is independent of the AI system that runs it. Benchmarks are authored in canonical YAML, compiled deterministically into JSON, and executed via a universal runtime contract that exposes a named-field input object and a named-field output object, ensuring consistent integration across AI systems.
By Francis F Daniel, Mauro Iba\~nez, Francis Perelman, Marian Basti
arXiv:2606. 10106v1 Announce Type: cross Abstract: The term agent harness now circulates widely in software engineering with generative artificial intelligence.
By Sanderson Oliveira de Macedo
arXiv:2608. 20201v1 Announce Type: new Abstract: Software form has undergone two paradigm shifts since its inception: Software 1.
By Wei Lin, Tao Zhou, Zhaofei Xie, Changgui Hong
arXiv:2606. 05608v2 Announce Type: replace-cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
By Zhenfeng Cao
arXiv:2605. 22093v3 Announce Type: replace Abstract: Knowledge graphs have become the primary vehicle for data integration and are critical to the success of modern AI, but the diversity of KG modelling practices, from lightweight vocabularies to richly axiomatised ontologies, makes integration and reuse expensive and brittle.
By Enrico Daga, Valentina Tamma, Terry Payne
The paper proposes a four‑dimensional formal framework—Semantic Expressivity, Agentic Discoverability, Task‑Relative Grounding, and Epistemic Trust Scope—to extend current KG metadata standards (VoID and DCAT). It introduces the Agentic Affordance Profile (AAP), a semantic layer that enables agents to select, compose, and diagnose failures in knowledge graphs at planning time. A scholarly‑search example illustrates the framework and outlines a five‑point research agenda for scaling AAP‑based affordance matching.
By Terry R. Payne, Valentina Tamma, Enrico Daga
arXiv:2608.29311v1 Announce Type: new
Abstract: Classic Formal Concept Analysis (FCA) primarily focuses on the positive relationships between objects and attributes and does not have mechanisms for h...
By Zhenghua Pan
arXiv:2602. 11198v2 Announce Type: replace-cross Abstract: Multi-agent frameworks (MAFs) promise to simplify LLM-driven software development, yet no principled metric captures how well AI coding assistants can generate correct, framework-specific code.
By Shafiuddin Rehan Ahmed, Sourabh Deshpande
The paper critiques the Semantic Web’s failure to deliver machine‑interpretable knowledge, arguing that its standards omitted key elements—conditions for claims, operational grounding, and coverage scope—making truth, applicability, and boundary recognition impossible. It proposes a new framework, Semantic Knowledge Technologies, with a seven‑layer architecture and five measurable tests of understanding (check, connect, derive, act, delimit). The authors introduce concepts such as Large Knowledge Models, SLKMs, and a falsifiable definition of Semantic Artificial General Intelligence, presenting a research agenda to address these gaps.
By Achille Zappa
arXiv:2606. 05608v1 Announce Type: cross Abstract: For over half a century, software engineering has operated on a foundational premise: human engineers decompose problems, encode decision logic into static code, and manually adapt that code as requirements evolve.
By Zhenfeng Cao
The paper discusses a third paradigm shift in software development, termed Software 3.0, where context and reasoning drive behavior. It proposes that Software 3.0 converges to three core components: a generalized database for all persistent state, a large model that performs reasoning and generation, and an agent that orchestrates the interaction between the two. The authors formalize this convergence, present a minimal reference architecture, and analyze its applicability and limits, noting that it applies best to task domains that are expressible, verifiable, externally stateful, and tool-complete.