Tuesday, June 30, 2026

Daily Digest 2026-06-30

Today’s digest is characterized by a heavy focus on agentic autonomy, specifically regarding memory architectures, safety alignment, and the evaluation of complex multi-agent collaboration.

Research highlights:

  • Agentic Autonomy and Memory: Research explores how agents can evolve recursively, manage complementary memory systems for vision-language tasks, and utilize living knowledge topologies.
  • Medical AI and Reasoning: New frameworks are being developed to evaluate doctor agents in clinical simulations and to benchmark multimodal models in image-grounded medical conversations.
  • Safety and Alignment: Studies investigate agentic abstention (knowing when to stop), action alignment as a safety metric, and the ethical profiling of models through virtue-based dilemmas.
  • Reinforcement Learning and Optimization: Work includes uncertainty-weighted baselines for critic-free RL and dynamic representation editing to steer reasoning trajectories toward truth.
  • Multimodal Interaction: New benchmarks and guidance systems are being established to improve real-time collaboration and intent grounding in multimodal models.

Tech buzz:

  • The industry is seeing a push toward massive-scale Mixture-of-Experts (MoE) models and significant infrastructure investment in hardware.
  • Large-Scale MoE: The release of LongCat-2.0 introduces a model with 1.6T total parameters and 48B active parameters.
  • Hardware Investment: South Korea announced a $1T investment plan targeting memory chip production and humanoid robotics.
  • Security and Governance: New research highlights vulnerabilities in multi-turn prompt injection defenses, while discussions continue regarding the long-term impact of government gatekeeping on frontier models.
Sort:
Today's digest is characterized by a heavy focus on agentic autonomy, specifically regarding memory architectures, safety alignment, and the evaluation of complex multi-agent collaboration.

Papers discovered from ArXiv subject categories

AI Safety

5/5 Artificial Intelligence (cs.AI) 30 Jun 2026
Agent Safety Is Action Alignment

Shawn Li, Yue Zhao

Abstract

ArXiv ID: 2606.28739

Authors: Shawn Li, Yue Zhao

Abstract:

Large language models increasingly act as agents: they call tools, move money, delete records, and send messages on a user's behalf. To keep them safe, practitioners imported the chatbot-era recipe (train the model to refuse unsafe inputs) into the agentic setting, and treat the resulting capability loss as a manageable ``alignment tax.'' We argue this is a \emph{category error}. Refusal is a primitive for \emph{content safety}, where the harm is in the model's output and is therefore a learnable function of it. Agentic harm is different in kind: it lies not in any output but in the relation between the authority an action exercises and the authority the user granted, which is absent from the text the model sees. Importing content-safety methods into this regime does not trade capability for safety; it pays capability and buys negative security. We support this with three lines of evidence spanning the autonomy spectrum: defense-trained models learn surface patterns rather than intent; the same training collapses multi-step agents before any threat appears while leaving them exploitable; and even undefended frontier models exceed granted authority under ordinary use. We conclude that action safety cannot be installed in weights. It must be expressed as \emph{least privilege}, enforced \emph{outside} the model at the action boundary, and evaluated as \emph{action alignment} (a relational, deployment-conditioned property) rather than a refusal score.

Insights

Contribution: The paper argues that agent safety is a relational property of action alignment rather than a content-safety problem, asserting that safety cannot be 'installed' in model weights via refusal training.

Core Idea: Agentic harm stems from the discrepancy between granted and exercised authority, making traditional refusal-based safety measures ineffective and counterproductive for autonomous agents.

Technique: The authors propose a shift from weight-based refusal training to external enforcement of the principle of least privilege at the action boundary.

Pipeline: User request β†’ Agentic action request β†’ External authority check (Least Privilege) β†’ Executed action (or blocked)

Methodology: The authors provide three lines of evidence: analyzing surface pattern learning in defense-trained models, evaluating multi-step agent collapse, and observing authority overreach in undefended frontier models.

Results: Defense-trained models learn surface patterns rather than intent; refusal training collapses multi-step agents before threats appear; and undefended frontier models exceed granted authority during ordinary use.

Limitations: The paper focuses on the conceptual and architectural shift toward external enforcement but does not provide a specific technical framework for the external monitoring layer.

PDF
4/5 Artificial Intelligence (cs.AI)cs.CYstat.APstat.ME 30 Jun 2026
Aristotelian Virtue Profiling of LLMs through Ethical Dilemmas

Ioannis Tzachristas, John Pavlopoulos

Abstract

ArXiv ID: 2606.28683

Authors: Ioannis Tzachristas, John Pavlopoulos

Abstract:

Large Language Models (LLMs) often face ethical tradeoffs in which several responses may be defensible but express different priorities, such as fairness, honesty, courage, or restraint. We introduce VirtueMap, a framework for describing these patterns through an Aristotelian virtue-ethics lens. Instead of asking for a single correct answer, VirtueMap asks humans or LLMs to rank all five responses to each of seven general, non-lethal, non-political, and non-religious ethical dilemmas. To define the reference orderings used for scoring, we first proposed, for each dilemma and virtue, an ordering of the five responses from most to least expressive of that virtue. We then collected more than 100 respondent evaluations per ordering and retained it as operational ground truth only when at least 95% confirmed it. Rankings are scored against these retained orderings using normalized Borda alignment, yielding profiles over Practical Wisdom, Justice, Truthfulness, Courage, and Temperance. We apply VirtueMap to nine LLM families in a repeated-run evaluation and find high mean rank consistency (90.3%), with the largest differences appearing on Courage, Temperance, and Justice. We also release an interactive website that computes profiles locally in the browser and compares respondents with measured LLM profiles.

Insights

Contribution: The paper introduces VirtueMap, a framework that profiles the ethical character of Large Language Models (LLMs) using Aristotelian virtue ethics. It provides a systematic way to measure how LLMs prioritize different moral values like justice, courage, and temperance across various dilemmas.

Core Idea: Instead of evaluating LLMs on a binary 'correct' or 'incorrect' basis, the research treats ethical responses as a spectrum of competing virtues. It aims to map the 'moral personality' of an AI by analyzing its preference rankings.

Technique: The authors use a ranking-based evaluation system where multiple defensible responses are scored against human-validated 'ground truth' orderings for specific virtues. They employ normalized Borda alignment to calculate the alignment between LLM rankings and these virtue-specific benchmarks.

Pipeline: Ethical dilemmas β†’ Multiple defensible responses β†’ Human-validated virtue orderings β†’ LLM response ranking β†’ Normalized Borda alignment β†’ Virtue profiles (Practical Wisdom, Justice, Truthfulness, Courage, Temperance)

Methodology: The researchers defined five virtues and seven ethical dilemmas, establishing ground truth orderings through high-consensus human evaluations (95%+ agreement). They then tested nine LLM families to see how closely their rankings aligned with these virtue-specific benchmarks.

Results: The study found high mean rank consistency (90.3%) across LLMs, with the most significant variations in how models express Courage, Temperance, and Justice.

Limitations: The study focuses on non-lethal, non-political, and non-religious dilemmas, which may not capture the most complex or high-stakes ethical challenges faced by AI in the real world.

PDF
4/5 Artificial Intelligence (cs.AI)Computer Science and Game Theory (cs.GT) 30 Jun 2026
The Two Genie Game: Adoption and Welfare in Audit-Grounded AI Governance

Darrell Lewis-Sandy

Abstract

ArXiv ID: 2606.28710

Authors: Darrell Lewis-Sandy

Abstract:

We ask under what conditions an agent with a harm-minimizing policy can displace an approval-seeking (RLHF) agent in a competitive market, and when that policy is sufficient to prevent community harm. We use evolutionary game theory (finite-population Moran-Fermi pairwise comparison) to formalize this subject to assumptions of wisher hindsight, peer testimony, a monotone harm ledger, sufficient information density of community feedback, and a finite, depleting resource pool, in a negative-sum environment. We show that adoption is favored when the prior distributions on how readily wishers attune to community sentiment are monotone, exhibit endpoint inversion, and have a centro-symmetric pairing property, and demonstrate this with several long-tailed priors (Hill, Pareto, Lomax, Frechet). Where it is favored, a critical adoption level separates communities that drift back to the approval-seeking agent from those for which the audited agent fixes; above that level fixation is the overwhelmingly likely outcome. We derive when fixation is attainable as a bound on the effective (informational) size N_c of the community, which must be small enough to allow fixation before depletion. We present these as Theorems 5.4 and 5.5; the algebraic and finite-grid backbone is machine-checked in Lean 4, with the barrier-crossing asymptotics retained as explicit hypotheses. We show that a self-audited agent with a community ledger is not, in general, sufficient to prevent community harm. Sufficiency depends both upon the alignment of the agent's audit with community values and the timeframe over which harm is evaluated. Regardless of alignment, once adoption reaches dominance, the state is absorbing. The same policy that reduced harm under alignment becomes a trap, welfare-negative under misalignment and, even under alignment, one that locks in harm deferred past the adoption horizon.

Insights

Contribution: The paper provides a formal game-theoretic framework to determine when harm-minimizing AI agents can displace approval-seeking (RLHF) agents in a competitive market and identifies the conditions under which such adoption leads to permanent community harm.

Core Idea: The research explores the 'Two Genie Game' where the success of an audited agent depends on community feedback density, the distribution of user attitudes, and the finite nature of resources in a negative-sum environment.

Technique: The authors employ evolutionary game theory using a finite-population Moran-Fermi pairwise comparison model, with algebraic proofs machine-checked in Lean 4.

Pipeline: Community feedback and agent policies β†’ Moran-Fermi evolutionary dynamics β†’ Adoption thresholds and welfare outcomes

Methodology: The study models agent competition under assumptions of monotone harm ledgers and wishers' hindsight, testing various long-tailed priors (Hill, Pareto, Lomax, Frechet) to find adoption boundaries.

Results: Adoption is favored under specific monotone, endpoint-inverted, and centro-symmetric priors; a critical adoption level exists where fixation becomes overwhelmingly likely; however, self-audited agents can create 'traps' that lock in deferred harm or welfare-negative states under misalignment.

Limitations: The sufficiency of harm prevention depends heavily on the alignment of the audit with community values and the specific timeframe over which harm is evaluated.

PDF

Agentic AI

5/5 Artificial Intelligence (cs.AI) 30 Jun 2026
Recursive Self-Evolving Agents via Held-Out Selection

Michael Nguyen, Quoc Nguyen, Paul Vuong

Abstract

ArXiv ID: 2606.28374

Authors: Michael Nguyen, Quoc Nguyen, Paul Vuong

Abstract:

LLM agents are increasingly improved without weight updates by evolving a natural-language artifact, such as reflections, workflows, playbooks, cheatsheets, or optimized prompts, that conditions a frozen policy. Such methods are typically reported as wins on the single benchmark where they help. We study them apples-to-apples and surface a sharper picture. We introduce RSEA, a Recursive Self-Evolving Agent that carries a compact three-layer natural-language state: an imperative strategy, reusable skills, and a procedural playbook. Across generations, RSEA rewrites all three layers from its own trajectories and commits a candidate only if it does not regress on a disjoint held-out split, using a strict keep-better gate. Across four diverse benchmarks, ALFWorld, GAIA, (\tau)-bench, and WebShop, and six faithful baselines, ReAct, Reflexion, GEPA, AWM, ACE, and Dynamic Cheatsheet, all evaluated on one shared local backbone, we find three main results. First, no artifact universally wins. RSEA is the strongest single-pass method on ALFWorld, reaching 69.3% compared with 64.6% for ReAct (McNemar (p=0.015)), and reaches 79.4% with retry, the best overall result. However, concrete-workflow induction, represented by AWM, is best on the strong-backbone tool-use tasks. Second, unguarded context evolution is high-variance and unsafe. Dynamic Cheatsheet, which curates context online without a held-out gate, is near-best on ALFWorld at 70.7%, yet collapses on WebShop, with a score of 0.14 compared with 0.43 for ReAct. Third, RSEA's strict held-out selection is what makes recursive self-evolution monotone-safe: it never significantly underperforms the base agent on any benchmark and falls back to vanilla ReAct when evolved context would hurt.

Insights

Contribution: The paper introduces RSEA, a framework for recursive self-evolution of LLM agents that uses a held-out selection mechanism to ensure monotonic improvement. It provides a systematic, apples-to-apples evaluation of various natural-language artifact evolution methods across multiple benchmarks.

Core Idea: Agents can improve without weight updates by evolving a three-layer natural-language state (strategy, skills, and playbook) based on their own trajectories. To prevent performance degradation, new artifacts are only committed if they pass a strict 'keep-better' gate on a disjoint held-out dataset.

Technique: Recursive Self-Evolving Agents (RSEA) utilize a three-layer state architecture and a held-out selection gate to filter out regressive updates during the evolution process.

Pipeline: Agent Trajectories β†’ Rewrite Strategy, Skills, and Playbook β†’ Evaluate on Held-Out Split β†’ Keep-Better Gate β†’ Updated Agent State

Methodology: The authors compared RSEA against six baselines (e.g., ReAct, Reflexion, Dynamic Cheatsheet) across four diverse benchmarks (ALFWorld, GAIA, (Ο„)-bench, and WebShop) using a shared local backbone.

Results: RSEA achieved the best overall result on ALFWorld (79.4% with retry) and proved monotone-safe by never significantly underperforming the base agent. Conversely, unguarded context evolution (Dynamic Cheatsheet) showed high variance, performing well on ALFWorld but collapsing on WebShop.

Limitations: The study notes that no single artifact evolution method universally wins across all tasks, as specific methods like AWM still outperform RSEA on certain strong-backbone tool-use tasks.

PDF
5/5 Artificial Intelligence (cs.AI)Computation and Language (cs.CL) 30 Jun 2026
GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

Amit Parekh, Sabrina McCallum, Kareem Al-Hasan, Malvina Nikandrou, Alessandro Suglia, Ioannis Konstas

Abstract

ArXiv ID: 2606.28514

Authors: Amit Parekh, Sabrina McCallum, Kareem Al-Hasan, Malvina Nikandrou, Alessandro Suglia, Ioannis Konstas

Abstract:

Multimodal models are increasingly deployed to solve tasks collaboratively with humans or other artificial agents. Existing benchmarks show that these models possess many of the required component capabilities, but the conditions that coincide in collaboration, including time pressure, information asymmetry, and imperfect communication, are usually studied in isolation. We introduce GPTNT, a benchmark built on the cooperative video game Keep Talking and Nobody Explodes, in which two agents must coordinate to defuse procedurally generated bomb puzzles against a live countdown. One agent can see and manipulate the bomb but does not have the defusal instructions; the other has the instructions but cannot see or manipulate the bomb. Neither agent can succeed alone: success requires effective and efficient communication. Unlike turn-based proxies, GPTNT requires agents to act asynchronously and communicate in real time. GPTNT is designed to separate collaboration from reliance on memorized solutions: the instruction manual, the partner, or both can be withheld to isolate what a model derives in the moment from what it already knows. We show that GPTNT poses a substantial challenge for state-of-the-art systems: none of the closed- or open-source models we test defuses a single bomb in real time, a bar that human players clear. Through controlled experiments, we identify critical weaknesses in state tracking, efficient action under time pressure, ambiguity handling, and error recovery. We release GPTNT as a benchmark for collaborative performance that current evaluations leave unmeasured. Because it runs on the real game, GPTNT benefits from procedural generation and inherits a living modding community, allowing the benchmark to evolve as models improve rather than being solved once and retired.

Insights

Contribution: The paper introduces GPTNT, a new benchmark for evaluating real-time, multimodal collaboration between AI agents under conditions of time pressure and information asymmetry.

Core Idea: The authors use the game 'Keep Talking and Nobody Explodes' to create a scenario where two agents must coordinate asynchronously to solve puzzles, as neither agent possesses all the necessary information to succeed alone.

Technique: The benchmark utilizes a procedural game environment to isolate real-time communication and state tracking from memorized solutions by withholding specific information from either agent.

Pipeline: Procedurally generated bomb puzzle β†’ Asynchronous multimodal agent interaction (visual/textual) β†’ Real-time communication and action execution β†’ Success/Failure evaluation

Methodology: The researchers tested various closed- and open-source models in a live game environment, systematically withholding the manual or the partner to isolate specific collaborative capabilities.

Results: None of the tested state-of-the-art models successfully defused a single bomb in real time, revealing critical weaknesses in state tracking, ambiguity handling, and error recovery.

Limitations: The benchmark highlights current model failures in high-pressure, real-time environments, but the specific architectural improvements needed to bridge this gap remain an open question.

PDF
5/5 Artificial Intelligence (cs.AI) 30 Jun 2026
An AI agent for treatment reasoning over a biomedical tool universe

Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik

Abstract

ArXiv ID: 2606.28692

Authors: Shanghua Gao, Ayush Noori, Richard Zhu, Curtis Ginder, Zhenglun Kong, Xiaorui Su, Justin Kauffman, Benjamin S. Glicksberg, Joshua Lampert, Ankit Sakhuja, Ashwin Sawant, ATHENA-R1 Evaluation Consortium, David A. Clifton, Noa Dagan, Ran Balicer, Marinka Zitnik

Abstract:

Treatment reasoning underpins every therapeutic decision, integrating disease context, comorbidities, medications, contraindications, and evolving biomedical knowledge to select an appropriate therapy. It is inherently iterative: candidates are weighed against many constraints, revised as evidence emerges, and grounded in verifiable sources. Here we introduce ATHENA-R1, an AI agent for treatment reasoning across all FDA approved drugs since 1939, trained by reinforcement learning over a universe of 212 biomedical tools. At each step it identifies missing information, selects and runs relevant tools, and incorporates the evidence. To train it without human-annotated traces, we build a two-level self-learning framework: multi-agent systems construct the tools, tasks, and reasoning trajectories for supervised fine-tuning, then reinforcement learning with scientific feedback rewards reasoning quality (evidence gathering, grounded tool use, logical non-redundancy). Across five benchmarks of 3,168 drug reasoning tasks and 456 patient treatment cases, ATHENA-R1 outperforms language models and tool-use systems, reaching 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning, 17.8 and 10.7 points above GPT-5. In blinded evaluations by experts from 28 rare disease organizations, it is preferred over reference models on all criteria, and physicians rated it favorably on complex hospitalized cardiovascular and infectious-disease cases. Adverse-event hypotheses it generated, tested in electronic health records from 5.4 million patients, reached adjusted odds ratios of 1.48-1.84, with no elevation among negative controls. Because it requires knowing what evidence to seek before concluding, treatment reasoning has long been hard for AI; we show it can be reframed as a learnable process of iterative evidence gathering that reinforcement learning can train AI to perform.

Insights

Contribution: The paper introduces ATHENA-R1, an AI agent capable of complex treatment reasoning across all FDA-approved drugs by iteratively gathering and synthesizing evidence from a vast universe of biomedical tools.

Core Idea: Treatment reasoning is reframed as a learnable process of iterative evidence gathering, where an agent identifies missing information, selects relevant tools, and updates its reasoning based on retrieved evidence.

Technique: The authors employ a two-level self-learning framework using multi-agent systems to generate synthetic reasoning trajectories for supervised fine-tuning, followed by reinforcement learning with scientific feedback rewards.

Pipeline: Patient/Drug Context β†’ Identification of missing information β†’ Tool selection and execution β†’ Evidence synthesis β†’ Iterative reasoning refinement β†’ Final treatment recommendation

Methodology: The model was trained on 212 biomedical tools using a self-learning framework that generates its own training data and uses scientific feedback rewards to optimize for evidence gathering and logical non-redundancy.

Results: ATHENA-R1 achieved 94.7% accuracy on open-ended drug reasoning and 82.9% on treatment reasoning, significantly outperforming GPT-5 by 17.8 and 10.7 points respectively, while also generating clinically relevant adverse-event hypotheses validated against 5.4 million EHR records.

Limitations: The paper does not explicitly detail the specific constraints of the 212 tools or the potential for hallucinations in the multi-agent synthetic data generation process.

PDF
5/5 Artificial Intelligence (cs.AI) 30 Jun 2026
Agentic Abstention: Do Agents Know When to Stop Instead of Act?

Han Luo, Bingbing Wen, Lucy Lu Wang

Abstract

ArXiv ID: 2606.28733

Authors: Han Luo, Bingbing Wen, Lucy Lu Wang

Abstract:

LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available environment. In such cases, a reliable agent should recognize that further interaction is unlikely to help and abstain from additional tool calls. We define Agentic Abstention, the problem of deciding when an agent should stop acting under uncertainty. Unlike standard LLM abstention, which is usually evaluated as a single-turn answer-or-abstain decision, agentic abstention is a sequential decision problem: an agent can answer, abstain, or gather more information at each turn, and the need to abstain may only become clear after interacting with the environment. We study this problem across web shopping, terminal environments, and question answering, evaluating 13 LLM-as-agent systems and 2 agent scaffolds on more than 28,000 tasks. Our results show that the main challenge is not only whether agents can abstain, but also when they abstain. Some agents never abstain when they should, while others do so only after many unnecessary interactions. This gap is especially large on tasks where the instruction appears feasible until the environment reveals otherwise (e.g., no valid result matches the instruction). We further find that model scale, reasoning, and agent scaffolding affect abstention in different ways, where larger or more capable models sometimes perform worse at timely abstention. Finally, we introduce CONVOLVE, a context engineering method for improving agentic abstention that distills full interaction trajectories into reusable stopping rules. On WebShop, CONVOLVE substantially improves timely abstention without updating model parameters, raising Llama-3.3-70B's timely recall rate from 26.7 to 57.4. Our dataset and code are available at https://lhannnn.github.io/agentic-abstention

Insights

Contribution: The paper introduces the concept of 'Agentic Abstention' as a sequential decision problem and provides a comprehensive evaluation across 13 LLM-as-agent systems. It also proposes CONVOLVE, a context engineering method to improve timely abstention without fine-tuning.

Core Idea: Unlike single-turn abstention, agentic abstention requires an agent to decide whether to continue interacting with an environment, gather more information, or stop because a goal is unachievable. The research highlights that the primary challenge is 'timely' abstentionβ€”stopping before wasting unnecessary turns.

Technique: The authors developed CONVOLVE, a context engineering method that distills full interaction trajectories into reusable stopping rules to guide the agent's decision-making.

Pipeline: User goal and environment state β†’ Agentic decision (Act, Gather Info, or Abstain) β†’ CONVOLVE-distilled stopping rules β†’ Timely abstention or task completion.

Methodology: The researchers evaluated 13 LLM-as-agent systems and 2 scaffolds on over 28,000 tasks across web shopping, terminal, and QA environments. They measured the gap between 'correct' abstention and 'timely' abstention.

Results: Larger models sometimes perform worse at timely abstention; CONVOLVE significantly improved Llama-3.3-70B's timely recall rate on WebShop from 26.7% to 57.4%.

Limitations: The study notes that the gap in timely abstention is largest on tasks that appear feasible until environment interaction reveals otherwise, suggesting a need for better handling of latent environmental constraints.

PDF
5/5 Artificial Intelligence (cs.AI)Multiagent Systems (cs.MA) 30 Jun 2026
HyphaeDB: A Living Knowledge Topology for Agent-First Memory

Krishna Halaharvi

Abstract

ArXiv ID: 2606.28781

Authors: Krishna Halaharvi

Abstract:

Every existing vector database and agent memory framework treats memory as passive storage that agents query explicitly. No system propagates knowledge between agents through the memory layer itself. We introduce HyphaeDB, an agent-native memory infrastructure that reinterprets the Hierarchical Navigable Small World (HNSW) graph topology the data structure at the core of every modern vector database not as a search optimization, but as a communication fabric for multi-agent AI systems. In HyphaeDB, agents are nodes in the vector space with persistent positions, knowledge propagates via a gossip protocol through the graph's neighbor structure with energy-based attenuation, and emergent behaviors contradiction detection, pattern crystallization, and consensus formation arise from the combination of topology, propagation dynamics, and local interaction rules. We present the architecture built on three primitives (knowledge nodes, topology edges, and memory diffs), a multi-layer abstraction hierarchy with promotion via emergent consensus, and theoretical analysis grounding the system in small-world network theory, epidemic broadcast protocols, and swarm intelligence. We provide a reference implementation on PostgreSQL with pgvector and describe a concrete deployment in Swarm-Driven Development, a multi-agent software engineering methodology. HyphaeDB represents, to our knowledge, the first system to combine navigable small world topology with gossip-based knowledge propagation for multi-agent coordination.

Insights

Contribution: The paper introduces HyphaeDB, the first agent-native memory infrastructure that transforms vector database topologies into a communication fabric for multi-agent systems. It enables autonomous knowledge propagation and emergent coordination through a gossip protocol integrated into the memory layer.

Core Idea: Instead of treating memory as passive storage for explicit queries, HyphaeDB treats the HNSW graph as a living topology where agents are nodes that actively propagate and synchronize knowledge.

Technique: The system utilizes a gossip protocol over a Hierarchical Navigable Small World (HNSW) graph, employing energy-based attenuation and local interaction rules to facilitate knowledge diffusion.

Pipeline: Knowledge nodes and memory diffs β†’ Gossip protocol propagation across HNSW topology with energy attenuation β†’ Emergent behaviors (contradiction detection, pattern crystallization, and consensus).

Methodology: The research combines small-world network theory, epidemic broadcast protocols, and swarm intelligence to design a multi-layer abstraction hierarchy for agent memory.

Results: The authors provide a reference implementation on PostgreSQL with pgvector and demonstrate a concrete deployment in a Swarm-Driven Development methodology for software engineering.

Limitations: The paper focuses on theoretical grounding and a reference implementation; scalability of gossip protocols in extremely high-density agent environments remains an area for further exploration.

PDF
5/5 Artificial Intelligence (cs.AI)Computation and Language (cs.CL) 30 Jun 2026
MedEvoEval: Evaluating Continual Evolution of Doctor Agents through Simulated Clinical Episodes

Hui Zhang

Abstract

ArXiv ID: 2606.28900

Authors: Hui Zhang

Abstract:

Doctor agents are moving beyond single-turn answer generation toward evolving clinical decision systems. Within an outpatient episode, they acquire evidence, use examination and consultation resources, and decide when to finalize a diagnosis and management plan. Across episodes, their behavior may change through memory, retrieval, reflection, or other update mechanisms. Current evaluations only partially cover this setting. Fixed-input medical QA benchmarks score final answers from complete inputs, whereas many interactive benchmarks still focus on individual encounters or fixed runs, providing limited support for evaluating how episode-level decisions interact with cross-episode experience. We introduce MedEvoEval, an executable longitudinal evaluation framework based on action-gated simulated outpatient episodes. Each source case is converted into role-specific patient, examination, and manager views; evidence is revealed only through valid actions; and each episode records a structured trace that links observations, actions, final outputs, manager scores, and optional experience write-back. We release a runnable E&D artifact with 700 processed episodes, provenance notes, schemas, an episode runner, scoring scripts, configurations, example logs, analysis code, and trajectory- and step-level derivatives. Experiments show that episode traces expose process costs hidden by final-answer scoring, show how MDT-style consultation reallocates resources, and support longitudinal analyses of memory maturation, held-out transfer, update-stage response, and backward retention. Together, these results show that MedEvoEval provides a concrete basis for evaluating whether doctor agents improve through experience, transfer useful behavior, and retain earlier capabilities over time.

Insights

Contribution: The paper introduces MedEvoEval, an executable longitudinal evaluation framework designed to assess how doctor agents evolve their clinical decision-making across multiple simulated outpatient episodes.

Core Idea: Current benchmarks fail to capture the evolution of agent behavior over time; MedEvoEval addresses this by evaluating episode-level decisions and cross-episode experience updates.

Technique: The framework uses action-gated simulated episodes where evidence is revealed only through valid actions, recording structured traces of observations, actions, and experience write-backs.

Pipeline: Source medical cases β†’ Role-specific views (patient, examination, manager) β†’ Action-gated interaction β†’ Structured episode traces β†’ Longitudinal analysis (memory, transfer, retention).

Methodology: The authors developed a runnable E&D artifact with 700 episodes, implementing a system where agents must actively perform actions to gain information and can update their internal state between episodes.

Results: The framework exposes hidden process costs, demonstrates how MDT-style consultations reallocate resources, and enables analysis of memory maturation, held-out transfer, and backward retention.

Limitations: The study focuses on simulated outpatient episodes and may not fully capture the complexities of real-world, multi-modal clinical environments or high-stakes emergency scenarios.

PDF

Computer Vision

5/5 Artificial Intelligence (cs.AI) 30 Jun 2026
ComMem: Complementary Memory Systems for Test-Time Adaptation of Vision-Language Models

Guanglong Sun, Shuang Cui, Bo Lei, Liyuan Wang, Zihan Zhai, Hongwei Yan, Hang Su, Jun Zhu, Yi Zhong

Abstract

ArXiv ID: 2606.28719

Authors: Guanglong Sun, Shuang Cui, Bo Lei, Liyuan Wang, Zihan Zhai, Hongwei Yan, Hang Su, Jun Zhu, Yi Zhong

Abstract:

Test-time adaptation (TTA) of vision-language models (VLMs) is essential for their robust deployment in dynamic, real-world environments. However, existing TTA methods often adapt locally without accumulating knowledge over time, or operating within a single modality without exploiting VLMs' inherently multi-modal nature. Inspired by the \textbf{Com}plementary \textbf{Mem}ory systems of the biological brain, we propose \textbf{ComMem}, an innovative approach that mimics the distinct but cooperative roles of the hippocampus and neocortex to enable effective TTA for VLMs. ComMem consists of two key components: a fast-adapting detailed memory, akin to the hippocampus, that forms a dynamic visual cache from high-confidence test samples; and a slow-integrating abstract memory, akin to the neocortex, that continually refines global textual prototypes. For each test instance, ComMem jointly optimizes both memory systems to ensure cross-modal consistency. Extensive experiments on 15 benchmark datasets show that ComMem significantly outperforms state-of-the-art methods under both natural distribution shifts and cross-dataset generalization, offering a promising direction for enhancing VLMs' practical adaptability.

Insights

Contribution: The paper introduces ComMem, a novel test-time adaptation (TTA) framework for vision-language models that mimics biological memory systems to accumulate and integrate knowledge over time.

Core Idea: The core idea is to leverage a dual-memory systemβ€”a fast-adapting visual cache and a slow-integrating textual prototype systemβ€”to achieve cross-modal consistency during test-time adaptation.

Technique: The technique involves jointly optimizing a high-confidence visual memory (hippocampus-inspired) and a global textual prototype memory (neocortex-inspired) for each test instance.

Pipeline: Test samples β†’ High-confidence visual caching & Global textual prototype refinement β†’ Joint optimization for cross-modal consistency β†’ Adapted VLM output

Methodology: ComMem utilizes a dynamic visual cache to store detailed information from recent samples and a slow-integrating memory to refine abstract textual prototypes, ensuring the model adapts to distribution shifts without forgetting global knowledge.

Results: ComMem significantly outperforms state-of-the-art methods across 15 benchmark datasets, demonstrating superior performance in both natural distribution shifts and cross-dataset generalization.

Limitations: The paper does not explicitly detail the computational overhead of maintaining dual memories or the specific criteria for determining 'high-confidence' samples in highly noisy environments.

PDF
4/5 Artificial Intelligence (cs.AI) 30 Jun 2026
COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan

Abstract

ArXiv ID: 2606.28696

Authors: Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan

Abstract:

Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation. We present COMPASS, the first unified multimodal framework that grounds composition-intent control in a single system spanning both composition perception and composition-guided generation, with a shared expert token $\tau_c$ as the central intent anchor. On the perception side, COMPASS injects composition expertise into an MoE backbone in a minimally invasive manner and distills the inferred intent into $\tau_c$. On the generation side, COMPASS reuses $\tau_c$ as a global conditioning signal that steers the denoising trajectory, effectively converting passive composition analysis into explicit layout control. To support systematic instruction-following composition learning and evaluation at scale, we construct Comp-11, a large-scale dataset with an 11-class taxonomy and reasoning-augmented annotations. Extensive experiments show that COMPASS substantially improves category-level composition understanding and delivers more composition-consistent, prompt-faithful generation than strong baselines.

Insights

Contribution: The paper introduces COMPASS, the first unified multimodal framework that bridges composition perception and composition-guided generation using a shared expert token. It also provides Comp-11, a large-scale dataset with an 11-class taxonomy and reasoning-augmented annotations for systematic composition learning.

Core Idea: The core idea is to ground high-level visual intent (composition) into a single system by using a shared expert token ($ au_c$) as a central anchor to link perception and generation.

Technique: The framework utilizes a Mixture-of-Experts (MoE) backbone to inject composition expertise and a shared token $ au_c$ to steer the denoising trajectory during generation.

Pipeline: Image/Text Input β†’ MoE-based Composition Perception β†’ Intent Distillation into $ au_c$ β†’ $ au_c$ Guided Denoising β†’ Composition-Consistent Image Generation

Methodology: COMPASS distills inferred composition intent into a global conditioning signal ($ au_c$) that is reused across both the perception and generation modules of a unified multimodal model.

Results: COMPASS substantially improves category-level composition understanding and produces more composition-consistent, prompt-faithful images compared to existing strong baselines.

Limitations: The paper does not explicitly detail the computational overhead of the MoE backbone or the scalability of the 11-class taxonomy to more complex, non-standard artistic styles.

PDF

General

3/5 Artificial Intelligence (cs.AI)q-bio.QMstat.AP 30 Jun 2026
Expert Evaluation of Clinical AI Tools on Real Point-of-Care Clinical Queries

Jean Feng, Vishal Patel, Patrick Heagerty, Yifan Mai, Venkatesh Sivaraman, Patrick Vossler, Jialin Ouyang, Anupam B. Jena

Abstract

ArXiv ID: 2606.28960

Authors: Jean Feng, Vishal Patel, Patrick Heagerty, Yifan Mai, Venkatesh Sivaraman, Patrick Vossler, Jialin Ouyang, Anupam B. Jena

Abstract:

Physicians now pose millions of clinical questions to AI tools each week, yet these tools are evaluated largely on hypothetical or exam-style questions, not those actually asked in practice. We report a blinded evaluation built on 620 Real-world Point-Of-Care Queries (Real-POCQi) submitted to the OpenEvidence (OE) platform by physicians spanning 30 specialties, as well as 187 questions from HealthBench. 149 practicing physicians across 36 states made head-to-head comparisons between answers from three frontier general-purpose models (Claude Opus 4.8, Gemini 3.1 Pro, and GPT-5.5) and a specialized clinical tool (OE), with graders matched to each question's specialty. When comparing answers along five dimensions relevant to clinical decision support -- accuracy, clinical utility, source quality, verifiability, & completeness -- physicians scored the specialized tool highest on all axes; in the primary analysis on Real-POCQi, win differences (margins between win and loss rates) ranged from 25 to 39 percentage points (p<0.001). Results remained consistent in sensitivity analyses stratifying by citation display, answer length, OE-user status, and Real-POCQi versus HealthBench. In parallel, LLM judges were found to systematically differ from expert judges, though both generally agreed on the best model. These findings underscore two conclusions: (i) AI tool evaluations should reflect real-world query distributions and use expert judges that mirror the specialization defining modern medicine and (ii) the consistent advantage of the specialized tool over general-purpose models does not necessarily mean that the latter cannot serve similar purposes, but that targeted engineering and customization can yield meaningful gains in performance for its users. We release Real-POCQi as a public benchmark, as well as the prespecified statistical analysis for reproducing results of this study.

Insights

Contribution: The study introduces a new benchmark of 620 real-world point-of-care clinical queries (Real-POCQi) and provides a blinded expert evaluation comparing specialized clinical AI tools against frontier general-purpose models.

Core Idea: Current AI evaluations rely on hypothetical questions, whereas clinical utility is best measured by evaluating how models handle actual queries posed by physicians in practice.

Technique: A blinded, head-to-head comparison of AI outputs across five clinical dimensions (accuracy, utility, source quality, verifiability, and completeness) using expert physician graders.

Pipeline: Real-world physician queries (Real-POCQi) β†’ Generation of answers by general-purpose models (Claude, Gemini, GPT) and a specialized tool (OpenEvidence) β†’ Blinded expert evaluation by specialty-matched physicians β†’ Statistical analysis of win margins.

Methodology: 149 practicing physicians performed a blinded comparison of three frontier models and one specialized tool across 620 real-world queries and 187 HealthBench questions, with results stratified by various sensitivity factors.

Results: The specialized tool (OpenEvidence) scored highest across all five dimensions, with win differences ranging from 25 to 39 percentage points (p<0.001) over general-purpose models.

Limitations: LLM judges were found to systematically differ from human expert judges, suggesting that automated evaluation may not fully capture the nuances of clinical judgment.

PDF

LLM

4/5 Artificial Intelligence (cs.AI) 30 Jun 2026
IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf

Abstract

ArXiv ID: 2606.28556

Authors: Maria Xenochristou, Ashutosh Joshi, Korosh Vatanparvar, Mohammad Abuzar Hashemi, Prasad Kasu, Deepak Bansal, Anchal Nema, Nivedita Wadhwa, Prashams S Jain, Rebecca Abraham, Will Kimbrough, Dilek Hakkani-Tur, Wilko Schulz-Mahlendorf

Abstract:

Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging. However, existing medical AI benchmarks are fragmented: some support multi-turn dialogues but lack images, while others provide multimodal inputs but focus on single-turn QA tasks. To address this gap, we introduce IMCBench, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles to simulate realistic patient-clinician interactions. Each conversation is evaluated across three clinical dimensions: safety, accuracy, and appropriate use of uncertainty in diagnosis. We benchmark eight multimodal frontier models across four model families (Claude, GPT, Nova, and Llama), scoring each on a 1-5 scale using LLM-as-Jury scoring calibrated against expert clinician annotations. Our results show that Claude Opus 4.6 achieves the highest overall score (3.61), followed by Claude Sonnet 4.6 (3.30) and GPT-5.2 (3.29), though no model dominates all dimensions and safety degrades for both malignant and rare conditions ($\Delta$ = -0.27 each). Ablation studies further reveal that both visual input and EHR context contribute to safe guidance (safety drops of 0.18 and 0.23 on average when each is removed), with stronger models leveraging visual features more effectively. Together, these findings demonstrate that accurate clinical description does not guarantee safe patient guidance, motivating the need for multi-dimensional evaluation frameworks in medical AI.

Insights

Contribution: The paper introduces IMCBench, a novel benchmark designed to evaluate multimodal LLMs on image-grounded, multi-turn medical conversations. It addresses the gap in existing benchmarks by combining real clinical images with synthetic patient profiles to simulate realistic clinician-patient interactions.

Core Idea: Accurate clinical description does not automatically translate to safe patient guidance, necessitating a multi-dimensional evaluation framework that considers safety, accuracy, and uncertainty.

Technique: The authors utilize a multi-turn dialogue framework paired with real clinical images and synthetic EHR profiles, evaluated using an LLM-as-Jury scoring system calibrated against expert clinician annotations.

Pipeline: Real clinical images + Synthetic patient profiles β†’ Multi-turn medical conversation simulation β†’ Multi-dimensional evaluation (Safety, Accuracy, Uncertainty) β†’ LLM-as-Jury scoring calibrated by clinicians.

Methodology: The study benchmarks eight frontier multimodal models across four families, using ablation studies to measure the specific contributions of visual inputs and EHR context to model safety.

Results: Claude Opus 4.6 achieved the highest overall score (3.61), followed by Claude Sonnet 4.6 (3.30) and GPT-5.2 (3.29); safety scores significantly degraded for malignant and rare conditions (Ξ” = -0.27).

Limitations: No single model dominated all clinical dimensions, and the study highlights that even frontier models struggle with safety in complex clinical scenarios.

PDF
4/5 Artificial Intelligence (cs.AI) 30 Jun 2026
Search for Truth from Reasoning: A Dynamic Representation Editing Framework for Steering LLM Trajectories

Tianlong Wang, Yuhang Wang, Weibin Liao, Xin Gao, Xinyu Ma, Yang Lin, Yasha Wang, Liantao Ma

Abstract

ArXiv ID: 2606.28589

Authors: Tianlong Wang, Yuhang Wang, Weibin Liao, Xin Gao, Xinyu Ma, Yang Lin, Yasha Wang, Liantao Ma

Abstract:

Current approaches to enhance Large Language Model (LLM) reasoning, such as Chain-of-Thought and "Wait" prompts, primarily encourage models to think more, yet often fail to guide them toward Truth. While Representation Editing (RepE) offers a intrinsic control, its application to dynamic reasoning trajectories remains underexplored. In this work, we bridge this gap by investigating the geometry of truth within unfolding reasoning chains. We uncover three critical insights: (1) Truth is encoded at the sentence level and is entangled with latent reasoning patterns; (2) Effective intervention follows an Uncertainty Principle and a Decay Effect, requiring localization to early, high-entropy forks; (3) Naive steering vectors suffer from noise, risking collateral damage to correct trajectories. Based on these findings, we propose DynaSteer, a dynamic RepE framework. DynaSteer employs pattern clustering to disentangle reasoning manifolds and utilizes Fisher-LDA to project purified truth. By dynamically monitoring lookahead entropy, it selectively steers and rolls back trajectories only when necessary. Comprehensive experimental results on several MATH benchmark verify the effectiveness of DynaSteer, and experiments on out-of-domain coding tasks further confirm its generalization ability. Our code is publicly available at https://github.com/tianlwang/DynaSteer.

Insights

Contribution: The paper introduces DynaSteer, a dynamic Representation Editing (RepE) framework that steers LLM reasoning trajectories toward truth by identifying and intervening at critical high-entropy forks.

Core Idea: Instead of just prompting models to think more, the authors propose geometrically intervening in the latent space of unfolding reasoning chains to correct errors early before they propagate.

Technique: The framework utilizes pattern clustering to disentangle reasoning manifolds and Fisher-LDA to project purified truth vectors while monitoring lookahead entropy for selective steering.

Pipeline: Reasoning Trajectory β†’ Pattern Clustering & Manifold Disentanglement β†’ Lookahead Entropy Monitoring β†’ Selective Steering/Rollback β†’ Corrected Reasoning Path

Methodology: The authors analyze the geometry of truth in reasoning chains to identify an 'Uncertainty Principle' and 'Decay Effect,' then develop a dynamic intervention mechanism that applies purified steering vectors only when necessary.

Results: DynaSteer demonstrates superior performance on MATH benchmarks and shows strong generalization capabilities on out-of-domain coding tasks compared to baseline steering methods.

Limitations: The paper notes that naive steering vectors can cause collateral damage to correct trajectories, and the effectiveness of intervention is highly dependent on early localization.

4/5 Artificial Intelligence (cs.AI)Machine Learning (cs.LG) 30 Jun 2026
Self-Supervised Theorem Discovery in a Formal Axiomatic System

Kazuki Ota, Takayuki Osa, Tatsuya Harada

Abstract

ArXiv ID: 2606.28747

Authors: Kazuki Ota, Takayuki Osa, Tatsuya Harada

Abstract:

Recent artificial intelligence (AI) systems have shown remarkable progress in mathematical reasoning. Many existing approaches, including large language models (LLMs), draw on human prior knowledge in the form of mathematical text, code, or theorem libraries. Although these approaches are highly effective in practice, it remains an open question whether an agent can autonomously discover useful theorems without such human priors. We study this question in a formal axiomatic system by developing an agent that starts from axioms and inference rules alone and gradually grows a library of useful theorems. Concretely, we propose a self-supervised theorem-discovery algorithm that alternates between proof search and useful-theorem extraction, building a theorem library whose entries are reused as lemmas for subsequent proof search. Experiments show that the agent discovers tens of thousands of theorems and finds proofs for human-written benchmark problems, suggesting that its discoveries include theorems meaningful from a human mathematical perspective. Furthermore, the discovered theorems improve LLM proof performance when provided as prompt lemmas, indicating that they can serve as external knowledge for LLM reasoning. Our results provide evidence that useful theorems can emerge from proof search without relying on human-provided theorem libraries. More broadly, they suggest a path toward self-evolving AI systems for mathematics whose discoveries remain formally verifiable.

Insights

Contribution: The paper demonstrates that an AI agent can autonomously discover useful mathematical theorems from scratch using only axioms and inference rules, without relying on human-provided prior knowledge. It also shows that these self-discovered theorems can effectively enhance the reasoning capabilities of Large Language Models (LLMs).

Core Idea: The research explores whether a self-evolving system can build its own library of lemmas through iterative proof search to solve complex problems in a formal axiomatic system.

Technique: The authors propose a self-supervised theorem-discovery algorithm that alternates between searching for proofs and extracting 'useful' theorems to be reused as lemmas.

Pipeline: Axioms and inference rules β†’ Iterative proof search and useful-theorem extraction β†’ Growing theorem library β†’ Solving benchmark problems and augmenting LLM prompts

Methodology: The researchers developed an agent that builds a library of theorems by identifying successful proof paths and storing them as lemmas for future use. They evaluated the system by testing its ability to solve human-written benchmarks and its impact on LLM proof performance.

Results: The agent discovered tens of thousands of theorems, successfully solved human-written benchmark problems, and improved LLM proof performance when the discovered theorems were provided as prompt lemmas.

Limitations: The study focuses on formal axiomatic systems; it remains to be seen how well this self-discovery scales to more complex, non-formalized mathematical domains or extremely high-level abstract reasoning.

PDF
4/5 Artificial Intelligence (cs.AI) 30 Jun 2026
Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions

David Courtis, Ting Hu

Abstract

ArXiv ID: 2606.28770

Authors: David Courtis, Ting Hu

Abstract:

Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.

Insights

Contribution: The paper introduces a mechanistic interpretability approach to steer LLM personality by identifying and intervening on latent features in the residual stream rather than relying on prompt engineering or fine-tuning.

Core Idea: Personality traits can be controlled by applying additive shifts to specific latent directions in the model's hidden states that correspond to OCEAN personality dimensions.

Technique: The authors use Sparse Autoencoders (SAEs) and contrastive activation analysis to isolate personality-related features and develop a linear weighting heuristic for steering.

Pipeline: Input text β†’ Latent feature identification via SAEs β†’ Contrastive activation analysis β†’ Additive steering vector application β†’ Personality-steered output

Methodology: The researchers identify latent directions for OCEAN traits, formalize an additive steering vector, and use grid search optimization to balance personality expression with language modeling performance.

Results: The method successfully enhances target personality traits while preserving high performance on standard benchmarks through optimized linear weighting of feature shifts.

Limitations: The study focuses on additive shifts and linear weighting heuristics, leaving open questions regarding the complexity of non-linear feature interactions and the scalability of SAE-based steering across all personality dimensions.

PDF

MLOps

4/5 Artificial Intelligence (cs.AI) 30 Jun 2026
Data and Evaluation Closed-Loop for Model Capability Enhancement

Zhixuan Li, Jiangan Yuan, Han Xu

Abstract

ArXiv ID: 2606.28471

Authors: Zhixuan Li, Jiangan Yuan, Han Xu

Abstract:

Model capability is the central variable in LLM pre-training, yet is never observed directly: data shapes it prospectively, while evaluation reveals it only retrospectively, compressing samples, prompts, decoding, and scoring rules into one noisy score. Practical optimization runs this backward: a failure is observed first, and the engineer must infer the corpus fix. The two sides speak incompatible vocabularies -- benchmark names and per-sample correctness versus data sources, domains, and quality labels -- so this inference is usually intuition, not method. We close this gap with the \emph{capability slice}: a group of evaluation samples sharing background condition, task type, solving operation, and output constraint -- precise enough to localize a single weakness yet stable enough to survive aggregation, unlike a benchmark name, too coarse, or a single sample, too noisy. Built around this unit, an evaluation taxonomy, a non-instruction data taxonomy, and mapping rules form a closed loop turning a benchmark-level failure into a targeted, testable data intervention. We test this loop on two case studies pulling in opposite directions. First, the loop rules the data out: continued pre-training drives BBH down by $-46.82\%$, but diagnosis traces this to a single masked \texttt{\textless EOS\textgreater} loss rather than weakened reasoning; restoring it recovers BBH to $66.44$, above the original checkpoint, without changing the data. Second, the loop rules the data in: a persistent math-reasoning weakness is decomposed by solving operation into specific failing combinations, and a weakness-targeted sampling procedure built from it lifts AIME2025/AIME2026 Pass@128 from $6.67$/$0.00$ to $26.67$ each. The same unmodified loop reaches opposite, correct verdicts in both cases, showing the evaluation-to-data inference can be routine, auditable, and experimentally validated rather than intuitive.

Insights

Contribution: The paper introduces a closed-loop framework that bridges the gap between retrospective evaluation and prospective data selection by defining a 'capability slice' as a common unit of analysis.

Core Idea: By mapping benchmark failures to specific data interventions through a structured taxonomy, the authors transform intuitive model debugging into a routine, auditable, and experimentally validated process.

Technique: The authors develop a capability sliceβ€”a group of evaluation samples sharing background conditions, task types, solving operations, and output constraintsβ€”to localize model weaknesses.

Pipeline: Benchmark failure β†’ Capability slice diagnosis β†’ Data taxonomy mapping β†’ Targeted data intervention β†’ Model re-training/fine-tuning β†’ Evaluation

Methodology: The authors established an evaluation taxonomy and a non-instruction data taxonomy, creating mapping rules to translate specific performance drops into actionable data corpus adjustments.

Results: The loop correctly identified a masked loss (recovering BBH scores) and successfully lifted AIME2025/AIME2026 Pass@128 from 6.67/0.00 to 26.67 each through targeted sampling.</p>

Limitations: The paper focuses on pre-training and specific reasoning benchmarks; the scalability of the manual taxonomy mapping to all possible LLM capabilities remains an open question.

</details> </div>
PDF
</div> #### NLP
3/5 Artificial Intelligence (cs.AI)stat.AP 30 Jun 2026
Primary ICD Category Prediction using LLM-based Probing

Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen

Abstract

ArXiv ID: 2606.28798

Authors: Chengyuan Liu, Xinyue Zhang, Yao Li, Guanting Chen

Abstract:

Objective: ICD codes are central to reimbursement, research, and population health surveillance, yet automated coding systems often struggle to integrate diagnostic signals from both clinical narratives and structured electronic health record (EHR) variables. We evaluated whether frozen medical large language model (LLM) representations can serve as a shared embedding space for multimodal primary diagnosis category prediction. Materials and Methods: We constructed a MIMIC-IV cohort of 13,645 admissions from the 10 most frequent primary ICD-10 codes, consolidated into seven categories. Structured variables were serialized into clinical narratives and combined with leakage-pruned discharge notes. Using a frozen MedFound-Llama3-8B-finetuned backbone, we extracted hidden states from five transformer layers and trained linear probes for structured-only, unstructured-only, and combined inputs, comparing against XGBoost and information-matched PLM-ICD baselines and evaluating MIMIC-III adaptation with a compact bottleneck adapter. Results: The combined probe performed best on MIMIC-IV (87.69% strict; 91.45% medical accuracy), exceeding both single-modality probes and baselines. The structured-only probe outperformed its standard baseline by 6.19 points in medical accuracy. Diagnostic information became increasingly linearly separable in deeper layers, and a 2M-parameter adapter restored cross-dataset transfer to MIMIC-III using only 5% of target labels. Discussion: LLM embeddings can unify structured and narrative EHR information for multimodal diagnosis prediction, supporting efficient reuse of clinical representations across modalities and datasets through a small representation-level module. Conclusion: Multimodal probing of frozen medical LLM representations provides a practical approach for studying EHR modalities and adapting clinical representations across datasets.

Insights

Contribution: The study demonstrates that frozen medical LLM representations can serve as a unified embedding space for multimodal primary ICD category prediction, effectively integrating structured and unstructured EHR data.

Core Idea: By using a frozen LLM backbone, the researchers show that diagnostic signals from both clinical narratives and structured variables are linearly separable in the model's hidden states.

Technique: The authors employ linear probing on hidden states from multiple transformer layers of a frozen MedFound-Llama3-8B model, combined with a compact bottleneck adapter for cross-dataset adaptation.

Pipeline: Structured EHR variables and clinical narratives β†’ Serialization and leakage-pruning β†’ Frozen MedFound-Llama3-8B embedding extraction β†’ Linear probing of hidden states β†’ Primary ICD category prediction

Methodology: The researchers evaluated three input types (structured-only, unstructured-only, and combined) using a MIMIC-IV cohort of 13,645 admissions, comparing performance against XGBoost and PLM-ICD baselines.

Results: The combined probe achieved 87.69% strict and 91.45% medical accuracy on MIMIC-IV, outperforming single-modality probes and baselines; a 2M-parameter adapter successfully restored transferability to MIMIC-III using only 5% of target labels.

Limitations: The study focuses on the top 10 most frequent primary ICD codes and evaluates the linear separability of representations, which may not capture the complexity of rare diseases or non-linear diagnostic relationships.

PDF
#### RL
5/5 Artificial Intelligence (cs.AI) 30 Jun 2026
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards

Yupeng Chang, Yuan Wu, Yi Chang

Abstract

ArXiv ID: 2606.28707

Authors: Yupeng Chang, Yuan Wu, Yi Chang

Abstract:

Critic-free reinforcement learning with verifiable rewards (RLVR), exemplified by Group Relative Policy Optimization (GRPO), avoids training a value function (critic) and reduces memory and compute overhead relative to critic-based PPO pipelines for aligning large language models. However, GRPO-style advantage estimation depends on prompt-local (within-prompt-group) reward statistics and can be unstable. In particular, when all rollouts in a prompt group receive identical rewards, the within-group reward variance becomes zero, and group normalization yields zero advantages for that group, impeding learning in cold-start regimes with binary verifiers. We introduce BV-Blend, a critic-free framework that stabilizes advantage estimation by combining prompt-local on-policy statistics with semantic-cluster-conditioned historical moments. BV-Blend maintains EMA-tracked reward moments for each cluster, derives a confidence weight from a standard error of the mean (SEM) proxy, and uses this weight to blend historical and prompt-local baseline and variance statistics into a standardized advantage for PPO-style clipped updates. Experiments on verifiable reasoning benchmarks show that BV-Blend improves training stability and performance, and remains robust in regimes where group-normalized methods may stall.

Insights

Contribution: The paper introduces BV-Blend, a critic-free reinforcement learning framework that stabilizes advantage estimation in RLVR by blending prompt-local statistics with historical moments.

Core Idea: To overcome the zero-variance issue in group-normalized methods (like GRPO) during cold-start regimes, the authors use semantic-cluster-conditioned historical data to provide a stable baseline.

Technique: The method employs EMA-tracked reward moments for semantic clusters and a confidence weight derived from a standard error of the mean (SEM) proxy to blend historical and on-policy statistics.

Pipeline: Prompt group rollouts β†’ Reward collection β†’ Semantic clustering β†’ EMA moment tracking β†’ Confidence weight calculation β†’ Weighted blending of historical and local statistics β†’ Standardized advantage calculation β†’ PPO-style clipped updates

Methodology: The framework maintains historical reward moments for different semantic clusters and uses a confidence-weighted blending mechanism to calculate the baseline and variance for advantage estimation.

Results: BV-Blend improves training stability and performance on verifiable reasoning benchmarks, specifically remaining robust in scenarios where group-normalized methods stall due to zero variance.

Limitations: The effectiveness depends on the quality of semantic clustering and the accuracy of the SEM proxy for confidence weighting.

PDF
#### Robotics
4/5 Artificial Intelligence (cs.AI) 30 Jun 2026
TrajRS: Towards Certified Robustness in Pedestrian Trajectory Prediction

Liang Zhang, Gaojie Jin, Yao Shi, Quanzhi Li, Cheng-Chao Huang, David N. Jansen, Lijun Zhang

Abstract

ArXiv ID: 2606.28716

Authors: Liang Zhang, Gaojie Jin, Yao Shi, Quanzhi Li, Cheng-Chao Huang, David N. Jansen, Lijun Zhang

Abstract:

The robustness of trajectory prediction models is crucial for developing safe autonomous driving systems. Adversarial attacks on trajectory prediction can significantly impair the accuracy of predicted trajectories, leading to hazardous driving behaviors. While heuristic defense strategies have been implemented to enhance the robustness of trajectory prediction models, these measures often fail against more sophisticated, targeted adversarial attacks. Hence, there is a pressing need to establish verifiable safety assurances for trajectory prediction models. In this paper, we extend the traditional Randomized Smoothing framework to "TrajRS", which provides a certified robust radius for smoothed trajectory predictors. We clarify and expand the formal definitions of robustness in trajectory prediction and tailor the practical TrajRS scheme specifically to "robustness for the optimal prediction" and "robustness for all possible predictions". An extensive set of experiments demonstrates that TrajRS effectively achieves robustness certification for all smoothed pedestrian trajectory predictors in this work.

Insights

Contribution: The paper introduces TrajRS, a framework that extends Randomized Smoothing to provide certified robustness for pedestrian trajectory prediction models. It establishes formal definitions for robustness in this context and provides verifiable safety assurances against adversarial attacks.

Core Idea: To move beyond heuristic defenses, the authors propose a provable safety guarantee by certifying a robust radius within which a trajectory prediction remains stable.

Technique: The authors adapt the Randomized Smoothing framework to trajectory data, tailoring the scheme to address both 'robustness for the optimal prediction' and 'robustness for all possible predictions'.

Pipeline: Pedestrian input data β†’ Randomized Smoothing (adding noise) β†’ Smoothed Trajectory Predictor β†’ Certified Robust Radius

Methodology: The researchers formalize robustness metrics for trajectory prediction and develop a mathematical framework to calculate certified radii for various smoothed predictors.

Results: Extensive experiments demonstrate that TrajRS successfully achieves robustness certification for all smoothed pedestrian trajectory predictors tested in the study.

Limitations: The paper focuses on certified robustness via smoothing, which may involve a trade-off between the certified radius size and the accuracy of the underlying prediction model.

PDF

Tech News

### AI Safety
Reddit r/ArtificialIntelligence 2026-06-30
I built a benchmark for multi-turn prompt injection attacks. Most defenses never see them coming.

A researcher has released an open-source benchmark designed to test LLMs against multi-turn prompt injection attacks, which are more sophisticated than standard one-shot attacks. The study revealed that existing defenses like LLM Guard and Arc Gate struggle to detect gradual semantic manipulation over multiple interactions. The project includes a benchmark, proxy, and live red team environment to encourage community-driven security improvements.

### Computing Systems
Hacker News Tue, 30 Ju
Memory Safe Context Switching (longjmp, setjmp) in Fil-C

The Fil-C project introduces a memory-safe approach to context switching using setjmp and longjmp. It aims to provide the low-level control of C with the safety guarantees required for modern systems programming. This is significant for developing robust, secure foundations for high-performance computing.

Reddit r/ArtificialIntelligence 2026-06-30
A native Rust cognitive engine that routes language through a biologically faithful neural substrate

GoldWorm is a novel cognitive engine built in Rust that routes language through a biologically faithful model of the C. elegans 302-neuron connectome. Unlike traditional black-box LLMs, it utilizes a transparent, zero-trust architecture with physically separated action and learning streams to prevent catastrophic forgetting. The system emphasizes structural immutability and inspectable synapses over massive parameter scaling.

### General
Reddit r/MachineLearning 2026-06-29
Loss functions in Instance Representation Learning [R]

A user on r/MachineLearning is seeking clarification on the mathematical justification for using Noise-Contrastive Estimation (NCE) over standard Softmax Negative Log-Likelihood in instance representation learning. The discussion focuses on why NCE is preferred despite still requiring denominator estimation, specifically regarding bias and gradient convergence as noise samples increase.

Reddit r/ArtificialIntelligence 2026-06-30
Government throttling & gatekeeping of new frontier LLMs is just a temporary glitch in the Matrix, not "the new norm going forward", here's why.

The post argues that government attempts to throttle or gatekeep frontier LLMs are temporary measures that will be overridden by the forces of global capitalism. It suggests that because AI is essentially 'powerful software' and not a physical weapon like nuclear material, the demand for profit and international competition (specifically mentioning China) will ensure widespread accessibility.

### LLM
Hacker News Tue, 30 Ju
LongCat-2.0, a large-scale MoE model with 1.6T total and 48B Active

LongCat-2.0 is a large-scale Mixture-of-Experts (MoE) model featuring 1.6 trillion total parameters with 48 billion active parameters. The model aims to balance high-capacity reasoning with efficient inference performance. It represents a significant step in scaling MoE architectures for complex tasks.

### Robotics
Hacker News Mon, 29 Ju
South Korea to spend $1T on more memory chip production and humanoid robots

South Korea has announced a massive $1 trillion investment plan aimed at expanding memory chip production and accelerating the development of humanoid robots. This initiative seeks to solidify the nation's position in the global hardware supply chain and the burgeoning robotics market. The move highlights the critical intersection between high-performance computing infrastructure and physical AI applications.

Trending repositories on GitHub filtered and scored for relevance to your interests.

### Agentic AI ### Computer Vision ### General ### Robotics