Daily Digest 2026-07-29
Global Trends
Personal Interests
Papers discovered through your interest topics.
Embodied AI
Abstract
ArXiv ID: 2607.24148
Authors: Zhuoran Song, Haozhe Jiang, Chunyu Qi, Minnan Pei, Gang Li, Xiaoyao Liang, Haibing Guan
Abstract:
Vision-Language-Action (VLA) models have demonstrated strong potential for embodied AI, yet their high inference latency on GPUs limits real-time deployment. Existing accelerators, such as Dadu-Corki, improve efficiency but treat VLA models as full-precision workloads, leaving substantial redundancy in both memory and computation underexploited. In this paper, we propose VQVLA, an algorithm-hardware co-design framework that accelerates VLA inference by exploiting weight similarity and execution dynamics. We first introduce MotionVQ, a motion-aware vector quantization scheme that dynamically adjusts quantization precision based on the robot's execution state, reducing memory access while preserving task success rate. We then propose a merged-centroid vectorized GEMM paradigm that operates on the codebook-index representation, eliminating redundant multiplications through spatial aggregation and temporal reuse of centroids. To realize these optimizations, we design an accelerator that efficiently supports dynamic precision selection and centroid-reuse computation. Experimental results show that VQVLA achieves 6.5x, 2.8x, 1.9x, 3.3x, and 4.3x speedup over the A100 GPU, Dadu-Corki, LUT-DLA, CodeGEMM, and ShiftAddLLM, respectively, with negligible accuracy degradation.
Human-Computer Interaction
Abstract
ArXiv ID: 2607.24430
Authors: Yifan Hu, Shuwei He, Rui Liu, Haizhou Li
Abstract:
Conversational Speech Synthesis is a fundamental component of human-computer interaction, aiming to generate contextually appropriate, expressive, and empathetic speech. However, facial expressions encode subtle and rich affective cues that are crucial for empathetic speech interaction, whereas existing approaches often overlook this important modality. In addition, the lack of large-scale natural conversational datasets with both speech and visual modalities also limits the development of visual affect understanding in conversational settings.To address these limitations, we propose FacialTalker, a facial-expression-aware CSS framework built upon a large language model backbone. To efficiently encode facial expressions, we propose AUTokenizer, a single-codebook visual tokenizer that discretizes each frame-level facial expression into a compact token, trained with supervision from combinations of facial Action Units. We further introduce a dual direct preference optimization (DualDPO) strategy, which extends the DPO by jointly imposing preference constraints on both visual and speech token sequences, to enhance the model's understanding of facial expressions and speech semantics in multimodal conversational contexts. Moreover, we construct VSDD-1K, a large-scale multimodal dialogue dataset collected through a fully automated pipeline from real-world Internet conversations, comprising over 1,033 hours of synchronized speaker videos and speech, with more than 85\% of frames containing valid faces. Extensive objective and subjective experiments demonstrate that FacialTalker consistently outperforms strong baselines in facial-expression perception and speech synthesis quality, generating speech that is more natural, expressive, and better aligned with the conversational context. The results also validate the effectiveness of our training strategy and dataset construction pipeline.
Multi-Agent Systems
Abstract
ArXiv ID: 2607.25564
Authors: Yuwen Ma, Sarah Spurgeon, Tao Li, Boli Chen
Abstract:
This paper proposes a differentially private distributed cooperative control scheme for multi-agent systems (MAS). Unlike conventional approaches that actively inject artificial noise for privacy protection, this work investigates whether inherent communication noise can itself serve as a natural privacy mechanism. A physically motivated communication-noise model is developed for mobile MAS by incorporating transmitter perturbation, receiver noise, path-loss attenuation, and log-normal shadowing. The resulting effective noise variance depends on inter-agent state differences, thereby capturing the distance-dependent signal perturbation arising in practice. Based on this model, a distributed finite-horizon Linear Quadratic Regulator (LQR) mechanism is designed to achieve formation tracking while protecting agents' private control preferences. Rather than protecting the full local cost function, the proposed privacy formulation focuses on the ratio of the LQR weighting matrices, which captures the trade-off between tracking accuracy and control effort when the quadratic cost structure is publicly known. A set-theoretic sensitivity analysis shows that this weighting-ratio adjacency formulation yields less conservative privacy bounds than gradient-based protection under the considered addition/removal adjacency relation. Theoretical analysis demonstrates that, under suitable design conditions, the proposed mechanism provides bounded cumulative (Ξ΅,Ξ΄)-differential privacy guarantees for the weighting ratios over an infinite horizon without artificial noise injection. Meanwhile, the cooperative tracking error is shown to converge almost surely and in mean square to a finite random limit, with its expectation remaining bounded. Numerical examples validate the theoretical results and illustrate the resulting privacy-performance trade-off.
Abstract
ArXiv ID: 2607.25316
Authors: Nagarani Brammanayagam, Devaprakash Muniraj
Abstract:
Collective intent prediction in multi-agent systems focuses on predicting the shared objectives and future behaviours of groups of interacting agents. The problem is particularly challenging because collective intent emerges from complex interactions, evolving cooperation structures, and long-term behavioural dependencies among heterogeneous agents operating in dynamic and partially observable environments. Furthermore, functional roles adopted by agents are often latent, may change over time, and are rarely available as explicit annotations, making the learning of coordinated group behaviours significantly more difficult. To address these challenges, this paper proposes a Meta-Role Temporal Graph Network (MR-TGN) framework for collective intent prediction in multi-agent systems. The proposed framework models agents as dynamically evolving graph entities and employs temporal memory mechanisms to encode historical interactions and coordination behaviours. To capture higher-level behavioural knowledge, MR-TGN introduces a memory-enhanced meta-role learning mechanism that derives latent role representations from agent-centric behavioral representations without requiring explicit role labels. An evaluation methodology for early collective intent prediction is proposed to assess prediction accuracy and timeliness, enabling realistic evaluation of the model's ability to anticipate collective objectives during the early stages of mission execution. Experimental results on representative multi-agent scenarios demonstrate that the proposed framework consistently outperforms competitive baselines and achieves effective early prediction of collective intents in dynamic and adversarial environments.
Abstract
ArXiv ID: 2607.25255
Authors: Haowen Dai, Zonghao Ying, Wenfeng Li, Xiangfan Wu, Yisong Xiao, Tianyuan Zhang, Jiaye Lin, Lei Wei, Guangyuan Dong, Xitong Ling, Xixun Lin, Quanchen Zou, Xiangzheng Zhang
Abstract:
Multi-agent systems improve capability through task decomposition and role specialization, but these same mechanisms introduce an important safety blind spot: a harmful objective can be fragmented into locally plausible subtasks, allowing malicious intent to evade detection by any single agent. This is a growing social-impact challenge: systems handling sensitive information or consequential tools can turn routine delegation into unauthorized disclosure or unsafe action. We argue that this failure mode is better understood as a semantic information-flow problem than as a single-turn prompt classification task. To address this, we propose SafeFlow, a defense framework for multi-agent systems that formalizes malicious cross-agent propagation as a semantic information-flow problem. SafeFlow attaches structured semantic taints to root requests, propagates them through a dynamic collaboration graph, and performs workflow-level validation to reconstruct the global risk context before irreversible actions are committed. Evaluated on four benchmarks spanning prompt injection, jailbreak-based unsafe tool use, risky code execution, and harmful web-agent behavior, SafeFlow reduces attack success rates compared to undefended baselines and external defenses while retaining high benign task completion and a high paired safe--harm success rate. Our findings show that multi-agent systems still lack mechanisms for preserving risk semantics across delegation boundaries. This gap can turn routine delegation into privacy harms or unsafe actions that affect people and organizations. SafeFlow keeps this risk visible throughout the workflow, before it results in harm.
Abstract
ArXiv ID: 2607.25098
Authors: Gilberto Gil F. G. Passos, Eber Assis Schmitz, Sildenir Alves Ribeiro
Abstract:
This article presents the development and validation of an artificial market for Brazilian Real Estate Investment Trusts (REITs), known as Fundos de Investimento Imobiliario (FIIs), using agent-based modeling methodology. The central contribution of this work is the integration, within a single multi-agent system, of the FII value chain, from the generation of real estate revenues subject to vacancy and operational costs, through dividend distribution, to the trading of shares by heterogeneous investors mediated by a double auction mechanism with an order book. The model incorporates endogenous macroeconomic variables, such as the Selic, the Brazilian benchmark interest rate, and inflation, and represents agent heterogeneity through a behavioral decomposition into fundamentalist, speculator, and noise trader components, modulated by individual financial literacy levels. The model was calibrated using the Method of Simulated Moments applied to the historical series of the IFIX index, the Brazilian REIT market index, between 2021 and 2025. The validation results, obtained using two distinct methods, demonstrate that the model reproduces the main stylized facts observed in the real market: (i) the coverage rate of calibrated moments exceeds 75 percent; (ii) 96 percent of simulated trajectories are structurally indistinguishable from real IFIX periods according to the nearest-neighbor criterion; and (iii) stylized facts such as the power law of autocorrelations of absolute returns and aggregational Gaussianity emerge spontaneously, without being incorporated into the calibration objective function. The results of the validation process indicate that the artificial market captures structural dynamics of the FII market, opening perspectives for its use as a computational laboratory for the analysis of regulatory policies and pricing mechanisms.
Vision-Language Models
Abstract
ArXiv ID: 2607.25479
Authors: Maria Rosaria Briglia, Igor Maljkovic, Antonio Emanuele CinΓ , Luca Oneto, Iacopo Masi, Fabio Roli
Abstract:
Vision--Language Models (VLMs) are increasingly deployed through a model supply chain in which pretrained checkpoints, architecture definitions, text encoders, and exported computation graphs are distributed by third parties and reused across downstream services. This reuse model creates a security-critical trust boundary: VLM deployments inherit not only learned parameters but also executable behavior encoded in shared model artifacts. In this paper, we show that a malicious provider can exploit this trust boundary by embedding architectural backdoors into VLM supply chains through representation steering. Our attack introduces dormant steering logic into the model architecture through a trigger-gated additive modification of an intermediate representation, without poisoning training data, controlling downstream fine-tuning, or modifying prompts at deployment time. When the trigger is absent, the modification reduces to zero and the model follows its normal computation, preserving clean utility. When the trigger is present, a steering direction shifts the internal representation toward an attacker-defined objective. We evaluate the attack across multiple VLM families and downstream tasks, including visual question answering, text-to-image generation, retrieval, and semantic response biasing. The results show that the proposed architectural steering backdoor compromises integrity, safety enforcement, and ranking fairness while preserving normal behavior on clean inputs. We further show that shared VLM artifacts can carry dormant steering logic against downstream services, and we propose an auditing defense that inspects the executable logic distributed with model artifacts rather than only their learned weights.