学术发表 | 智慧治理学院数智技术领域高水平论文成果动态
中国人民大学智慧治理学院自成立以来
始终围绕智慧治理交叉学科建设
系统推进师资队伍建设、人才培养
与有组织科研
锚定数智技术、智能经济、智慧治理等
核心前沿领域
开展师生共研共创
近两年来在国内外权威期刊和学术会议上
发表了一系列高水平论文
着力构建智慧治理自主知识体系
为推动学术交流和知识传播
特刊发系列成果动态
第01期 数智技术领域成果介绍数智技术以数据、算法与智能模型为基础,推动信息感知、知识发现、智能分析与决策支持不断发展。学院师生围绕数智技术发展的关键问题,从数据资源、智能模型与计算方法等多个层面开展研究,探索复杂信息的有效获取、知识的深度利用以及智能分析能力的持续提升,推动数智技术在不同场景中的应用与发展。
在技术方法研究的基础上,进一步关注数智技术在复杂环境中的适应性与可靠性,研究数据差异、任务变化和应用情境对智能模型运行效果的影响,探索面向真实场景的模型优化、能力增强与可信应用方法。随着数智技术不断融入社会治理、公共服务和产业发展,相关研究也进一步关注数据使用、算法决策与社会影响之间的关系,探索数智技术应用中的公平、可信与治理机制,为数智技术的创新发展与规范应用提供理论基础和方法支撑。
1 人工智能与大模型
论文标题:
Expert Heads: Robust Evidence Identification for Large Language Models
作者:
Qi Wu, Jianfeng Qu, Ximing Li, Zhixu Li
出版物名称:International Conference on Learning Representations
出版时间:2026-04-23
摘要:
Large language models (LLMs) exhibit strong abilities in multi-document reasoning, yet their evidence identification is highly sensitive to input order. We trace this limitation to attention mechanisms, where many heads overemphasize sequence boundaries and neglect central content. We systematically analyze attention distributions under document permutations and discover a small subset of heads that consistently prioritize task-relevant documents regardless of position. We formalize these as Expert Heads, identified via activation frequency and stability across permutations. Experiments on LLaMA, Mistral, and Qwen reveal architecture-specific patterns: mid-layer heads in LLaMA and Mistral dominate semantic integration, while deeper-layer heads in Qwen specialize in evidence selection. Moreover, Expert Heads exhibit concentrated focus during understanding and more distributed engagement during generation. Their activation strongly correlates with answer correctness, providing diagnostic signals for hallucination detection. Leveraging Expert Heads for document voting significantly improves retrieval and ranking on HotpotQA, 2WikiMultiHopQA, and MuSiQue, outperforming dense retrievers and LLM-based ranking with minimal overhead. Ablations confirm that even a small subset achieves robust gains. Our findings establish Expert Heads as a stable and interpretable mechanism for evidence integration, offering new directions for context pruning, hallucination mitigation, and head-guided training of LLMs.
论文标题:
Reading Attribution from Attention: Evidence Heads as Latent Attribution Mechanisms in LLMs
作者:
Qi Wu, Jianfeng Qu, Peng-Fei Zhang, Yanzhe Ji, Siyu Li, Wei Chen, Zhixu Li, Jiajie Xu
出版物名称:The Fortieth Annual Conference on Neural Information Processing Systems
出版时间:2026-12-6(会议录用)
摘要:
Large language models (LLMs) show strong performance in multi-document question answering, but their practical deployment is limited by unreliable and unfaithful attribution to supporting evidence. Existing prompting and training-based methods often suffer from hallucinated citations and lack interpretability in how evidence is selected. In this work, we investigate whether attribution signals are inherently encoded within transformer attention mechanisms. We introduce a sensitivity-based diagnostic that identifies a small subset of attention heads, termed Evidence Heads, which are highly responsive to perturbations in supporting documents. Through causal interventions and semantic analysis, we show that these heads play a significant role in evidence identification and exhibit alignment with document-level entailment signals. Building on this finding, we propose a training-free Attention-based Attribution framework that extracts evidence signals directly from Evidence Head attention using global and local strategies. Extensive experiments show that our method consistently outperforms strong baselines while remaining lightweight and interpretable. Overall, our results suggest that structured attribution signals are implicitly encoded in LLM attention, and can be effectively leveraged for faithful multi-document reasoning.
论文标题:
Towards Practical LLM Unlearning: Efficient, Modular, and Retain-Free
作者:
Peng Liu, Peng-Fei Zhang, Jianfeng Qu, Ximing Li, Zhixu Li, Pengpeng Zhao
出版物名称:WWW '26: Proceedings of the ACM Web Conference 2026
出版时间:2026-04-13
摘要:
Large Language Models (LLMs), trained on vast web corpora and now widely integrated into web services like search engines and chatbots, have raised growing concerns about their privacy and security. This has spurred the development of Machine Unlearning techniques, which aim to effectively remove the influence of specific data (such as private user information, copyrighted web content, or harmful knowledge) from trained models while preserving general performance. However, existing LLM unlearning approaches face significant challenges, including extensive parameter modifications, a heavy reliance on retain sets to preserve utility, and limited scalability for handling sequential unlearning requests, impeding their real-world deployment in dynamic web environments. In this work, we propose Semantic Redirection for Unlearning (SRU), a lightweight framework that fine-tunes only the output word embedding layer. This targeted adjustment reshapes the model’s semantic-to-lexical mapping to block undesired concepts without disturbing deeper representations, thereby preserving the model’s general performance and eliminating dependency on retain sets. Furthermore, SRU’s modular design enables independent, sequential unlearning tasks—a vital feature for live web services handling continuous data removal requests. Experiments on the MUSE benchmark demonstrate that SRU achieves state-of-the-art efficiency, reducing computational cost by approximately 98%, while maintaining competitive unlearning performance, making it a practical and efficient solution for building compliant and trustworthy LLM-powered web applications.
论文标题:
Code LLMs Still Fall Short of Top Programmers: Evaluating Algorithmic Code Generation Through Computational Thinking
作者:
Shisong Chen, Ziyu Zhou, Yicong Zhao, Chengyi Yang, Zhixu Li, Yanghua Xiao, Xin Lin, Xiaojun Meng, Jiansheng Wei, Kuien Liu
出版物名称:Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining
出版时间:2026-02-21
摘要:
Evaluating the coding capabilities of models through algorithmic code generation is challenging, as it requires deep problem understanding and complex algorithm design. Current benchmarks suffer from a narrow focus on final execution results (such as pass@k), neglecting the crucial reasoning and problem-solving processes inherent in code generation. To address this limitation, we introduce a multi-phase algorithmic code generation benchmark, MUPA, structured around human computational thinking. MUPA dissects the evaluation into four distinct phases: example understanding, algorithm selection, solution description, and code generation. This framework facilitates a comprehensive assessment by providing insights into the model’s intermediate problem-solving steps, rather than just the final code. We manually curated 197 high-quality competitive programming problems from Codeforces. Utilizing an LLM-as-a-judge paradigm with specialized prompts, our rigorous evaluation of several existing code generation LLMs reveals significant across-the-board challenges. Notably, we establish a positive correlation, indicating that proficiency in an earlier phase directly impacts performance in subsequent phases, underscoring the interdependency of these algorithmic skills.
论文标题:
Trust the Eye or the Crowd? Unveiling Conformity in Multimodal Large Language Models
作者:
Haoran Luo, Yuhan Niu, Hengxian Liu, Yuanfei Sun, Zhihao Yang, Tong Chen, Dexin Liu, Yiyi Miao, Jingshi Zhou, Haiyang Zhang, Jionglong Su, Huixin Zhong, Yanan Liu, Wei Wang, Zimu Wang, Qi Chen
出版物名称:2026 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD)
出版时间:2026-05-13
摘要:
Multimodal large language models (MLLMs) have been propelled from passive information processors to active participants in complex social interactions. In such contexts, their ability to preserve independent judgment amidst potentially conflicting social cues becomes crucial. While recent studies have scrutinized conformity in text-only LLMs, the phenomenon remains largely unexplored in multimodal systems. To address this gap, we introduce MM-BenchForm, the first benchmark designed to systematically evaluate conformity behaviors in MLLMs. MM-BenchForm assesses model resilience across multiple cognitive capabilities, ranging from simple visual perception to advanced logic and mathematical reasoning, by leveraging established multimodal datasets such as CLEVR and PlotQA. We further propose five evaluation protocols that simulate different forms of social influence. Using this framework, we conduct extensive experiments on state-of-the-art MLLMs, including GPT-4o mini, DeepSeek-VL2, and GLM-4.5V. Our results reveal pronounced conformity tendencies: models frequently abandon correct visual evidence and instead hallucinate answers that align with induced erroneous opinions. We further analyze the impact of interaction history in shaping conformity and perform a preliminary attention-based analysis to examine how models distribute attention across visual inputs and socially provided information.
2 数据智能与知识计算
论文标题:
Caf4AVC: LLM-Enhanced Collaborative Framework for Attribute Value Canonicalization in Open KBs
作者:
Ying He, Qiang Yang, Yaxin Wang, Dan Gao, Zhouhong Gu, Zhixu Li, Yanghua Xiao
出版物名称:IEEE Transactions on Knowledge and Data Engineering
出版时间:2026-06-10
摘要:
Open Knowledge Bases (Open KBs) are fundamental to knowledge-driven applications, including semantic search, knowledge reasoning, and recommendation systems. However, the presence of redundant and ambiguous expressions within Open KBs significantly hinders their application. This highlights the urgent need for Open KB canonicalization, particularly of attribute values, which comprise nearly 40% of the facts within Open KBs. Unlike entities and predicates, attribute values are inherently sparse and diverse, posing unique challenges for their canonicalization. However, existing studies mainly focus on entities or predicates, leaving attribute value-level noun phrase canonicalization (NPC-AV) underexplored. Large language models (LLMs), with their strengths in common-sense reasoning and fault tolerance, have shown promise in Open KB canonicalization. Yet, current LLM-based approaches often rely heavily on LLM responses, overlooking their high computational cost and potential errors. In this paper, we introduce Caf4AVC, a collaborative framework that integrates clustering-based methods and LLMs for the NPC-AV task. We further propose an innovative two-factor authentication correction mechanism and an adaptive threshold-based selection strategy to address these limitations. Extensive experiments on multiple real-world Open KB datasets demonstrate the effectiveness of our framework, achieving a 17.52% reduction in LLM call costs and a 6.3% average performance improvement compared to competitive methods.
论文标题:
Tabular Data Wrangling in the Era of Large Language Model: A Survey
作者:
Ying He, Qiang Yang, Yuxiao Yang, Yubo Zhou, Zhecheng Hu, Sheng Yuan, Yanghua Xiao, Zhixu Li
出版物名称:IEEE Transactions on Knowledge and Data Engineering
出版时间:2026-07-24
摘要:
Data quality is a critical factor in data-centric artificial intelligence (AI), as “dirty data” can significantly hinder analysis and degrade downstream model performance. Among various data formats, tabular data is a ubiquitous and essential form of structured information, widely used across domains such as business, finance, healthcare, and scientific research. Tabular data wrangling, encompassing tasks, such as error correction, missing value imputation, data normalization, semantic interpretation, and the integration of heterogeneous sources, is a foundational step in the data preparation pipeline. While traditional statistical and deep learning approaches have shown competitive results, they often face limitations in robustness, manual effort, and generalization. Recent advances in large language models (LLMs), whose capabilities have expanded from unstructured text to structured data through techniques like prompt engineering and supervised fine-tuning, offer promising alternatives for addressing these challenges. Although prior surveys have examined LLM applications in downstream tabular tasks, such as question answering and fact verification, limited attention has been given to their potential in improving data quality. This survey aims to fill that gap by providing a comprehensive overview along four dimensions: 1) formal definitions and categorization of tabular wrangling tasks, 2) the typical workflow of applying LLMs to tabular data, 3) the roles LLMs can play in various wrangling methods, and 4) a comparative analysis of LLM-based versus traditional approaches, highlighting their respective strengths and limitations.
论文标题:
KUG: Joint Enhancement of Internal and External Knowledge for Retrieval-Augmented Generation
作者:
Mingyang Li, Shisong Chen, Shengkun Tu, Ziyi Du, Jinghao Zhang, Zhixu Li, Yanghua Xiao
出版物名称:Proceedings of the 34th ACM International Conference on Information and Knowledge Management
出版时间:2025-11-10
摘要:
Query enhancement, a pivotal methodology in Retrieval-Augmented Generation (RAG) for addressing information scarcity in queries, has garnered increasing research attention. Nevertheless, existing approaches overlook the inherent distinctions between domain-specific knowledge and external factual sources during integration. To bridge this gap, we propose KUG (Knowledge-Update-Generation), a novel RAG framework that leverages internal knowledge semantics to ensure query enhancement efficacy, validates and dynamically updates knowledge representations using external evidence, and achieves systematic integration through knowledge graph embeddings. Extensive experiments on six standard BEIR benchmarks demonstrate that KUG outperforms the state-of-the-art methods, achieving an improvement of 1%–2% in recall metrics. Notably, the framework demonstrates significant performance gains in multi-hop reasoning tasks, advancing the development paradigm for RAG systems. The code will be public soon.
论文标题:
ReCode: Improving LLM-based Code Repair with Fine-Grained Retrieval-Augmented Generation
作者:
Yicong Zhao, Shisong Chen, Jiacheng Zhang, Zhixu Li
出版物名称:Proceedings of the 34th ACM International Conference on Information and Knowledge Management
出版时间:2025-11-10
摘要:
Recent advances in large language models (LLMs) have demonstrated impressive capabilities in code-related tasks such as code generation and automated program repair. Despite their promising performance, most existing approaches for code repair suffer from high training costs or computationally expensive inference. Retrieval-augmented generation (RAG), with its efficient in-context learning paradigm, offers a more scalable alternative. However, conventional retrieval strategies, which are often based on holistic code-text embeddings, fail to capture the structural intricacies of code, resulting in suboptimal retrieval quality. To address the above limitations, we propose ReCode, a fine-grained retrieval-augmented in-context learning framework designed for accurate and efficient code repair. Specifically, ReCode introduces two key innovations: (1) an algorithm-aware retrieval strategy that narrows the search space using preliminary algorithm type predictions; and (2) a modular dual-encoder architecture that separately processes code and textual inputs, enabling fine-grained semantic matching between input and retrieved contexts. Furthermore, we propose RACodeBench, a new benchmark constructed from real-world user-submitted buggy code, which addresses the limitations of synthetic benchmarks and supports realistic evaluation. Experimental results on RACodeBench and competitive programming datasets demonstrate that ReCode achieves higher repair accuracy with significantly reduced inference cost, highlighting its practical value for real-world code repair scenarios.
论文标题:
Large Language Model Judged Self-Training for Named Entity Recognition
作者:
Shisong Chen, Jiaan Wang, Chengyi Yang, Yanghua Xiao, Zhixu Li, Xin Lin
出版物名称:Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining
出版时间:2026-02-21
摘要:
Self-training for Named Entity Recognition (NER) aims at identifying named entities and their types in the text using self-training to fully make use of the limited labeled data and a large amount of unlabeled data. The major challenge in self-training is confirmation bias where incorrect pseudo-labels increase errors. Many efforts have been made to address this challenge, but few labeled data limit their performance. In this paper, we introduce Large Language Model (LLM) into self-training to select high-quality pseudo-labels leveraging its rich knowledge and few-shot learning capability. Specifically, we design a comprehensive prompt to improve the judgment performance of LLM, where the prompt incorporates task rules mined by LLM itself to fully leverage labeled data. In addition, to reduce the impact of LLM’s hallucinations, we adopt a collaborative pseudo-label selection based on combined confidence and calibration-guided probability smoothing. Our empirical study conducted on several NER datasets shows that our method outperforms state-of-the-art approaches.
3 多模态智能与智能交互
论文标题:
Selecting Tangible Media for Immersive Exploration of Volumetric Scientific Data
作者:
Zhouhao Wu, Huiting Kong, Mingming Zhou, Qichen Liu, Shuai Chen, Chufan Lai, Richen Liu
出版物名称:Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems
出版时间:2026-04-13
摘要:
Immersive scientific data exploration faces challenges in precise and efficient interaction. Tangible media offer a potential solution; but designers lack clear guidance on choosing the appropriate physical dimensionality (1D, 2D, or 3D) for different tasks. To address this problem, we present a design space structuring the relationship between the representative techniques on scientific data visualization and exploration, tangible interactions, and media dimensionality. We further developed a prototype to empirically explore these relationships according to our design space. In a controlled user study, we compared 1D, 2D, and 3D tangible media across seven core techniques. The results demonstrated that the 3D media (e.g., a box) were preferred when tasks required manipulating the entire volumetric data and acted as a proxy. Regarding the tasks requiring 2D operations or interior localization, the 2D media (e.g., a card) offered superior performance. For single-parameter techniques like histogram-based filtering, the 1D media (e.g., a pen) were overwhelmingly preferred for their simplicity and perceived ease of use.
论文标题:
CausalConflictBench: Can Multimodal Models Follow Local Mechanisms That Conflict with Commonsense?
作者:
Bo Tian, Jianfeng Qu, Peng-Fei Zhang, Siyu Li, Zhixu Li, Kaiye Yu
出版物名称:The Fortieth Annual Conference on Neural Information Processing Systems
出版时间:2026-12-6(会议录用)
摘要:
Large multimodal models often benefit from commonsense priors, but these priors can mislead reasoning when a task specifies a local mechanism that conflicts with real-world regularities. Existing scientific VQA and visual reasoning benchmarks tend to align problem mechanisms with commonsense, so a correct answer may reflect either mechanism following or prior-based answering. We introduce CausalConflictBench, a diagnostic benchmark that makes this ambiguity observable by constructing samples where the current cause-to-effect mechanism contradicts the default commonsense mechanism. CausalConflictBench contains two modules: Textual Rule Override (TRO), which provides explicit commonsense-conflicting textual rules in real science image questions, and Visual Counter-Commonsense Induction (VCI), which requires models to induce a conflicting mechanism from a three-frame visual sequence. Beyond sample-level accuracy, we report Group-Strict Accuracy, factual-prior fallback metrics, and output-level process diagnostics. Across 14 proprietary and open-source multimodal models, we find that high sample-level accuracy can mask unstable mechanism following; errors concentrate strongly on factual-prior answers; and failure modes differ across rule-delivery paths, with TRO revealing rule-application failures and VCI revealing visual rule-induction failures. CausalConflictBench therefore provides a controlled diagnostic setting for uncovering commonsense fallback hidden beneath aggregate accuracy.
论文标题:
Meta-Illustrator: Transferring Illustrations from 2D Interactive Image Space to 3D Immersive Exploration Space
作者:
Richen Liu, Lingyu Sun, Xuefeng Huang, Yiran Li, Jiang Zhang, Siru Chen, Zhouhao Wu, Ayush Kumar, Chufan Lai
出版物名称:Proceedings of the 33rd ACM International Conference on Multimedia, MM 2025
出版时间:2025-10-27
摘要:
Interactive data illustrations in an immersive environment are challenging due to their inherent ambiguities during the interaction. These challenges are introduced by visual clutter and 3D occlusions resulting from depth information, as well as the relatively inefficient fine-grained manipulations required by handle controllers on immersive devices. In this paper, we propose Meta-Illustrator, an illustration transfer tool to generate immersive 3D illustrations for a volumetric data with the 2D illustrated results transferred from its one or multiple 2D slices (images). Initially, the slices can be illustrated by users expressively, owing to the plenty of the existing mature 2D sketching techniques and image processing algorithms. Then the 2D illustrated results on the slices can be intelligently transferred from their 2D image space to 3D volumetric space by Meta-Illustrator. Compared to the state-of-the-art image-to-image style transfer neural networks, which are either computation-intensive or memory-intensive, the proposed 2D-to-3D transferring approach can be built on a desktop PC without training. We demonstrate the usability, expressiveness, and effectiveness of Meta-Illustrator by both quantitative and qualitative evaluations.
论文标题:
Enhancing Chinese Multimodal Entity Linking with CLIP-RoBERTa and Contrastive Learning
作者:
Chun Wang, Chunyan An, Qiang Yang, Zhixu Li
出版物名称:Database Systems for Advanced Applications — DASFAA 2025
出版时间:2026-06-20
摘要:
Multimodal Entity Linking (MEL) is a vital task in natural language processing, aiming to associate ambiguous mentions in multimodal data—such as text and images with entities in a knowledge base (KB). However, existing methods focus primarily on English corpora, limiting their applicability to Chinese data. This challenge is further exacerbated by the scarcity of Chinese text-image datasets and semantic inconsistencies between English and Chinese languages, making it difficult to adapt English-based models for Chinese contexts. Short text labels often lack semantic richness, while noise in visual data introduces irrelevant features, reducing linking accuracy. To address these challenges, we propose MCR, a contrastive learning-based Multimodal Entity Linking method tailored for Chinese datasets. MCR leverages CLIP-RoBERTa for deep feature learning and incorporates contrastive learning to strengthen feature relationships, enabling more accurate linking of multimodal data. Additionally, we contribute a high-quality Chinese MEL dataset focused on ethnic minority elements, addressing a critical resource gap. Experimental results on this new dataset and other benchmarks demonstrate that MCR outperforms existing methods across multiple metrics, establishing it as a robust solution for MEL in both Chinese and English domains.
论文标题:
Attention distillation for low-cost driver behavior recognition
作者:
Hang Gao, Mengting Hu, Kaiye Yu, Han Xing
出版物名称:Applied Soft Computing
出版时间:2026-02-01
摘要:
Deploying artificial intelligence (AI) models for driver behavior recognition has become increasingly popular for improving driving safety. However, as vehicle intelligence advances, the rising number of AI models imposes heavy demands on onboard hardware. This study aims to develop low-cost driver behavior recognition models to alleviate such burdens, where "low-cost" refers to smaller model size and faster inference speed. To this end, we construct shallower convolutional neural networks (CNNs) by removing deeper convolutional layers and propose an innovative Attention Distillation (AD) mechanism to enhance their performance. Compared to previous distillation mechanisms, attention distillation provides more detailed knowledge transfer: it leverages the category-specific attention of deeper CNNs as the teacher to guide the attention of shallower CNNs. Experimental results show that MobileNetV2-11AD, which retains only 11 residual bottlenecks and is trained with AD, achieves higher accuracy than the original MobileNetV2 on the State Farm and AUC V2 datasets, while reducing parameters by 89.29 % and boosting the frame rate on CPU and GPU by 57.52 % and 72.65 %, respectively. Consistent gains are observed on the diverse 100-Driver dataset. These results confirm that the proposed AD mechanism provides a promising solution for low-cost driver behavior recognition. The code is available at https://github.com/gaohangcodes/AD4DBR.
4 智能计算与信息技术
论文标题:
Lizard: Bandwidth-Adaptive Real-Time Video Analytics through Content-Aware Packet Discarding at Last-Mile Edge Routers
作者:
Shan Yu, Yu Chen, Yifan Qiao, Sheng Zhang, Ravi Netravali, Harry Xu
出版物名称:11th ACM/IEEE Symposium on Edge Computing (SEC 2026)
出版时间:2026-10-13(会议录用)
摘要:
The timeliness and accuracy of edge-based video analytics can be hindered by drastic reductions in available bandwidth (ABW) at last-mile edge routers, causing prolonged queuing delays. This work proposes Lizard, a system that leverages video-content-aware packet discarding to mitigate the negative effects of drastic ABW degradation that may frequently occur at a last-mile edge router by judiciously discarding packets that contain frame blocks less important to the analytics at the destination. To achieve this, we first devise a frame-block-aware RTP header extension to effectively decouple packet dependencies to encode frame blocks. Second, Lizard uses a priority-based feedback mechanism that dynamically evaluates packet priorities based on relative accuracy impacts. Third, we develop an adaptive phase-transition-based packet discarding strategy at the router to discard packets that represent unimportant blocks. Our evaluation of Lizard shows improvements over existing methods are substantial: 53.2% reduction in latency and 27.1% increase in analysis accuracy.
论文标题:
Learning 6-DoF Fine-Grained Grasp Detection Based on Part Affordance Grounding
作者:
Yaoxian Song, Penglei Sun, Piaopiao Jin, Yi Ren, Yu Zheng, Zhixu Li, Xiaowen Chu, Yue Zhang, Tiefeng Li, Jason Gu
出版物名称:IEEE Transactions on Automation Science and Engineering
出版时间:2025-05-02
摘要:
Robotic grasping is a fundamental ability for a robot to interact with the environment. Current methods focus on how to obtain a stable and reliable grasping pose in object level, while little work has been studied on part (shape)-wise grasping which is related to fine-grained grasping and robotic affordance. Parts can be seen as atomic elements to compose an object, which contains rich semantic knowledge and a strong correlation with affordance. However, lacking a large part-wise 3D robotic dataset limits the development of part representation learning and downstream applications. In this paper, we propose a new large Language-guided SHape grAsPing datasEt (named LangSHAPE) to promote 3D part-level affordance and grasping ability learning. From the perspective of robotic cognition, we design a two-stage fine-grained robotic grasping framework (named LangPartGPD), including a novel 3D part language grounding model and a part-aware grasp pose detection model, in which explicit language input from human or large language models (LLMs) could guide a robot to generate part-level 6-DoF grasping pose with textual explanation. Our method combines the advantages of human-robot collaboration and LLMs’ planning ability using explicit language as a symbolic intermediate. To evaluate the effectiveness of our proposed method, we perform 3D part grounding and fine-grained grasp detection experiments on both simulation and physical robot settings, following language instructions across different degrees of textual complexity. Results show our method achieves competitive performance in 3D geometry fine-grained grounding, object affordance inference, and 3D part-aware grasping tasks.
论文标题:
Parsimonious Decomposition-Based Model Search Engine for Monthly Shipping Freight Rate Interval Forecasting
作者:
Jixian Mo, Jiayu Gong, Ruobin Gao, Kum Fai Yuen
出版物名称:Applied Soft Computing
出版时间:2026-07-01
摘要:
Forecasting freight rates in the shipping market is challenging due to high volatility and complex cyclical dynamics. This paper proposes a novel model, named ternary parsimonious decomposition-based model search engine (TPDMSE) for predicting monthly freight rate intervals by integrating ternary interval-valued time series (TITS) transformation, multivariate variational mode decomposition (MVMD), and a parsimonious intelligent model search engine (PIMSE) ensemble. Weekly freight rate data are first aggregated into monthly interval-valued observations to retain intra-month variability. These interval series are then fused with relevant exogenous factors and decomposed into multi-scale components using MVMD. Each component is forecasted with a rolling PIMSE strategy that adaptively selects a robust predictive model for that mode. Finally, the predicted components are summed to reconstruct forecasts of the monthly interval freight rates. Our proposed model effectively captures both trend and volatility information, yielding improved accuracy and stability over comparative models across different ship types. Further sensitivity analysis reveals that adaptive hyperparameter tuning through grid search significantly improves forecasting accuracy by precisely accommodating the distinctive volatility patterns inherent to each vessel segment. Robustness checks examining alternative weighting schemes for interval components indicate that the model consistently maintains high accuracy across different risk and preference settings. Our work contributes to bridging the gap of shipping freight rate forecasting, highlighting the benefits of combining ternary interval transformation with multivariate decomposition in maritime time series prediction.
论文标题:
Distributed quadratic interpolation estimation for large-scale quantile regression
作者:
Ziqian Qin, Yue Chao, Xuejun Ma
出版物名称:Journal of Parallel and Distributed Computing
出版时间:2026-04-01
摘要:
A number of statistical learning approaches for large-scale quantile regression (QR) have been rapidly developed to address the optimization issues arising from massive data computations. However, the principal idea behind most distributed QR estimation procedures for solving the nondifferentiable quantile loss problem is to approximate the check function using kernel-based smoothing approaches with bandwidth. In this article, we develop a new communication-efficient distributed QR estimation procedure called Distributed Quadratic Interpolation estimation strategy for QR (DQIQR) to tackle the issue posed by the limited memory constraint on a single computer machine. Specifically, we implement a quadratic function in a small neighborhood around the origin, which transforms the nondifferentiable check function into a convex and smooth quadratic loss function without using kernel-based methods. The minimizer, named the DQIQR estimator, is obtained through an approximate multi-round reweighted least squares aggregations procedure under the divide-and-conquer (DC) framework. Theoretically, we establish the asymptotic normality for the DQIQR estimator and show that our estimator achieves the same efficiency as the QR estimator computed on the entire data. Furthermore, a regularized version of DQIQR (DRQIQR) for processing distributed variable selection procedure is also investigated. Finally, the synthetic and real datasets are used to evaluate the effectiveness of the proposed approaches.
论文标题:
Personalized federated learning on quantile regression for heterogeneous missing data
作者:
Yi Lu, Zhiwei Nie, Chen Liu, Xuejun Ma
出版物名称:Knowledge-Based Systems
出版时间:2026-06-15
摘要:
To address the challenges of data heterogeneity and missingness, we propose a novel personalized federated learning quantile regression (PFLQR) method that is robust and captures the conditional quantiles of the response variable. In PFLQR, the model weights are learned collaboratively by minimizing a quadratic interpolation quantile loss with a sparse fusion penalty. Furthermore, we extend the approach to the missing-data setting by developing PFLQR-SA, which employs a simultaneous inverse probability weighting (IPW) scheme. We establish that the model weights of PFLQR and PFLQR-SA converge at rates of O (1/rpP) and O (1/rpP2), respectively. The performance of both methods is evaluated through extensive simulation studies and real-data analyses. The results demonstrate that PFLQR offers superior generalization under data heterogeneity, whereas PFLQR-SA excels in handling heterogeneous missing data.
5 人工智能与医疗健康
论文标题:
Rethinking Closed-Ended Medical VLM Evaluation: From Stable Wrongness to Counterfactual Answer Sensitivity
作者:
Qi Wu, Xiaoyu Liu, Ao Wang, Rongsheng Wang, Wenting Chen, Wenxuan Wang, Qingsong Yao
出版物名称:IEEE International Conference on Bioinformatics and Biomedicine
出版时间:2026-12-03(会议录用)
摘要:
As medical VLMs are increasingly evaluated for clinical decision-support scenarios, knowing when to trust their predictions is critical. Closed-ended medical VQA has become a common benchmark format because it enables automated scoring. Yet, systematic investigation into whether this format changes the reliability problem remains scarce. We identify a failure mode that we call stable wrongness: models can produce incorrect yes/no answers that remain confident and repeatable even when patient-specific visual evidence is removed or mismatched. This failure violates the premise behind sampling-stability uncertainty methods, including recent hallucination-aware calibration (HAC) variants: disagreement is informative only when errors are unstable. When shortcut-driven errors are themselves stable, sampling agreement can remain deceptively high. Motivated by this limitation, we introduce Counterfactual Answer Sensitivity (CAS), which shifts reliability estimation from answer stability to evidence sensitivity. CAS uses cross-modal counterfactual interventions, such as text-only, black, and corrupted images, to audit whether a committed yes/no answer remains supported when visual evidence is invalidated. Across 7 VLMs and 3 benchmarks, CAS outperforms the best HAC-family scores across all evaluated settings, achieving a macro-averaged AUROC gain of +0.150 for error detection. Mechanistic analysis further shows that CAS recovers hidden errors in the HAC-confident region and that its strongest standalone signal is counterfactual decision-boundary shift rather than hard prediction stability. Our findings suggest that closed-ended evaluation should test visual dependence, not only answer stability, in safety-critical medical contexts.
论文标题:
When Foundation Models Disagree: Representational Heterogeneity for Calibrated Uncertainty in Whole-Slide Image Classification
作者:
Yekai Shen, Yuqian Wang, Zishun Liao, Rongsheng Wang, Wenting Chen, S. Kevin Zhou, Qingsong Yao
出版物名称:IEEE International Conference on Bioinformatics and Biomedicine
出版时间:2026-12-03(会议录用)
摘要:
Pathology foundation models supply strong representations for whole-slide image classification, yet the confidence attached to their predictions remains hard to trust. Prevailing uncertainty estimators act within a single feature space and thereby overlook the uncertainty that the choice of foundation-model representation itself induces. This work intro duces CALIFUSE, a calibration-oriented framework that fuses several pathology foundation models for the specific purpose of uncertainty estimation. Departing from accuracy-driven fusion, CALIFUSE reads representational heterogeneity across encoders as an explicit confidence cue. Slide-level decision consensus is coupled with chunk-level evidence agreement quantified through Jensen–Shannon divergence, so that both prediction stability and fine-grained evidence consistency inform the estimate. All foundation models and MIL backbones stay frozen, no auxiliary calibration split is consumed, and confidence scores emerge entirely at inference. Evaluated on seven WSI classification tasks under several foundation-model combinations, the framework sharpens calibration, misclassification detection, and selective prediction relative to both the strongest single-encoder baseline and naive ensemble averaging, all while preserving competitive balanced accuracy.
论文标题:
Structure-Guided Self-Supervised Matching for One-Shot Medical Landmark Detection
作者:
Qingsong Yao, Zhen Huang, Ao Wang, Rongsheng Wang, Wenting Chen, Jianji Wang, and S. Kevin Zhou
出版物名称:IEEE International Conference on Bioinformatics and Biomedicine
出版时间:2026-12-03(会议录用)
摘要:
Medical landmark detection usually requires accurate expert annotations, which are laborious and difficult to scale across anatomical regions. In this work, we study an extreme annotation-efficient setting where only a single annotated template image is available. We propose SGB-Match, a structure-guided global-to-local self-supervised matching framework for one-shot medical landmark detection. The framework first learns dense anatomical correspondence from unlabeled augmented image pairs, and then transfers the landmark definition from the annotated template to each target image through feature matching. Different from standard contrastive correspondence learning, where negative candidates are penalized by a structure-agnostic rule, we introduce a structure-guided bias into the contrastive objective. The bias is constructed from relative distance and edge-aware anatomical cues, and explicitly reweights the negative gradients: nearby structure-relevant candidates are weakly repelled, while distant or structure-irrelevant negatives are strongly suppressed. As a result, the learned feature space better preserves local anatomical structures around template landmarks and reduces confusing responses from repeated textures. We further adopt a global-to-local design, where a global encoder provides coarse landmark localization and a local encoder refines the prediction in a cropped region. Extensive experiments on four 2D radiological landmark datasets demonstrate that SGB-Match achieves strong one-shot performance across both public and newly collected datasets and consistently benefits from both structure-guided bias and two-stage refinement.
6 人工智能与智慧治理
论文标题:
Guardians of Tomorrow: Leveraging Responsible AI for Early Detection and Response to Criminal Threats
作者:
Xiaotong Sun,Qili Wang,Liangfei Qiu,Wei Xu
出版物名称:INFORMS Journal on Computing
出版时间:2025-05-29
摘要:
Crime detection is crucial for creating and sustaining peaceful societies. In this study, we introduce Internet of Things (IoT) technology into the development of crime detection systems. Utilizing the situational crime prevention (SCP) theory from criminology, we propose a feature engineering method to extract IoT-based features that describe, explain, and predict criminal activities. In addition to commonly used features in traditional crime detection, we derive four groups of features based on SCP: criminal efforts, criminal risks, anticipated rewards, and excuses. In addition, to address growing concerns about IoT privacy issues, we incorporate a data synthesizer into our framework to generate privacy-preserving data similar to the original, allowing predictive models to be trained without accessing private information. The synthesizer uses a Bayesian network model combined with differential privacy techniques through the Laplace mechanism. Our results, based on real-world datasets, demonstrate that the proposed IoT-enabled crime detection system can achieve high-performance crime detection and has the potential to increase border surveillance efficiency with limited police resources. Our study highlights the power of artificial intelligence (AI) analytics and provides a viable framework solution for the responsible development of AI-based systems.
论文标题:
Are Large Language Models Reliable Reviewers? A Benchmark for Error Detection in Financial Documents
作者:
Ying He, Zhouhong Gu, Zhecheng Hu, Yubo Zhou, Hao Shen, Jiaqing Liang, Zhaoqian Dai, Shuguang Ma, Fei Yu, Yanghua Xiao, Zhixu Li
出版物名称:Findings of the Association for Computational Linguistics: ACL 2026
出版时间:2026-07-02
摘要:
Ensuring the accuracy of financial documents is critical for economic analysis, regulatory compliance, and corporate decision-making. Several studies have shown that Large Language Models (LLMs) perform well in many financial tasks, such as stock price movements and financial analytics. However, a critical task remains unexplored: the ability of LLMs to identify errors in financial documents. In this paper, we introduce FinED-Bench, the first publicly Benchmark for Financial Error Detection across three levels of cognitive complexity. FinED-Bench covers nine real-world financial scenarios, and includes over 900 documents reported in 2025 that are unseen by existing language models. We detail the benchmark construction process and evaluate several advanced LLMs (e.g., GPT-4o, Qwen3-14B) on this tasks, which requires both financial domain knowledge and reasoning capabilities. Experimental results show that current LLMs still struggle with this task, especially in high-complexity cases. Besides, supervised fine-tuning can significantly improve the performance of weaker LLMs on this task.
论文标题:
Geo-textual rumor detection in location-based social media by decomposing spatial subspaces
作者:
Chunyang Jiang, Bing Wang, Jianfeng Qu, Ximing Li
出版物名称:GeoInformatica
出版时间:2026-06-13
摘要:
The rapid proliferation of location-based social media platforms has greatly accelerated the dissemination of geo-tagged information, but it has also facilitated the widespread propagation of localized rumors. Geo-textual Rumor Detection (GRD) has therefore become an important research topic in geoinformatics aimed at automatically identifying deceptive content tied to specific geographical contexts. However, most existing GRD methods rely on learning static patterns from offline datasets, which limits their ability to generalize to emergent local events characterized by rapidly evolving spatial-temporal information distributions. To better understand this limitation, we conduct preliminary analyses of model fitting behaviors during training and identify two critical issues: imbalanced fitting between real and fake classes, and low-rank feature representations caused by the model's tendency to overfit to homogeneous real patterns. These phenomena directly lead to the severe loss of vital spatial information, which significantly constrains the model's capacity to capture the diverse spatial and textual patterns inherent in localized rumors. To address these challenges, we propose a novel framework named Decomposing Orthogonal Spatial Subspaces for Emergent Geo-textual rumor detection (Doseg). Our approach decomposes model transformation matrices via singular value decomposition, explicitly separating linguistic semantic, geographical spatial information-aligned, and localized event-specific spatial components while enforcing orthogonality constraints to enhance spatial feature diversity. Extensive experiments on benchmark geo-textual datasets with strict spatial-temporal splits demonstrate that our method substantially improves detection performance and increases the number of dominant principal components in feature representations, leading to stronger generalization for emergent geo-textual rumor scenarios within the geospatial ecosystem.
论文标题:
Decoupling Static Context and Dynamic Change: A Spatiotemporal Decoder With Time-Averaged Priors for Multidecadal Heritage Monitoring
作者:
Lixian Zhang, Binxiao Liu, Pengming Feng, Tianxiang Hao, Huiyue Tang, Zijian Ouyang, Ying Yang, Shaoheng Tan, Siyu Yue
出版物名称:IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing
出版时间:2026-03-03
摘要:
The long-term preservation of Urban Cultural Heritage (UCH), a key directive of the United Nations' Sustainable Development Goal 11.4, requires robust monitoring of multidecadal landscape dynamics driven by urbanization. However, analyzing the necessary long-term satellite archives reveals a significant spatiotemporal challenge. This multidecade data presents two interwoven problems: first, complex, nonstationary temporal dynamics that require processing the entire sequence history to distinguish evolving characteristics; and second, a severe spatial ambiguity in recent imagery, where heritage sites become spectrally indistinguishable from the surrounding, dynamically evolved urban sprawl. This dual challenge renders conventional methods, which are often purely spatial or purely temporal, ineffective at resolving this ambiguity. To address this, we first introduce a novel, multidecadal (1984-2024) spatiotemporal dataset for Xi'an, China, specifically curated to study this problem. We then propose a novel architectural framework designed to solve this conundrum by decoupling static spatial context from dynamic temporal evolution. Our primary contribution is a new decoding strategy that directly confronts the spatial ambiguity. Instead of sourcing skip connections from a single, ambiguous, recent frame, our decoder is fed a time-averaged composite derived from the entire sequence. This composite acts as a robust spatial prior that inherently enhances the separability between stable heritage sites and their evolved urban surroundings. Evaluated on our dataset, the proposed framework achieves state-of-the-art performance (leading by 2.6%-10.7% in mIoU) and yields a 3.2% mIoU gain attributable to the static-prior decoder in ablation studies, contributing significant improvement independent of the specific temporal model used.
免责声明:
① 凡本站注明“稿件来源:教育在线”的所有文字、图片和音视频稿件,版权均属本网所有,任何媒体、网站或个人未经本网协议授权不得转载、链接、转贴或以其他方式复制发表。已经本站协议授权的媒体、网站,在下载使用时必须注明“稿件来源:教育在线”,违者本站将依法追究责任。
② 本站注明稿件来源为其他媒体的文/图等稿件均为转载稿,本站转载出于非商业性的教育和科研之目的,并不意味着赞同其观点或证实其内容的真实性。如转载稿涉及版权等问题,请作者在两周内速来电或来函联系。




教育在线
