本文介绍了一种名为RESCUE-BENCH的关系感知多方情感支持对话系统基准,旨在评估大型语言模型(LLM)在捕捉和利用人际关系动态以提供更有效的情感支持方面的能力。该基准基于真实情侣和家庭访谈对话构建,包含丰富的社会情感和支持相关动态标注,定义了六个任务来评估关系理解和关系敏感支持两种核心能力。
要点:
1. RESCUE-BENCH基准包含191个样本、7,079个标注的对话轮次和1,064.8分钟的访谈视频。
2. 该基准定义了六个任务,涵盖关系理解和关系敏感支持两个核心能力。
3. 实验表明,当前LLM在依赖局部情感或干预线索的任务上表现较好,但在关系密集型任务上如关系模式预测、观点预测和支持策略预测方面存在困难。
4. 研究揭示了当前LLM在建模人际关系和做出关系敏感支持决策方面的局限性。
5. 该基准有助于推动LLM在情感支持对话系统中的发展,特别是在多方面人际关系处理方面。
Existing emotional support conversation systems mainly focus on one-on-one seeker-supporter interactions and individual emotional states, leaving interpersonal relations in multi-party scenarios underexplored. In this work, we introduce relation-aware emotional support conversation, a new task that evaluates whether LLMs can capture and utilize the evolving dynamics of relationships to offer more effective emotional support. We construct RESCUE (Relation-aware Emotional Support Conversation Understanding and Evaluation Benchmark) from real couple and family interview conversations, containing 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Based on rich annotations of socio-emotional and support-related dynamics, RESCUE defines six tasks that evaluate two core capabilities required for relation-aware emotional support: Relational Understanding and Relation-Sensitive Support. Experiments with ten LLMs show that current models perform relatively well on tasks relying on local emotional or intervention cues, but struggle with relation-intensive tasks such as relation pattern prediction, viewpoint prediction, and support strategy prediction. These findings reveal the limitations of current LLMs in modeling interpersonal relations and making relation-sensitive support decisions.
原始发布方:HuggingFace Daily Papers(社区热门论文)
原文时间:2026-09-09 08:00:00 +08:00
