论文解读 · 作者已确认

AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents

作者: , Linchao Zhu, Yi Yang

arXiv · 2024

LLM agentsmulti-agent interactionsocial interactionevaluationinformation exchange

论文信息

作者
Yuanzhi Liang, Linchao Zhu, and Yi Yang
推荐论文引用
Yuanzhi Liang, Linchao Zhu, and Yi Yang. “AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents.” arXiv (2024). arXiv:2401.06509.

版本日期

arXiv 首次提交
2024-01-12
arXiv 最近修订
2024-03-05
来源核验日期
2026-07-31
核验状态
作者已于 2026-07-31 核验

论文概要

AntEval 通过多 Agent 交互框架和两个定量指标评估 LLM agent 的社交互动能力:Information Exchanging Precision(IEP)与 Interaction Expressiveness Gap(IEG)。

研究问题

如果评估只关注语言是否流畅或能否闲聊,而不能量化 agent 是否交换了相关信息、是否表达了自身意图,就很难诊断和改进复杂多角色社交能力。

论文贡献

  • 设计用于激发多角色信息交换和意图表达的交互框架。
  • 提出 Information Exchanging Precision(IEP)评估信息交换。
  • 提出 Interaction Expressiveness Gap(IEG)评估交互对意图信息的表达效果。

证据与评测范围

论文通过 agent 交互实验说明 IEP 与 IEG 对互动能力诊断的用途;准确 agent 配置、交互任务和指标行为应以 arXiv v3 为准。

适用范围与局限

IEP 与 IEG 只操作化了社交互动的部分维度,并不能等价为完整的社会智能、安全性、关系质量或跨文化开放环境表现。

Related work 定位

AntEval 是面向 LLM 多 Agent 社交互动的评测工作。引用时应使用当前标题和 IEP/IEG 指标名,旧的“informativeness and expressiveness”标题已经过时。

Related Work 表述

Liang 等提出 AntEval,利用 Information Exchanging Precision 与 Interaction Expressiveness Gap 定量评估 LLM-driven agent 互动中的信息交换和意图表达。

这是一段用于说明论文定位的简洁中性表述。

论文官方英文摘要

Large Language Models (LLMs) have demonstrated their ability to replicate human behaviors across a wide range of scenarios. However, their capability in handling complex, multi-character social interactions has yet to be fully explored, primarily due to the absence of robust, quantitative evaluation methods. This gap has slowed the development of agents proficient in more nuanced interactions beyond simple exchanges, for example, small talk. To address this challenge, we introduce the Multi-Agent Interaction Evaluation Framework (AntEval), encompassing a novel interaction framework and evaluation methods. The interaction framework aims to foster an complex interaction environment that bolsters information exchange and intention expression within social interactions. Furthermore, we introduce evaluation methods, including two metrics: Information Exchanging Precision (IEP) and Interaction Expressiveness Gap (IEG), designed for the quantitative and objective assessment of agents' interaction competencies. Our findings highlight the utility of these evaluative methods and show significant potential for improving LLMs' ability to construct agents that interact in a more natural manner with human-like intricacy.

摘要仅用于学术识别,版权仍归论文作者或出版方所有,不属于本页 CC BY 许可范围。

依据与出处

核对内容论文中的位置
问题陈述Abstract
方法与贡献Abstract; interaction framework and IEP/IEG evaluation sections
评测结论Abstract; agent-interaction evaluation experiments

主要核验来源: arXiv record (arXiv:2401.06509v3, revised 2024-03-05).

如何引用

科研结论应引用论文本身;只有在复用本站原创解读时才引用本页。

Yuanzhi Liang, Linchao Zhu, and Yi Yang. “AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents.” arXiv (2024). arXiv:2401.06509.

复用许可

本页原创解读采用 CC BY 4.0:复用时须署名并链接本页。论文标题、摘要、图表和书目信息不在此许可范围内,仍保留原有权利。 Creative Commons Attribution 4.0 International.

一手来源与资源