基本信息

要点解读

这是什么

EarthVerse 是一个用于评估科学智能体在动态地球系统和自然灾害分析中表现的基准平台,提供可复现的任务、细粒度评估标准和错误定位机制。

用在哪里

适用于开发、测试和比较科学智能体的研究团队,以及需要验证模型在复杂多源数据处理、证据选择和推理链条上可靠性的实际应用场景。

可以推断的

推测:智能体若仅关注单步准确率而忽视整体证据链的一致性,可能导致对灾害规模或影响的误判。
推测:该基准的公开有助于推动模型在科学执行层面的提升,尤其是跨证据、单位和计算的整体一致性。

来源摘要/节选

Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, and mechanism. We introduce EarthVerse, a benchmark that evaluates scientific agents through package-scoped investigations. Its 405 reproducible tasks are grounded in 199 documented events and 19 hazard families. Agents inspect heterogeneous event packages, choose compatible evidence, execute transparent calculations, reconcile source differences, and preserve provenance in the final answer. We provide executable ground truth that decomposes each task into fine-grained answer units, together with task-specific rubrics that assess the supporting research process while allowing multiple valid paths. We evaluate 25 model and agent systems under a controlled tool-using protocol, then use controlled studies to locate failures in evidence access, tool selection, memory, reasoning, interaction, and scientific execution. Across systems, the best mean answer-unit accuracy is 84.65%, while the highest Strict@95 is only 34.81%. The gap shows that current agents often complete individual steps without maintaining a consistent chain across evidence, scales, units, calculations, and physical interpretation. EarthVerse provides a reproducible basis for measuring end-to-end scientific reliability in dynamic Earth systems.

来源说明

当前保存的是来源摘要,不代表论文全文。请以原始来源为准。

「要点解读」由 AI Stack 依据上方已保存内容整理,不代表来源的完整表述;标注「推测:」的判断来自编辑,不是来源陈述。