來源:arXiv · cs.CL查看原文 ↗
原文著作權歸來源方所有,本站僅作收錄、翻譯或格式整理。
事實脈絡
arXiv:2608.07529v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existi
解讀與影響
固体废物管理涉及法规政策、工程技术、环境科学等多学科交叉,通用评测基准往往难以反映模型在此类专业任务中的实际水平。研究团队指出,现有评测体系在固体废物管理领域存在明显缺口,导致难以判断大语言模型作为技术助手的可靠性 arXiv:2608.07529。WuYuEval 的设计正是为了填补这一空白,通过构建多层次的测试任务,覆盖从基础概念理解到复杂决策支持的不同难度层级。
从方法论角度看,该基准的提出呼应了当前大语言模型评测领域的整体趋势。正如近期一篇关于评测框架的综述所指出的,有效的评估需要关注答案的忠实度、上下文精确度以及回答相关性等多个维度 machinelearningmastery.com。WuYuEval 在固体废物管理这一特定语境下,同样需要考量模型是否基于可靠信息给出建议、是否能准确理解专业术语与法规条文,以及其回答是否切实解决了用户提出的具体问题。
值得注意的是,人工智能算法在固体废物管理中的应用本身就是一个活跃的研究方向。一项近期发表的比较性系统综述梳理了多种算法在市政固体废物管理中的表现,其结论可作为行业基准参考 researchgate.net。WuYuEval 的推出,有望为这一领域的大语言模型能力评估提供标准化的工具,帮助研究者和从业者更客观地比较不同模型在废物分类、处理路径规划、政策合规咨询等任务上的优劣。关于该基准的具体任务构成、数据集规模及评估指标等细节,原文未提供更多信息。
參考來源
來源原文
arXiv:2608.07529v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as technical assistants, but their competence in solid waste management (SWM) remains difficult to assess because existing benchmarks emphasize general knowledge rather than professional decisions under engineering, environmental, and policy constraints. We introduce WuYuEval, a multi-level benchmark for evaluating LLMs in SWM across foundational knowledge, domain reasoning, and expert decision-making. After quality auditing, WuYuEval contains a Foundation Module with 4,590 closed-ended multiple-choice questions across six task types and eight domain categories, together with an Expert Module with 247 scenario-based open-ended questions involving multi-objective optimization, constraint trade-offs, and system design. For expert tasks, we combine anchor-calibrated LLM-as-a-Judge scoring with Elo-based pairwise comparison. Across 33 LLMs, performance varied widely. The leading model reached 94.64% accuracy on the Foundation Module, but average accuracy still fell from 84.14% on easy questions to 42.50% on hard questions, with lower performance concentrated in calculation, experimental design, urban planning, and open-ended expert tasks. Reasoning-oriented Thinking modes improve most matched model pairs after auditing, but the gains depend on baseline capability and are not uniformly positive. These results suggest that visible deliberation helps only when it remains anchored to units, assumptions, and engineering constraints; otherwise, it may drift from decisive answer boundaries. WuYuEval therefore provides both an evaluation resource and an empirical basis for developing SWM-oriented foundation models with professional reasoning chains and explicit constraint control.