MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping AgentsarXiv:2607.29002v1 Announce Type: new Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in te·AI/AI 应用, 多模态, 智能体·arxiv.org ↗arxiv.org ↗
Scaling Scientific Discovery Environments for Turn-Level Agentic RLarXiv:2607.28990v1 Announce Type: new Abstract: Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces·Research/Agent, LLM, 智能体, +1·arxiv.org ↗arxiv.org ↗
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce OperationsarXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world depl·AI/LLM, 智能体·arxiv.org ↗arxiv.org ↗
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial ObservabilityarXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scie·AI/Agent, LLM, 智能体·arxiv.org ↗arxiv.org ↗
Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental DesignarXiv:2607.28894v1 Announce Type: new Abstract: Computational cognitive modeling seeks to infer latent cognitive mechanisms underlying observed behavior. Bayesian inverse planning provides a principled framework for such·Research/推理优化·arxiv.org ↗arxiv.org ↗
Fragility of Value under Imperfect AlignmentarXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is·AI/AI 应用·arxiv.org ↗arxiv.org ↗
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI CompanionsarXiv:2607.28818v1 Announce Type: new Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that eit·AI/AI 应用·arxiv.org ↗arxiv.org ↗
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresarXiv:2607.28802v1 Announce Type: new Abstract: Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This·Research/Agent, 智能体·arxiv.org ↗arxiv.org ↗
EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter DiagnosesarXiv:2607.28788v1 Announce Type: new Abstract: Clinical diagnosis at hospital admission must be made rapidly from limited, incomplete evidence. Existing diagnosis-prediction benchmarks are poorly suited to this setting:·Research/临床医学·arxiv.org ↗arxiv.org ↗
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool AcquisitionarXiv:2607.28692v1 Announce Type: new Abstract: Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance·AI/Agent, LLM, 智能体·arxiv.org ↗arxiv.org ↗
Safety, or Just Capability? A Validity Audit of Agent-Safety BenchmarksarXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm·Research/Agent, 智能体·arxiv.org ↗arxiv.org ↗
Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic DiscoveryarXiv:2607.28684v1 Announce Type: new Abstract: Existing benchmarks for scientific equation discovery are largely composed of well-known equations available in the public domain, making it difficult to determine whether·Research/综合科学·arxiv.org ↗arxiv.org ↗
ViSAGE: Constructing Self-Correcting Memories for Long-Form Video UnderstandingarXiv:2607.28678v1 Announce Type: new Abstract: Multimodal agents operating in long-horizon environments must build and continually update multimedia memories to support entity-consistent, temporally grounded reasoning.·Research/多模态, 推理优化, 智能体·arxiv.org ↗arxiv.org ↗
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision SupportarXiv:2607.28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for sym·AI/LLM, 临床医学, 推理优化, +1·arxiv.org ↗arxiv.org ↗
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought TrajectoriesarXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods re·AI/LLM, 推理优化·arxiv.org ↗arxiv.org ↗
An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous DocumentsarXiv:2607.28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same pers·Research/LLM·arxiv.org ↗arxiv.org ↗
Empowering Cross-Domain Sequential Recommendation with Hybrid Tokenization and Serial-Parallel DecodingarXiv:2607.28659v1 Announce Type: new Abstract: Cross-domain sequential recommendation (CDSR) aims to model users' dynamic interest transitions and sequential patterns across multiple domains. Recently, generative recomm·Research/深度学习, LLM·arxiv.org ↗arxiv.org ↗
TAPR: Enhancing LLM Performance with a Task-Aware Prompt RewriterarXiv:2607.28657v1 Announce Type: new Abstract: Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the ch·AI/LLM·arxiv.org ↗arxiv.org ↗
ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon ReasoningarXiv:2607.28642v1 Announce Type: new Abstract: Long chain-of-thought reasoning improves performance on complex problems, but it also introduces redundancy accumulation, context overflow, and error anchoring. We argue th·Research/推理优化·arxiv.org ↗arxiv.org ↗
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann HypothesisarXiv:2607.28632v1 Announce Type: new Abstract: Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial ma·AI/AI 应用, LLM·arxiv.org ↗arxiv.org ↗