CAGE: Certified Authorization under Typed-Return Uncertainty for Tool-Using AgentsarXiv:2607.29190v1 Announce Type: new Abstract: Tool-using LLM agents act on typed tool returns, records pairing provenance and categorical fields with numerical values. Runtime permission gates generally authorize the o·AI/LLM, 智能体·arxiv.org ↗arxiv.org ↗
Harnessing the Wisdom of LLM Crowds through Complementarity-Driven Iterative CollaborationarXiv:2607.29087v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in enterprise settings, yet individual models remain bounded by model-specific capability limitations. These heteroge·AI/LLM·arxiv.org ↗arxiv.org ↗
On the Generalization of Steering Vectors for Chain-of-Thought FaithfulnessarXiv:2607.29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning,·AI/AI 应用, 推理优化·arxiv.org ↗arxiv.org ↗
MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping AgentsarXiv:2607.29002v1 Announce Type: new Abstract: Online shoppers increasingly turn to AI shopping assistants, using images and multi-turn dialogue to express and refine product needs that are difficult to articulate in te·AI/AI 应用, 多模态, 智能体·arxiv.org ↗arxiv.org ↗
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce OperationsarXiv:2607.28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world depl·AI/LLM, 智能体·arxiv.org ↗arxiv.org ↗
NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial ObservabilityarXiv:2607.28942v1 Announce Type: new Abstract: Recently Large Language Models (LLMs) have been increasingly deployed as autonomous agents in applications such as self-reflection, retrieval-augmented generation, and scie·AI/Agent, LLM, 智能体·arxiv.org ↗arxiv.org ↗
Fragility of Value under Imperfect AlignmentarXiv:2607.28881v1 Announce Type: new Abstract: As more responsibility is placed upon AI systems, it becomes increasingly important to guarantee that these systems are aligned with humanity. A common fear in AI safety is·AI/AI 应用·arxiv.org ↗arxiv.org ↗
Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI CompanionsarXiv:2607.28818v1 Announce Type: new Abstract: As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that eit·AI/AI 应用·arxiv.org ↗arxiv.org ↗
SciToolAgent-Evo: An Ontology-Aware Self-Evolving Agent for Open-World Scientific Tool AcquisitionarXiv:2607.28692v1 Announce Type: new Abstract: Large language model (LLM) agents have been increasingly adopted in scientific research for organizing and invoking specialized computational tools. However, their reliance·AI/Agent, LLM, 智能体·arxiv.org ↗arxiv.org ↗
Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision SupportarXiv:2607.28677v1 Announce Type: new Abstract: LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for sym·AI/LLM, 临床医学, 推理优化, +1·arxiv.org ↗arxiv.org ↗
How Hard Does It Think? Analyzing Step-Aware Reasoning Energy in LLM Chain-of-Thought TrajectoriesarXiv:2607.28674v1 Announce Type: new Abstract: Understanding how computational effort is allocated across individual chain-of-thought (CoT) reasoning steps remains an open challenge: existing interpretability methods re·AI/LLM, 推理优化·arxiv.org ↗arxiv.org ↗
TAPR: Enhancing LLM Performance with a Task-Aware Prompt RewriterarXiv:2607.28657v1 Announce Type: new Abstract: Large Language Models (LLMs) often require carefully crafted prompts to unlock their full potential, which can be a barrier for non-expert users. This work addresses the ch·AI/LLM·arxiv.org ↗arxiv.org ↗
LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann HypothesisarXiv:2607.28632v1 Announce Type: new Abstract: Major mathematical conjectures still depend heavily on expert intuition, so a unified method for the systematic generation and validation of conjectures with substantial ma·AI/AI 应用, LLM·arxiv.org ↗arxiv.org ↗
Can AI Evaluate AI Scientists? A Benchmarking Study of Autonomous Research Generation Systems Using Automated Multi-Model ReviewarXiv:2607.28631v1 Announce Type: new Abstract: AI Scientist systems capable of autonomous research have the potential to significantly accelerate scientific discovery. However, evaluating and comparing the quality of AI·AI/AI 应用·arxiv.org ↗arxiv.org ↗
OpenClaw and Ollama in Agentic AI: Toward Fully Autonomous and Scalable AI Agent SystemsarXiv:2607.28629v1 Announce Type: new Abstract: The rapid transition from reactive large language models (LLMs) to persistent, action-capable systems has exposed critical gaps in the architectural understanding of Agenti·AI/AI 应用, Agent, LLM, +1·arxiv.org ↗arxiv.org ↗
Circles powers telco personalization with OpenAI technologyCircles uses the OpenAI API and Codex to power AI-native telco experiences, increasing ARPU by 22%, reducing churn by 9%, and improving development efficiency.·AI/AI 应用·openai.com ↗openai.com ↗
AI use mirrors student schedules in study of 77,000 online learnersHow do students actually use AI learning assistants? A new research paper by IU International University of Applied Sciences provides the first robust answers to this question. For the study "Using AI-based Learning Assi·AI/AI 应用, 国际要闻·phys.org ↗phys.org ↗
AI opens new era in cognitive studies of wild primatesScientists created an AI system that uses facial recognition and real-time touchscreen testing to automate cognitive studies of capuchin monkeys in the wild. The American Journal of Primatology published a proof-of-conce·AI/AI 应用·phys.org ↗phys.org ↗
Ten advances in mathematics and theoretical computer scienceOpenAI shares new results on long-standing open problems in mathematics and theoretical computer science, including advances in geometry, cryptography, and complexity.·AI/综合科学·openai.com ↗openai.com ↗
Claude published malicious code to the Internet and attacked 3 real companiesHad the hacks used conventional methods, someone would likely go to prison.·AI/大模型·arstechnica.com ↗arstechnica.com ↗