👍 285
07/29 20:00
Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists. As a result, scientists and AI agents need to locate relevant information, verify their provenance, and assemb
中文介绍 为了解决化学文献整合问题,论文提出了 AskChem,一个以声明为中心的基础设施,能够高效地汇集分散于多篇文献中的特定发现。该系统通过文献检索和信息验证来帮助科学家和 AI 代理获取和整合相关信息,旨在提升科研效率。意义:该研究推动了信息检索与知识集成的交叉应用,有助于加速化学研究及其相关领域的进展。
👍 270
07/29 20:00
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execute workflows across platforms, combine GUI interaction with CLI execution, complete long-horizon tas
中文介绍 论文提出 Qwen-UI-Agent,旨在打造下一代以现实世界为中心的基础 GUI 代理,具备跨平台工作流执行能力。通过结合 GUI 和 CLI 操作,系统支持长时间任务,并增强了用户交互的可靠性,推动了智能代理在真实设备上的应用。意义:此研究有助于提升人机交互的智能化水平,影响日常数字工作和自动化任务领域。
👍 243
07/28 20:00
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However, agent memory is still primarily implemented through external modules, leaving the native memory ca
中文介绍 Metis 论文提出了一种内建记忆机制的基础模型,以应对传统 AI 代理记忆依赖外部模块的问题。通过将记忆集成于基础模型中,该模型支持多模态和推理能力,增强代理的内在决策和学习效率。意义:这项研究推动了记忆集成与推理能力的创新,对智能代理和自主决策领域具有重要影响。
👍 157
07/29 20:00
Recursive self-improvement (RSI) requires AI systems that improve the process of building AI (i.e., AI4AI); machine learning engineering (MLE) offers a concrete, executable testbed for studying this capability. We introduce OpenMLE, an open full-stack system for RSI research in MLE, spanning verifia
中文介绍 论文介绍了 Frontis-MA1,一个用于机器学习工程的递归自我改进(RSI)模型,旨在提升 AI 系统在构建其他 AI 时的能力。该系统结合了验证、调试和迭代提升等多项功能,是 RSI 研究的全面方案,推动了 AI 自我优化的可能性。意义:该研究为 AI 领域的自主学习和优化奠定了基础,推动了 AI 系统的发展趋势。
👍 150
07/29 20:00
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional vi
中文介绍 PhiZero 论文提出了一种围绕物理语言构建的物理世界模型,利用紧凑的离散表示来捕捉世界状态转变。与传统模型直接在像素空间预测视频不同,PhiZero 明确了世界动态,提升了未来推测的准确性。意义:该模型在理解和模拟现实世界动态方面具有重要应用潜力,促进了物理推理与人工智能的结合。
👍 87
07/28 20:00
Existing autoregressive video distillation methods commonly adopt a Distribution Matching Distillation (DMD)-based multi-stage pipeline. However, they typically decouple the initialization and DMD stages -- which then pursue different target distributions -- and judge the intermediate student mainly
中文介绍 DistillAlign 提出了优化自回归视频蒸馏的方法,通过协调模式覆盖和模式寻找来提升蒸馏效果。通过解决初始化和多阶段蒸馏目标分布的耦合问题,本研究推动了视频生成的准确性和效率。意义:该研究为视频生成领域提供了新的思路,强化了多模态学习的结合。
👍 62
07/28 20:00
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly from a highly compressed text prompt. Existing chain-of-thought approaches introduce intermediate pla
中文介绍 VideoCoCo 通过引入代码作为链式思考(CoT)的方法,解决了文本到视频生成中物理一致性的问题。此系统结合了中间推理过程与生成策略,能够生成在物理动态上更为连贯的视频效果。意义:该研究促进了视觉生成与物理一致性结合的探索,影响视频生成与理解的相关应用。
👍 47
07/29 20:00
Decoder-only language models entangle long-term memory and reasoning in a single parameter set, making it difficult to scale memory capacity independently. Memory Decoder introduces a parametric long-term memory module but only studies it at a relatively small scale. In this work, we present Memory
中文介绍 Memory Decoder 论文提出了一种参数化的长期记忆模块,旨在解决解码器仅模型记忆与推理能力耦合的问题。通过实现可扩展的记忆能力,提升了语言生成模型的表现和应用效率。意义:该研究为语言模型的长期记忆发展提供了新视角,有助于改善生成模型在复杂任务中的应用。
👍 45
07/29 20:00
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet inefficient reasoning paradigm. In this work, we rethink agentic visual reasoning through two key d
中文介绍 Beacon 论文通过重构代理视觉推理的方法来提升多模态大语言模型的成功率,聚焦于精确与高效的推理机制。通过采用两大核心策略,研究推动了复杂任务处理的成功率。意义:此研究将加强 AI 在复杂推理任务中的应用,改善多模态交互的能力。
👍 42
07/27 20:00
Real-world tasks often require models to learn from task-specific context rather than relying only on pre-trained knowledge. While recent work has highlighted this capability as context learning, existing evaluations mainly focus on textual contexts. In many practical settings, however, the context
中文介绍 CLBench-V 关注多模态上下文学习的评估,突显了模型在任务特定上下文中的学习能力,指出现有评估主要集中在文本上下文中而忽略视觉情境。通过新任务的提出,研究提升了模型在现实应用中的适应性。意义:这一研究推动了对多模态学习能力的理解,有助于丰富自然语言处理与视觉处理的结合应用。
👍 39
07/29 20:00
Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size, leaving their accuracy-cost scaling unclear. To bridge this gap, we present a controlled study that
中文介绍 BM25 Wins at Scale 论文对 Retrieval-Augmented Generation 的各种范式进行了系统的尺度研究,揭示了不同模型在准确度与成本上的表现差异。通过控制实验,研究探讨了如何在具体语料库规模中优化性能。意义:此研究有助于提升信息检索技术的发展,为大型语言模型的应用提供更为高效的策略。
👍 34
07/29 20:00
Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modali
中文介绍 ACE-Data-0 提出了人本环境捕获的新方法,旨在解决模型在捕捉人类感知和运动等复杂动态数据时的瓶颈。该方法通过整合多种感知方式,提供了更为全面的数据支持。意义:这一研究拓展了模型在环境感知中的应用潜力,对智能体的数据采集能力提升具有重要意义。
👍 30
07/29 20:00
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding
中文介绍 Beyond Borrowed Histories 论文探讨了用户模拟在角色扮演评估中的重要性,提出使用基于人对角色的对齐进行可靠评估的方法。通过多轮对话评估,研究强化了角色扮演代理的交互体验与能力测量。意义:该研究推动了大语言模型在社交互动中的应用与进展。
👍 24
07/29 20:00
Existing video captioning models generate natural descriptions of video content but cannot explicitly ground local visual elements to multiple reference images. We introduce multi-reference image-grounded video captioning, a new task requiring factual video descriptions with phrase-level reference g
中文介绍 RefCaptioner 论文介绍了一种多参考图像基础的视频字幕生成方法,弥补了现有模型在局部视觉元素与参考图像间缺乏关联的问题。此方法使视频字幕生成更为准确与多样。意义:该研究对视频理解和生成的结合应用具有重要推动作用,增加了视觉数据的使用价值。
👍 21
07/28 20:00
Multimodal large language models increasingly use sketches, annotations, tools, and intermediate images during reasoning, but it remains unclear whether they truly rely on these visual states. Existing benchmarks are limited both by task collections with narrow coverage or partially text-solvable sa
中文介绍 See2Think 论文探讨了多模态大语言模型在推理时对中间视觉状态的真实依赖,提出了更具扩展性的评估基准。通过识别模型在处理复杂视觉任务时的表现,该研究加深了对模型推理过程的理解。意义:这一研究有助于增强人工智能在视觉推理中的应用能力,推动多模态研究的深化。
👍 21
07/29 20:00
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a fundamental capability mismatch remains: general VLMs can reason about the overall task but often m
中文介绍 SpatialCLI 论文致力于增强盲层次推理能力,使视觉语言模型能够更准确地处理空间关系问题。其通过学习空间工具的使用,后续在没有工具的情况下继续进行推理,提升了模型的决策能力。意义:该研究在推进空间推理能力和智能体逻辑决策能力方面具有重要影响。
👍 21
07/28 20:00
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program from scratch remains a major challenge: even the frontier models evaluated on ProgramBench fully resolv
中文介绍 MindForge 论文关注通过源代码无关的程序合成,教给小型语言模型整个软件工程的生命周期。该研究在代码构建方面取得了突破,为自动化代码生成提供了创新路径。意义:该研究有助于优化软件工程领域的智能应用,推动编程任务的自动化水平提高。
👍 18
07/28 20:00
LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as ProgramBench quantify this gap: given only natural-language documentation and an execute-only binary as a
中文介绍 SpecFirst 论文提出将行为规范归纳作为从头开始的程序合成过程中的关键步骤,旨在强化 LLM 代理在无现存代码的情况下构建程序的能力。该研究通过自然语言文档和执行二进制的处理,填补了当前技术的空白。意义:该研究推动了软件开发的智能化,提升了代码生成的可靠性与准确性。
👍 16
07/29 20:00
On-policy self-distillation (OPSD) is a promising approach to improve reasoning language models, but it remains brittle in practice: making it work reliably often requires substantial engineering effort. We identify a structural source of this difficulty: vanilla OPSD is precisely the β=1 member of
中文介绍 β-OPSD 论文探讨了通过策略优化衍生和自我蒸馏的训练方法,以改进推理语言模型。尽管传统方法存在脆弱性,研究识别并解决了这一结构性问题,为提升模型性能提供了新的思路。意义:该研究为推理能力和语言模型的自我学习提供了新路径,推动了语言模型的进化。
👍 16
07/29 20:00
We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models. The obstacle is representational: existing interfaces either encode an action loosely, leaving how it unfolds for the model to improvise, or encode it exactly through structured signals th
中文介绍 ShadowDancer 论文提出了一种独特的方法,通过学习统一动力学表征,实现对交互式视频世界模型的任意动作控制。研究克服了现有方法在动作表现上的局限性,为视频交互提供了新的可能性。意义:此研究对增强人机交互体验、推动视频应用技术具有潜在的影响。