Meituan's deep-research system: an enhanced LongCat model plus a multi-agent harness whose organizing idea is that an open-ended research plan has to be executable and revisable, because the evidence a report needs cannot be specified before the investigation starts. The workflow has three stages. Explore and plan: several planners inspect the question and early sources, a judge merges their proposals, a critic names the gaps and a reviser emits an executable plan object, the ResearchSpec. Research sections: independent researchers investigate and draft their assigned sections in separate contexts, gathering more evidence as their analysis develops while retaining the full specification. Assemble and edit: sections are assembled directly, a global editor flags ownership and consistency problems, and local editors apply section-scoped revisions — so improving the report does not mean rewriting it end to end.

Reported results: 55.25 on DeepResearchBench, 51.35 on DeepResearchBench II and 79.83 on ResearchRubrics, which the report puts at +0.30, +3.17 and +5.62 points over the strongest of the three commercial systems it ran alongside (Gemini-, ChatGPT- and Claude-DeepResearch), following each benchmark's official protocol — GPT-5.5 (medium) judges the two DeepResearchBench suites, Gemini 2.5 Pro judges ResearchRubrics. The gains are concentrated in substance rather than prose: on DeepResearchBench it leads comprehensiveness (56.21), insight and depth (56.48) and instruction following (55.18) but trails on readability (49.96 against ChatGPT-DeepResearch's 51.51), and on an in-house benchmark it ranks second of four at 76.04, behind ChatGPT-DeepResearch's 76.59 and ahead of Claude-DeepResearch's 61.42 and Gemini-DeepResearch's 42.49. The same stage interfaces double as a data engine: questions and task-specific rubrics are built from an independently licensed review article or a frozen multi-source brief, passed through a bounded searchability gate that keeps a criterion only when an alternative source supports it, and the resulting teacher trajectories feed the mid-training and post-training of LongCat's general-purpose models. The report does not name the specific checkpoint behind the system. The repository (MIT) open-sources the harness only — it depends on nothing but the Python standard library and expects the user to supply an object with llm, web_search and web_fetch methods; the data-construction pipeline, generated datasets, hidden rubrics, training trajectories and training code are withheld. Authored by the Meituan LongCat Team (31 authors); the code landed on 2026-09-22 and the report on 2026-09-28.

Paper

Authors: He Zhu · Yue Xu · Wanli Wu · Haolin Ren · Yuxin Bian · Jiarui Zhao

Library

Language Python
License MIT
agentsagent-harnessagenticretrievalopen-sourceresearch

Related