.png)
Key Takeaways关键要点
This macroeconomic research agent analyzes GDP data across all 27 EU member states, detects anomalies, investigates structural and cyclical drivers at the sector level, and produces a 13-section cited briefing in approximately 45 minutes. Deep Agents orchestrates each research layer, LangSmith captures every step, and every finding traces back to the primary source that produced it.该宏观经济研究代理分析所有 27 个欧盟成员国的 GDP 数据,检测异常,调查部门层面的结构性和周期性驱动因素,并在约 45 分钟内生成一份包含 13 部分并带有引用的简报。Deep Agents 编排每个研究层,LangSmith 捕获每一步,所有发现均可追溯到产生它的原始来源。
The You.com Finance Research API scores 87.29% on FinSearchComp (arXiv 2509.13160), a public financial services benchmark, with a full 27-country GDP run costing roughly $2.20 in API calls. It combines licensed structured data from providers including S&P Global with live web intelligence across central bank commentary, regulatory signals, and sector-level analysis.You.com 金融研究 API 在 FinSearchComp(arXiv 2509.13160)这一公共金融服务基准上得分 87.29%,对全部 27 国的 GDP 运行约耗费 $2.20 的 API 调用费用。它结合了包括标普全球在内的授权结构化数据以及跨央行评论、监管信号和行业层面分析的实时网络情报。
Macro research desks need to know, on a regular basis, which countries in a given set are performing anomalously and why. The underlying data exists but it is fragmented. A single GDP figure might require reconciling a Eurostat release against a national statistics office publication that arrived on a different schedule and uses a different methodology. Getting from raw data to a usable, sourced briefing can be as time consuming as the analysis itself. To show this in practice, we built an agent and ran it against 2025 GDP data for all 27 EU member states.宏观研究部门需要定期了解给定集合中哪些国家表现异常以及原因。底层数据虽有,但碎片化。单一的 GDP 数字可能需要将 Eurostat 发布与国家统计局在不同时间表发布、采用不同方法的数据进行对账。将原始数据转化为可用、带来源的简报的过程耗时甚至与分析本身相当。为展示实际效果,我们构建了一个代理,并对 2025 年所有 27 个欧盟成员国的 GDP 数据运行了它。
Ireland came back as the single largest outlier, with 12.3% GDP growth that looked like a boom. Per-country investigation identified it as a pharma-led export surge front-loaded ahead of US tariffs, with the industrial sector alone contributing +6.55pp to the print. Modified GNI showed a far more modest number. Germany was flagged for the opposite reason: structural contraction driven by automotive exposure and construction collapse, not a cyclical dip. The agent produced that distinction, sourced and cited, in 45 minutes costing $2.20 in API calls.爱尔兰被识别为最大异常值,GDP 增长 12.3% 看似繁荣。按国家调查显示,这源于制药出口激增,提前于美国关税实施,工业部门单独贡献了 +6.55 个百分点。经修正的 GNI 显示的数字要温和得多。德国因相反原因被标记:结构性收缩由汽车行业敞口和建筑业崩溃驱动,而非周期性下滑。代理在 45 分钟内、耗费 $2.20 的 API 调用,生成了带来源和引用的这一区分。
The findings are only half the story. In financial services, the ability to explain how a conclusion was reached matters as much as the conclusion itself. AI agents create a gap here: without explicit instrumentation, the decisions an agent makes during a run are lost once the run completes. This architecture preserves the decision log: every query issued, every response received, and every intermediate result produced before the final report is written. LangSmith captures the complete execution trace as the agent runs, so anyone reviewing the output can follow any data point in the final report back to the source that produced it.发现只是故事的一半。在金融服务业,解释结论的形成过程与结论本身同等重要。AI 代理在这方面存在缺口:如果没有显式的仪表化,代理在运行期间的决策在运行结束后就会丢失。此架构保留了决策日志:每一次查询、每一次响应以及在最终报告撰写前产生的每一个中间结果。LangSmith 在代理运行时捕获完整的执行轨迹,任何审阅输出的人都可以将最终报告中的任意数据点追溯到产生它的来源。
The prompt提示词
Using the latest available GDP data for 2025, analyze each country within the EU economic zone. Highlight those that are increasing or decreasing at an anomalous rate. Specify and break down which industries are causing these shifts and investigate macroeconomic trends within each country that are contributing.
What the output looks like输出示例
The query has two primary questions: which EU-27 countries are growing or contracting anomalously, and what structural and cyclical forces are driving those deviations?查询有两个主要问题:哪些欧盟 27 国的增长或收缩异常,以及哪些结构性和周期性力量驱动这些偏差?
You get a structured briefing: GDP trajectory, anomaly drivers, and second-order implications including rates sensitivity, FX exposure, sovereign risk signals, and sector positioning. Every step is visible and auditable.您将获得结构化简报:GDP 轨迹、异常驱动因素以及二阶影响,包括利率敏感性、外汇敞口、主权风险信号和行业定位。每一步都是可见且可审计的。
The report follows a standard format:报告遵循标准格式:
- Executive Summary: Headline numbers, key patterns, most important finding执行摘要:头条数字、关键模式、最重要的发现
- Methodology & Data Notes: Sources used, data vintage, known caveats方法论与数据说明:使用的来源、数据时效、已知注意事项
- Regional Overview: Aggregate GDP, average growth rate, macro context地区概览:总体 GDP、平均增长率、宏观背景
- Country-by-Country GDP Table: All countries ranked by growth rate, anomaly flags, delta from mean逐国 GDP 表:按增长率排序的所有国家,异常标记,偏离均值的差值
- Multi-Year Growth Context: 3–5 year growth trajectory多年增长背景:3–5 年增长轨迹
- Anomaly Analysis, High Growth: Per-country deep dives异常分析——高增长:逐国深度剖析
- Anomaly Analysis, Low Growth / Contraction: Per-country deep dives异常分析——低增长/收缩:逐国深度剖析
- GDP Decomposition: Expenditure-side and sector-side breakdown tablesGDP 分解:支出侧和行业侧的分解表
- Structural vs Cyclical Analysis: Classification of each anomaly结构性 vs 周期性分析:对每个异常的分类
- Macroeconomic Themes & Root Causes: Cross-cutting forces宏观经济主题与根本原因:跨领域力量
- Policy Context: Monetary, fiscal, EU-level政策背景:货币、财政、欧盟层面
- Risks & Forward-Looking Assessment: Implications for the next 1-2 years风险与前瞻评估:对未来 1-2 年的影响
- Sources: Unified sequential [[n]] numbering from all workpapers来源:统一的顺序 [[n]] 编号,来源于所有工作文件


Key findings:关键发现:
- Ireland's 12.3% is driven by multinational pharma output and IP effects, not domestic activity. Modified GNI would show a far more modest number.爱尔兰的 12.3% 增长源于跨国制药产出和知识产权效应,而非国内活动。经修正的 GNI 会显示更为温和的数字。
- Major laggards share a common thread: exposure to US tariffs, Chinese competition in manufacturing, and high-rate-lag drag on construction.主要落后者有共同点:受美国关税、制造业的中国竞争以及建筑业高利率拖累的影响。
- Spain, Poland, Bulgaria, and Croatia outperformed on real wage recovery and EU fund disbursements.西班牙、波兰、保加利亚和克罗地亚在实际工资恢复和欧盟基金拨付方面表现突出。
See the full report, subagent workpapers, country-by-country breakdown, industry attribution, macroeconomic root causes, and all cited sources. View in the GitHub repository查看完整报告、子代理工作文件、逐国细分、行业归因、宏观根本原因以及所有引用来源。请在 GitHub 仓库中查看。
What Deep Agents and LangSmith make possible hereDeep Agents 和 LangSmith 在此实现的可能性
The Finance Research API handles data retrieval, reasoning, and synthesis. Give it a complex research query and it returns an answer grounded in public and private data, with inline citations. Deep Agents and LangSmith provide the engineering tools and infrastructure to build around it: context engineering, subagent management, tool execution, observability, and production deployment.金融研究 API 负责数据检索、推理和综合。给它一个复杂的研究查询,它会返回基于公共和私有数据、带内联引用的答案。Deep Agents 和 LangSmith 提供工程工具和基础设施:上下文工程、子代理管理、工具执行、可观测性和生产部署。
Context engineering. System prompts, subagents, Skills and file system management (via Backends) ensure that each subagent strictly receives only the context it needs. This allows for designing repeatable and reliable subagent behaviors.上下文工程。系统提示、子代理、技能和文件系统管理(通过后端)确保每个子代理仅收到其所需的上下文。这使得子代理行为可重复且可靠。
Subagent management. Five predefined subagents and one general-purpose subagent built into Deep Agents by default. Some run once; others fan out in multiples. The country-investigator runs one instance per anomalous country. Define it once; Deep Agents handles delegation, concurrency, failure isolation, and result aggregation.子代理管理。Deep Agents 默认内置五个预定义子代理和一个通用子代理。有些只运行一次;有些则多实例并行。country‑investigator 为每个异常国家运行一个实例。定义一次,Deep Agents 负责委派、并发、故障隔离和结果聚合。
Tool execution. The Finance Research API is one tool call. MCP servers, REST endpoints, and internal data feeds plug in the same way, scoped per subagent. In a few lines of code, you can add new tools or data sources to a particular subagent.工具执行。金融研究 API 是一次工具调用。MCP 服务器、REST 端点和内部数据流以相同方式接入,按子代理范围限定。只需几行代码,即可为特定子代理添加新工具或数据源。
Production deployment. LangSmith Deployment handles scaling, persistent storage via StoreBackend, and environment management. The same agent runs in local dev and in production without changes.生产部署。LangSmith 部署处理扩展、通过 StoreBackend 的持久化存储以及环境管理。同一代理在本地开发和生产环境中运行无需更改。
Observability. Every you_finance_research call, built-in tool call (todo list, file reads, workpaper write) and orchestrator decision is captured in LangSmith. The trace is the audit trail. It is easily accessible via CLI, MCP, as JSON export and the LangSmith UI.可观测性。每一次 you_finance_research 调用、内置工具调用(待办列表、文件读取、工作文件写入)以及编排决策都被 LangSmith 捕获。追踪即审计轨迹,可通过 CLI、MCP、JSON 导出和 LangSmith UI 轻松访问。
Implementation实现
.png)
Defining the Finance Research API tool定义金融研究 API 工具
Each subagent gets one tool: the Finance Research API. This API is itself an agent: it runs multi-step research, ingests structured public data (World Bank, IMF, OECD, Eurostat, FRED) and licensed private data, verifies sources across parallel branches, and returns cited answers with [[n]] source tags. Wrapping it as a LangChain tool makes it callable by Deep Agents.每个子代理获得一个工具:金融研究 API。该 API 本身也是一个代理:它进行多步骤研究,摄取结构化公共数据(世界银行、IMF、OECD、Eurostat、FRED)和授权私有数据,在并行分支中验证来源,并返回带 [[n]] 来源标签的引用答案。将其包装为 LangChain 工具后即可被 Deep Agents 调用。
@tool(parse_docstring=True)
async def you_finance_research(
input: str,
research_effort: Literal["deep", "exhaustive"] = "deep",
) -> str:
"""Research financial and macroeconomic topics with cited sources.
Args:
input: The research question (max 40,000 characters).
research_effort: How thorough the research should be.
"""
body = {"input": input, "research_effort": research_effort}
headers = {"Content-Type": "application/json", "X-API-Key": os.environ["YDC_API_KEY"]}
async with httpx.AsyncClient(timeout=HTTP_API_TIMEOUT) as client:
response = await client.post(HTTP_ENDPOINT, headers=headers, json=body)
data = response.json()
output = data.get("output", {})
content = output.get("content", "")
sources = output.get("sources", [])
result = content
if sources:
result += "\n\n### Sources\n"
for i, src in enumerate(sources, 1):
title = src.get("title", "Untitled")
url = src.get("url", "")
result += f"[[{i}]] {title}: {url}\n"
return result
The tool sends a research question with an effort level, pulls out the content field (with inline [[n]] citation tags) and the sources array, and appends them in a format the agent carries through to the final report. The read=None timeout is deliberate since the API can take several minutes on complex queries. The reference implementation also retries with exponential backoff on transient connection failures.工具发送研究问题和工作量级别,返回内容字段(带内联 [[n]] 引用标签)和 sources 数组,并以代理在最终报告中携带的格式追加。read=None 超时是有意为之,因为复杂查询可能需要数分钟。参考实现还在瞬时连接失败时使用指数退避重试。
You can also load the tool via MCP instead of direct HTTP. You.com exposes a hosted MCP server at https://api.you.com/mcp?tools=you-finance that works with langchain-mcp-adapters.您也可以通过 MCP 加载该工具,而不是直接使用 HTTP。You.com 在 https://api.you.com/mcp?tools=you-finance 提供托管的 MCP 服务器,可配合 langchain-mcp-adapters 使用。
Understanding the API's budget model了解 API 的预算模型
The Finance Research API has a finite compute and retrieval budget per call. It splits that budget across everything you ask in a single query:金融研究 API 对每次调用都有有限的计算和检索预算。该预算在单次查询中被分配到您请求的所有内容:
- Focused queries (one entity, one analytical question) get the full budget and return rich, quantitative answers聚焦查询(单一实体、单一分析问题)获得完整预算,返回丰富的定量答案
- Overloaded queries (many entities, many analytical dimensions) split the budget and return thin, qualitative-only answers超载查询(多实体、多分析维度)会分割预算,仅返回薄弱的定性答案
- Data-retrieval queries (e.g., "GDP growth for all 27 EU countries") are cheap per entity once the API finds the right database endpoint, so batching many countries works well数据检索查询(例如“所有 27 个欧盟国家的 GDP 增长”)在 API 找到正确数据库端点后,每个实体的成本很低,因此批量处理多个国家效果良好
This is why the agent issues focused queries rather than batching everything into a single call. Each focused call handles one analytical job and produces a discrete, attributable result, which matters as much for traceability (and hence compliance) as it does for result quality.这就是代理为何发起聚焦查询而不是将所有内容一次性调用的原因。每个聚焦调用处理一个分析任务并产生离散、可归因的结果,这对可追溯性(以及合规性)和结果质量同样重要。
Query shapes that work可行的查询形态
Three query shapes work reliably with this budget model. Each is encoded in its subagent's system prompt. The full prompts appear in the subagent definitions below.三种查询形态在此预算模型下可靠工作。每种形态都在子代理的系统提示中编码。完整提示见下文的子代理定义。
Shape A — Data Tables: "[Metric] for all [N] countries in [year(s)]"形态 A — 数据表格:“[指标] 对所有 [N] 个国家在 [年份] 的数据”
Structured data across many countries in a single call. Data retrieval is cheap per entity, so batching all 27 EU member states works fine:在一次调用中获取多个国家的结构化数据。数据检索对每个实体成本低,批量处理所有 27 个欧盟成员国完全可行:
"Provide real GDP growth rates (annual percent change, chain-linked volumes) for all 27 EU member states for each year from 2020 to 2025." "Provide current account balances as a percentage of GDP for all 27 EU member states in 2025."
Shape A calls are also your primary source layer for compliance purposes. When the Finance Research API returns GDP figures from Eurostat or IMF databases, those source URLs are included in the response and carried forward into the workpaper. The claim chain that a MiFID II records review or an EU AI Act audit requires starts here.形态 A 调用也是合规的主要来源层。当金融研究 API 返回来自 Eurostat 或 IMF 数据库的 GDP 数值时,响应中会包含这些来源的 URL,并在工作文件中保留。MiFID II 记录审查或欧盟 AI 法案审计的索赔链正从这里开始。
Shape B — Per-Country Qualitative Context: "Here are the numbers from Eurostat. What explains them?"形态 B — 单国定性背景:“以下是 Eurostat 的数字。它们的解释是什么?”
The story behind the numbers. The agent feeds in the Eurostat data it already has from Shape A and asks for causal explanations from well-indexed sources:数字背后的故事。代理将形态 A 已获取的 Eurostat 数据输入,并请求来自良好索引来源的因果解释:
"Ireland's Industry (B-E) GVA grew 29.1% in 2025 and GFCF contributed +6.32pp to GDP growth. What explains this? Was there front-loading of pharma exports ahead of US tariffs?"
"Germany's manufacturing GVA fell -0.8% and construction fell -2.9% in 2025. What specific factors explain this? Focus on: automotive production levels vs 2019, VW Group restructuring announcements."
Shape C — Mechanism Comparisons: "Compare [mechanism] across [2-3 closely related countries]"形态 C — 机制比较:“比较 [机制] 在 [2-3 个密切相关的国家] 中的表现”
How a shared mechanism played out differently across 2-3 related countries:共享机制在 2-3 个相关国家中的不同表现:
"How did ECB rate hikes in 2022-2023 affect Sweden and Denmark through their variable-rate mortgage markets? Compare with France's fixed-rate market."
Each subagent's system prompt also specifies what to avoid: don't batch 4+ countries into a single analytical query, don't combine data retrieval with interpretation in one call, and don't escalate to exhaustive when deep fails. Narrow the scope or rephrase instead.每个子代理的系统提示还规定了要避免的情况:不要在单个分析查询中批量处理 4+ 个国家,不要在一次调用中同时进行数据检索和解释,也不要在深度失败时升级为穷尽查询。应缩小范围或重新表述。
Defining the research subagents定义研究子代理
Subagents don't inherit tools from the orchestrator. Each one is explicitly configured with an LLM of our choosing, a specific task, and exclusively the Finance Research API tool. By scoping the main task down to smaller units of work via subagents, we reduce context bloat, improve predictability and optimize overall cost and speed.子代理不会继承编排器的工具。每个子代理都显式配置了我们选择的 LLM、特定任务以及唯一的金融研究 API 工具。通过子代理将主任务拆分为更小的工作单元,可降低上下文膨胀,提高可预测性并优化整体成本和速度。
landscape_scanner_subagent = {
"name": "landscape-scanner",
"description": "Retrieve structured macroeconomic data tables for all EU member states via Shape A queries.",
"system_prompt": """You are a macroeconomic data specialist...
Run 2-4 Shape A queries at `deep` effort to build complete data tables.
Write ALL results to /workpapers/landscape_scan.md as structured markdown
tables with all citations preserved.""",
"tools": [you_finance_research],
"model": "fireworks:accounts/fireworks/models/minimax-m2p5", # subagents can use a different model than the orchestrator
}
anomaly_analyst_subagent = {
"name": "anomaly-analyst",
"description": "Analyze landscape data to compute regional mean, flag anomalous countries, and recommend investigation targets. Pure statistical analysis; no Finance Research API calls.",
"system_prompt": """You are a quantitative analyst...
Read /workpapers/landscape_scan.md. Compute the unweighted mean.
Flag countries deviating by >=2.0 percentage points. Group by mechanism.
Write your full analysis to /workpapers/anomaly_analysis.md, including
a fenced JSON block at the end with investigation_targets.""",
"tools": [], # Only uses filesystem (provided by middleware)
"model": "fireworks:accounts/fireworks/models/minimax-m2p5",
}
The remaining subagents (expenditure-decomposer, sector-decomposer, country-investigator) each follow the same structure: a focused system prompt, tools=[you_finance_research], and a dedicated workpaper path. The country-investigator is fanned out once per anomalous country identified by anomaly-analyst. See the full subagent definitions in prompts.py →.其余子代理(支出分解器、行业分解器、国家调查员)均遵循相同结构:聚焦系统提示、tools=[you_finance_research],以及专用的工作文件路径。country‑investigator 会针对异常分析器识别的每个异常国家展开一次。完整子代理定义见 prompts.py →。
Creating the orchestrator agent创建编排代理
With the subagents defined, the orchestrator is assembled with create_deep_agent(). The orchestrator's system prompt has the workflow coordination logic and analytical frameworks. Query construction knowledge lives in the subagent prompts.在定义好子代理后,使用 create_deep_agent() 组装编排器。编排器的系统提示包含工作流协调逻辑和分析框架。查询构造知识位于子代理提示中。
from deepagents import create_deep_agent
from deepagents.backends import CompositeBackend, StateBackend
from deepagents.backends.filesystem import FilesystemBackend
from langgraph.checkpoint.memory import MemorySaver
backend = CompositeBackend(
default=StateBackend(),
routes={"/": FilesystemBackend(root_dir=reports_dir, virtual_mode=True)},
)
agent = create_deep_agent(
model="fireworks:accounts/fireworks/models/minimax-m2p7", # swap any LangChain-compatible model string here
tools=[],
system_prompt=system_prompt,
subagents=[
landscape_scanner_subagent,
anomaly_analyst_subagent,
expenditure_decomposer_subagent,
sector_decomposer_subagent,
country_investigator_subagent,
],
backend=backend,
checkpointer=MemorySaver(),
)
Two things to note:需要注意的两点:
CompositeBackend routes agent-internal state to StateBackend (in-memory); file writes go to FilesystemBackend on disk. Subagents write workpapers and the final report there, keeping that content out of message history, which would get unwieldy. The orchestrator reads the workpapers back during synthesis; the final report lands at /final_report.md.CompositeBackend 将代理内部状态路由到 StateBackend(内存);文件写入则使用 FilesystemBackend 写入磁盘。子代理将工作文件和最终报告写入该位置,避免将内容塞入消息历史导致膨胀。编排器在合成阶段读取这些工作文件;最终报告位于 /final_report.md。
subagents gives the orchestrator a task tool. It dispatches subagents by calling task(subagent_type="landscape-scanner", description="..."). To run subagents in parallel, the orchestrator emits multiple task calls in a single message and Deep Agents runs them concurrently.子代理为编排器提供任务工具。它通过调用 task(subagent_type="landscape-scanner", description="…") 来分发子代理。要并行运行子代理,编排器在单条消息中发出多个 task 调用,Deep Agents 会并发执行。
The multi-layer workflow多层工作流
The orchestrator starts each run by calling write_todos to lay out a research plan as a checklist, an explicit artifact it tracks against throughout the run rather than relying solely on the system prompt.编排器首先调用 write_todos 列出研究计划作为检查清单,这是它在整个运行期间跟踪的显式工件,而不是仅依赖系统提示。
1. [ ] Layer 1: Dispatch landscape-scanner for data tables
2. [ ] Layer 2: Dispatch anomaly-analyst to flag outliers
3. [ ] Layer 3a: Fan out expenditure-decomposer and sector-decomposer in parallel
4. [ ] Layer 3b: Fan out country-investigator per anomalous country
5. [ ] Cross-reference: Check workpapers for contradictions
6. [ ] Synthesize: Write final report
Each layer's results feed the next:每一层的结果喂给下一层:
Layer 1: Landscape Scan. landscape-scanner fires 2-4 Shape A calls to the Finance Research API, building data tables for all 27 EU member states. Results go to /workpapers/landscape_scan.md.第 1 层:景观扫描。landscape‑scanner 发起 2‑4 次形态 A 调用金融研究 API,为所有 27 个欧盟成员国构建数据表。结果写入 /workpapers/landscape_scan.md。
Layer 2: Anomaly Detection. anomaly-analyst reads the landscape workpaper, computes the regional mean, flags countries deviating by 2+ percentage points, and writes a full analysis to /workpapers/anomaly_analysis.md with a fenced JSON block at the end. The step involves no Finance Research API calls; it is pure analysis and computation. The orchestrator reads the workpaper and parses investigation_targets from the JSON block to decide which countries get deep follow-ups.第 2 层:异常检测。anomaly‑analyst 读取景观工作文件,计算区域均值,标记偏离 2+ 个百分点的国家,并将完整分析写入 /workpapers/anomaly_analysis.md,文件末尾附带 JSON 块。此步骤不调用金融研究 API,纯粹是分析和计算。编排器读取该工作文件并解析 JSON 块中的 investigation_targets,以决定哪些国家进入第 3 层深度跟进。
Layer 3a: Quantitative Decomposition. expenditure-decomposer and sector-decomposer run in parallel: two task calls in a single message. Each fires one Shape A query and writes its workpaper. This gives the agent the numerical backbone for the whole analysis.第 3a 层:定量分解。expenditure‑decomposer 与 sector‑decomposer 并行运行:在单条消息中发出两个 task 调用。每个调用一次形态 A 查询并写入其工作文件,为整个分析提供数值骨架。
Layer 3b: Country Investigation Fan-Out. The orchestrator reads the decomposition workpapers, picks the most interesting anomalies, and dispatches one country-investigator per country, all in parallel. Each investigator gets the country name, its key data points, and a workpaper path in the task description. Each runs Shape B queries independently and writes to its own file (/workpapers/country_ireland.md, /workpapers/country_germany.md, etc.).第 3b 层:国家调查分发。编排器读取分解工作文件,挑选最有趣的异常,并为每个国家并行分发一个 country‑investigator。每个调查员获得国家名称、关键数据点和工作文件路径。它们各自执行形态 B 查询并写入各自文件(/workpapers/country_ireland.md、/workpapers/country_germany.md 等)。
Cross-Reference. The orchestrator reads every workpaper and checks for contradictions: does the Ireland GDP figure match across the landscape scan, the expenditure decomposition, and the country investigation? If not, it dispatches the general-purpose subagent with a targeted verification query.交叉引用。编排器读取所有工作文件,检查是否存在矛盾:爱尔兰的 GDP 数据在景观扫描、支出分解和国家调查中是否一致?若不一致,则调度通用子代理进行有针对性的验证查询。
Synthesis. The orchestrator applies its analytical frameworks (expenditure decomposition, structural vs. cyclical classification, policy channel analysis), classifies each anomaly, identifies macro themes, and writes the final report to /final_report.md with unified [[n]] citation numbering.合成。编排器应用其分析框架(支出分解、结构 vs 周期分类、政策渠道分析),对每个异常进行分类,识别宏观主题,并将最终报告写入 /final_report.md,使用统一的 [[n]] 引用编号。
How the agent runs代理运行方式
A full run takes 45 minutes and ~20 API calls. The agent builds complete quantitative coverage cheaply in Layers 1 and 3a, where Shape A queries batch well; identifies the interesting stories in Layer 2; and concentrates its budget on those in Layer 3b. Every country gets hard numbers in the final report; only the real anomalies get the deep treatment.完整运行约需 45 分钟和约 20 次 API 调用。代理在第 1 层和第 3a 层以低成本构建完整的定量覆盖;第 2 层识别有趣的故事;第 3b 层将预算集中在这些故事上。每个国家在最终报告中都有硬数据;只有真正的异常获得深度处理。
Running the agent运行代理
import asyncio
from finance_research.agent import run_finance_research
report = asyncio.run(run_finance_research(
query="Using the latest available GDP data for 2025, analyze each country "
"within the EU economic zone. Highlight those that are increasing or "
"decreasing at an anomalous rate.",
preset="gdp",
))
Why observability matters for this agent为何可观测性对该代理重要
A full run involves roughly 20 Finance Research API calls and dozens of orchestrator decisions: which Shape A queries to fire, which countries get Shape B follow-ups, whether Ireland's GDP print reflects domestic activity or MNC distortion.
The trace is the record that survives the run. Any claim in the final report traces backward to the specific Finance Research API call that produced it, and from there to the primary source URL. Three regulatory frameworks make this non-negotiable for FSI deployments:一次完整运行大约涉及 20 次金融研究 API 调用和数十次编排决策:哪些形态 A 查询要发起,哪些国家进行形态 B 跟进,爱尔兰的 GDP 数据是反映国内活动还是跨国公司扭曲。追踪是运行后仍然保留的记录。最终报告中的任何主张都可以追溯到产生它的具体金融研究 API 调用,再追溯到原始来源 URL。以下三大监管框架使其在金融服务行业部署时不可协商:
- MiFID II: records obligations require firms to document the basis for investment recommendations, including AI-assisted research inputsMiFID II:记录义务要求公司记录投资建议的依据,包括 AI 辅助的研究输入
- DORA: third-party ICT oversight requires ongoing monitoring of what each vendor returned, on what input, and with what confidence; incident reporting windows require fast root-cause accessDORA:第三方 ICT 监督要求持续监控每个供应商返回的内容、输入以及置信度;事故报告窗口要求快速根因访问
- EU AI Act (Article 12): high-risk AI systems must maintain automatic event logs sufficient for post-hoc review欧盟 AI 法案(第 12 条):高风险 AI 系统必须保留自动事件日志,以供事后审查

What LangSmith capturesLangSmith 捕获的内容
The trace is built automatically without writing any instrumentation code. Every run produces a nested trace tree: orchestrator → task dispatch → subagent → you_finance_research call → response. At each node, LangSmith records input/output content, token counts (input, output, cache read, cache creation), latency, and cost. For the LLM calls, that includes the full prompt, completion, and model parameters. For tool calls, it's the arguments and return value.追踪在无需编写任何仪表化代码的情况下自动构建。每次运行产生一个嵌套的追踪树:编排器 → 任务分发 → 子代理 → you_finance_research 调用 → 响应。在每个节点,LangSmith 记录输入/输出内容、令牌计数(输入、输出、缓存读取、缓存创建)、延迟和成本。对于 LLM 调用,还包括完整提示、完成内容和模型参数。对于工具调用,则记录参数和返回值。
The practical upshot: you can click into any subagent's you_finance_research call and see the exact query, the effort level, the full cited response, and the source URLs. You can then click up one level to see how the subagent used that response in its workpaper. Any claim in the final report traces backward to the specific API call that produced it, and from there to the primary source URL.实际意义在于:您可以点击任意子代理的 you_finance_research 调用,查看具体查询、工作量级别、完整引用响应以及来源 URL。随后再点击上一级,查看子代理如何在工作文件中使用该响应。最终报告中的任何主张都可追溯到产生它的具体 API 调用,再追溯到原始来源 URL。
LangSmith's dashboards give you this view aggregated across runs, not just within a single trace. Out of the box, every project gets charts for trace count, latency percentiles (p50/p90/p99), error rates, total cost, token breakdown, and tool call frequency by name. You can build custom dashboards on top. For example, tracking the cost of Layer 3b country investigations over time, or the error rate of you_finance_research calls grouped by subagent.LangSmith 的仪表盘提供跨运行的聚合视图,而不仅限于单个追踪。默认情况下,每个项目都会生成追踪计数、延迟百分位(p50/p90/p99)、错误率、总成本、令牌分布以及按名称统计的工具调用频率等图表。您可以在此基础上构建自定义仪表盘。例如,跟踪第 3b 层国家调查随时间的成本,或按子代理分组的 you_finance_research 错误率。

What the trace shows追踪显示的内容
The first thing in any run is the orchestrator's write_todos plan, the research strategy laid out before any subagent fires. 任何运行的第一件事是编排器的 write_todos 计划,即在任何子代理启动前制定的研究策略。

Layer 1 dispatches landscape-scanner. Click into the task node and you see the subagent's Shape A queries, the data tables that came back, and the write_file call that committed results to /workpapers/landscape_scan.md.第 1 层调度 landscape‑scanner。点击任务节点即可看到子代理的形态 A 查询、返回的数据表以及写入 /workpapers/landscape_scan.md 的 write_file 调用。
Then, the orchestrator dispatches anomaly-analyst. Its trace shows the read_file loading the landscape workpaper, the statistical computation, and the write_file saving the analysis with a JSON block that the orchestrator parses for investigation targets.随后,编排器调度 anomaly‑analyst。其追踪显示读取景观工作文件、统计计算以及写入带 JSON 块的分析文件,编排器随后解析该块以获取调查目标。
Layer 3a shows expenditure-decomposer and sector-decomposer running concurrently, each with their own Finance Research API calls. Layer 3b shows the country-investigator fan-out: multiple task nodes in parallel, one per country, each with its own Shape B queries and workpaper write.第 3a 层显示支出分解器和行业分解器并行运行,各自包含自己的金融研究 API 调用。第 3b 层显示国家调查的分发:多个并行任务节点,每个国家一个,分别包含形态 B 查询和工作文件写入。
The cross-referencing step shows the orchestrator reading every workpaper, comparing figures, and deciding whether to fire verification queries. Synthesis shows the final write_file to /final_report.md.交叉引用步骤显示编排器读取所有工作文件、比较数字并决定是否发起验证查询。合成阶段显示最终写入 /final_report.md 的 write_file。
What lands in /workpapers//workpapers/ 中的内容
Every run produces 14 files:每次运行会生成 14 个文件:
- landscape_scan.md: GDP tables for all 27 member states with inline citations and the JSON block the orchestrator parses to select investigation targets.landscape_scan.md:所有 27 个成员国的 GDP 表,带内联引用和编排器用于选择调查目标的 JSON 块。
- anomaly_analysis.md: Outlier classification, structural vs. cyclical flags, and the ranked list of countries dispatched to Layer 3.anomaly_analysis.md:异常分类、结构 vs 周期标记以及被分配到第 3 层的国家排名列表。
- expenditure_decomposition.md and sector_decomposition.md: Parallel Layer 3a workpapers, each with a full breakdown table.expenditure_decomposition.md 与 sector_decomposition.md:并行的第 3a 工作文件,每个包含完整的分解表。
- country_[name] (one per anomalous country): GDP trajectory, expenditure and sector decomposition, principal finding with named mechanism, forward-looking risk assessment, and [[n]] citations throughout.country_[name](每个异常国家一个):GDP 轨迹、支出和行业分解、主要发现及命名机制、前瞻风险评估以及贯穿全文的 [[n]] 引用。
See a full example workpaper in the GitHub repository →.在 GitHub 仓库中查看完整工作文件示例 →。
File system access is handled by Deep Agents' Backends: a pluggable filesystem interface that gives each agent read_file, write_file, edit_file, ls, glob, grep backed by whatever storage you configure. For local development: FilesystemBackend(root_dir="."). For production: StoreBackend routes to Redis or Postgres via LangGraph's store interface.文件系统访问由 Deep Agents 的后端处理:可插拔的文件系统接口为每个代理提供 read_file、write_file、edit_file、ls、glob、grep,后端可自行配置。本地开发使用 FilesystemBackend(root_dir="."),生产环境使用 StoreBackend 路由至 Redis 或 Postgres(通过 LangGraph 的存储接口)。

Evaluating the agent评估代理
A research desk runs this agent on a recurring schedule: weekly GDP updates, monthly sector rotations, ad-hoc deep dives before allocation meetings. Over dozens of runs, pattern-level questions emerge that no single trace resolves: is the Finance Research API returning thinner results on certain countries? Does the 2.0 pp anomaly threshold flag too many countries in volatile quarters? Would Sonnet perform as well as Opus for the orchestrator at half the cost?研究部门会定期运行此代理:每周 GDP 更新、每月行业轮动、分配会议前的临时深度调研。经过多次运行后,会出现单个追踪无法解决的模式性问题:金融研究 API 在某些国家返回的结果是否更薄?2.0 个百分点的异常阈值是否在波动季度标记了过多国家?Sonnet 是否能以半价提供与 Opus 相同的编排性能?
LangSmith's evaluation framework is built for this, and applies in five ways:LangSmith 的评估框架专为此设计,体现在五个方面:
Offline experiments. Build a dataset of test queries with reference outputs. This can include past reports the team has validated. Run the agent against the dataset, score with evaluators, and get aggregate results. Then swap the orchestrator's core LLM, or change the anomaly threshold from 2.0 pp to 1.5 pp, and run the same dataset again. LangSmith's comparison view shows the two experiments side by side, with regressions highlighted in red and improvements in green. You can drill into any row to see the traces from both runs next to each other.离线实验。构建包含参考输出的测试查询数据集,可包括团队已验证的过去报告。对该数据集运行代理,使用评估器打分并获取汇总结果。随后更换编排器核心 LLM,或将异常阈值从 2.0pp 调整为 1.5pp,再次运行相同数据集。LangSmith 的对比视图并排展示两次实验,红色标出回归,绿色标出改进。您可以点击任意行,查看两次运行的追踪并排对比。
Custom evaluators. You can write scoring functions tailored to this workflow. For example, checking whether every [[n]] citation in the final report maps to a valid source URL, or counting how many you_finance_research calls returned "insufficient" results. These run as part of the experiment and produce scores you can track over time.自定义评估器。您可以编写针对该工作流的评分函数。例如,检查最终报告中的每个 [[n]] 引用是否映射到有效的来源 URL,或统计有多少 you_finance_research 调用返回了“不足”结果。这些评估在实验中运行,并产生可随时间跟踪的分数。
Online evaluation. Attach evaluators to production traffic. LangSmith can automatically score a sample of live runs. For example, checking that the report includes required sections, that citation numbering is sequential, or flagging runs where a subagent hit a rate limit. Runs that match evaluation criteria get extended retention for investigation.在线评估。将评估器附加到生产流量。LangSmith 可自动对实时运行的样本进行评分。例如,检查报告是否包含必需章节、引用编号是否连续,或标记子代理触发速率限制的运行。符合评估标准的运行会获得延长保留以便进一步调查。
Annotation queues. When human review is necessary, like a report that's going to a risk committee, or an output where the agent's structural-vs-cyclical classification looks borderline, runs can be routed to an annotation queue. Reviewers score against a rubric, add corrections, and those corrections feed back into the evaluation dataset for future runs. Pairwise queues let reviewers compare two versions of the same report side by side.标注队列。当需要人工审阅时,例如报告要提交给风险委员会,或代理的结构‑周期分类存在争议,运行可被路由至标注队列。审阅者依据评分标准打分、添加修正,这些修正会反馈到评估数据集用于未来运行。成对队列让审阅者并排比较同一报告的两个版本。
Tool-level analytics. Filter runs across the project by tool name to aggregate you_finance_research performance: how often it returns useful results vs. rate limits vs. "insufficient" responses, average latency by query shape, and cost per call. This is how you'd notice that Shape B queries on Nordic countries consistently come back thin, or that one subagent is burning a disproportionate share of the API budget.工具层分析。按工具名称过滤项目中的运行,以聚合 you_finance_research 的表现:返回有用结果、速率限制或“不足”响应的频率,按查询形态的平均延迟以及每次调用的成本。这可以帮助您发现北欧国家的形态 B 查询始终返回薄弱结果,或某个子代理消耗了不成比例的 API 预算。
For production, you can set alerts on these metrics: flag if error rate exceeds 5% in a 15-minute window, if average latency spikes, or if per-run cost crosses a threshold. Alerts go to Slack, PagerDuty, or a custom webhook.在生产环境中,您可以对这些指标设置警报:如果错误率在 15 分钟窗口内超过 5%,或平均延迟飙升,或单次运行成本超过阈值,则触发警报。警报可发送至 Slack、PagerDuty 或自定义 webhook。
Getting started入门指南
# Clone the reference template
git clone https://github.com/youdotcom-oss/langchain-deepagents-finance-research
The Finance Research API is available via the langchain-youdotcom package, or as a hosted MCP server at https://api.you.com/mcp (docs). Get your API key → See the integration docs →金融研究 API 可通过 langchain‑youdotcom 包获取,或在 https://api.you.com/mcp(文档)使用托管的 MCP 服务器。获取您的 API 密钥 → 查看集成文档 →
# Your You.com API key
export YDC_API_KEY=you.com_api_key
# At least one model provider
# Choose from several other LLM providers
export FIREWORKS_API_KEY="your_api_key_here"
# Enable LangSmith traces
export LANGCHAIN_API_KEY=langchain_api_key
export LANGSMITH_TRACING=true
export LANGSMITH_ENDPOINT=https://aws.api.smith.langchain.com
export LANGSMITH_PROJECT="My LangSmith project"
# Install dependencies
pip install deepagents langchain-youdotcom langchain-mcp-adapters langchain-fireworks
# Run the agent.
python examples/eu_gdp_analysis.py
To run this example, you’ll need a LangSmith account (start for free →), a You.com API key (sign up at you.com →, all new accounts come with $100 in free API credits), and an Fireworks API key. 运行此示例需要一个 LangSmith 账户(免费开始 →),以及一个 You.com API 密钥(在 you.com 注册 →,所有新账户均赠送 $100 免费 API 额度),还有 Fireworks API 密钥。
You can use several other models already available inside LangChain. To swap in a different model, set its API key, install the corresponding LangChain package, and update the model string in the agent and subagent definitions.您可以使用 LangChain 中已提供的多种模型。要切换模型,只需设置相应的 API 密钥,安装对应的 LangChain 包,并在代理和子代理定义中更新模型字符串。
Full documentation, including how to configure the anomaly detection threshold and customize the country scope, is in the integration docs →.完整文档,包括如何配置异常检测阈值和自定义国家范围,见集成文档 →。
Who this is for适用人群
This architecture fits any team running structured multi-step research on financial subjects: deal screening at PE firms, credit underwriting at banks, KYB onboarding at compliance teams, macro positioning at asset managers. The five-subagent structure here is a starting point. Add or swap tracks to fit your workflow: management background checks for compliance-heavy diligence, IP portfolio analysis for M&A screening, earnings signal aggregation for equity research. Each new track is one additional subagent dict with a focused system prompt and the same you_finance_research tool.
A no-code version is coming to Fleet, LangChain's UI-driven platform for building and managing agents.
For benchmark methodology and accuracy details, see the Finance Research API overview →
Ready to build? Get your API key → Finance Research API docs → Reference implementation on GitHub →此架构适用于任何在金融主题上进行结构化多步骤研究的团队:私募公司的项目筛选、银行的信用承销、合规团队的 KYB 入职、资产管理公司的宏观定位。这里的五子代理结构是起点。可添加或替换轨道以适配您的工作流:合规密集的尽职调查背景审查、并购筛选的 IP 组合分析、股票研究的收益信号聚合。每个新轨道都是一个额外的子代理字典,包含聚焦系统提示和相同的 you_finance_research 工具。无代码版本即将推出于 Fleet,LangChain 的 UI 驱动平台,用于构建和管理代理。有关基准方法论和准确性细节,请参阅金融研究 API 概览 →准备构建?获取您的 API 密钥 → 金融研究 API 文档 → GitHub 上的参考实现 →
Additional Resources其他资源
- LangSmith Fleet documentation →LangSmith Fleet 文档 →
- You.com integration page in LangChain docs →You.com 在 LangChain 文档中的集成页面 →
- Learn more about Deep Agents了解更多关于 Deep Agents 的信息
- Learn more about You Finance Research API了解更多关于 You Finance Research API 的信息



.png)





