Dark Dwarf Blog background

Agent Tool 设计

Agent Tool 设计

Agent 框架里最容易被低估的模块往往是工具层。Agent Tool 是模型与外部世界之间的唯一合法接口,接口设计得好不好,直接决定了模型能不能稳定地做事、能不能不做危险范围外的事情、能不能在复杂任务里保持可控。

这篇文章主要集中讨论工具设计下面的问题:

  1. 模型一回合发出 5 个工具调用,到底能不能并行?
  2. 同一个函数签名,怎么同时服务 LLM 和运行时两类消费者?
  3. 当工具数量变多、业务场景变复杂时,怎么防止模型”越权”或”无限循环”?

关键的设计原则是:同一份工具声明,应该被 LLM、调度器、审批层按各自的视角消费。

1. 工具是确定性与非确定性之间的契约

Anthropic 在 Writing effective tools for agents 里给了一个很准的定义:传统函数调用是确定性系统之间的契约,而工具是确定性系统与非确定性 Agent 之间的契约。getWeather("NYC") 对程序来说永远做同一件事,但 Agent 可能选择调用它、可能用常识回答、可能先反问地点,也可能根本不知道该怎么用。

这个定义解释了为什么工具设计要比普通 API 设计更保守:

  • API 的消费者是另一个程序,它会按文档传参、按状态机处理错误;
  • Agent 的消费者是一个会”想”的模型,它可能看错描述、传错参数、在不应该调用的时候调用,也可能在应该停止的时候继续调用。

所以工具层不能只负责”把函数暴露出去”,它还要解决下面的问题:

问题对应设计
这个工具会不会改外部状态?副作用元数据
模型看到的接口和运行时拿到的上下文是不是一回事?双契约签名
模型能不能随便调用任何工具?工具范围与预算约束

下面分别讲解这三个问题对应的设计。

2. 副作用元数据——工具能并发调用吗?

Function Calling 协议允许 LLM 在一轮响应里返回多个 tool_calls,但协议本身不规定这些调用要不要按顺序执行。对框架来说,这就是一个未定义的语义:如果 5 个调用都是只读的,并发能省时间;如果其中有一个会写数据库、有一个依赖前一个的结果、还有一个是”翻页”信号,并发就会出乱子。

一个简明的做法是:不在协议层猜,而在工具声明层显式标注:

# nonoka/core/execution.py
@dataclass(frozen=True)
class ToolExecution:
  """Declare the side-effect semantics of a capability.

  A capability must opt in to parallel execution with ``read_only=True``.
  Missing metadata is deliberately treated as stateful and serial; this is a
  safe compatibility default for third-party tools whose effects are unknown.
  """

  read_only: bool = False
  mutates_workspace: bool = False
  exclusive: bool = False
  stateful_action: bool = False
  pagination: bool = False

  @property
  def parallel_safe(self) -> bool:
    return (
      self.read_only
      and not self.mutates_workspace
      and not self.exclusive
      and not self.stateful_action
    )

这里每个字段都是一种声明,而不是运行时检测:

  • read_only:不修改任何外部状态;
  • mutates_workspace:会改工作区(文件、git、环境变量等);
  • exclusive:需要独占执行;
  • stateful_action:调用结果依赖或影响后续调用顺序;
  • pagination:这是一个翻页信号,通常需要等下一轮决策。

只有当 read_only=True 且其他三个”危险标志”都是 False 时,parallel_safe 才为真。最关键的默认值是:未声明的 ToolExecution 会被当成 UNKNOWN_EXECUTION,即 stateful_action=True,强制串行。

UNKNOWN_EXECUTION = ToolExecution(stateful_action=True)

a.a. 调度策略:副作用决定批次

有了副作用声明之后,调度器就可以把一回合的多个调用切成若干批次分别执行。每个批次内部是并发的,批次之间必须保持顺序:

# nonoka/core/execution.py
class ToolExecutionCoordinator:
  async def execute(self, calls, capability_for, invoke):
    results = [None] * len(calls)
    index = 0
    while index < len(calls):
      capability = capability_for(calls[index])
      if not execution_for(capability).parallel_safe:
        # 有状态 / 写入 / 未知语义:单点串行
        results[index] = await invoke(calls[index])
        index += 1
        continue

      # 连续 read_only 调用合并成一个 batch
      end = index + 1
      while end < len(calls) and execution_for(capability_for(calls[end])).parallel_safe:
        end += 1

      sem = asyncio.Semaphore(self.max_concurrency)
      async def run_one(call_index):
        async with sem:
          return call_index, await invoke(calls[call_index])

      for call_index, value in await asyncio.gather(*[
        run_one(i) for i in range(index, end)
      ]):
        results[call_index] = value
      index = end
    return results

3. 一个签名,两个受众

a.a. LLM 看到的 JSON Schema vs 运行时拿到的上下文

工具函数在 Python 里通常长这样:

@tool
async def search_records(ctx: RunContext[AppDeps], query: str, top_k: int = 5) -> list[Record]:
  ...

这个签名同时服务两类消费者:

  1. LLM:它只需要知道”这个工具叫什么、有什么用、参数是什么”。ctx 不是参数,因为 LLM 没法构造一个运行时上下文;
  2. 运行时:它需要在调用函数时注入 RunContext,让工具能访问 deps、memory、session_id 等。

nonoka 的 Tool 类在构建时就把这两件事分开了:

# nonoka/core/tool.py
class Tool(Capability):
  def __init__(self, func, ...):
    self._sig = inspect.signature(func)
    self._type_hints = get_type_hints(func)

    # 1. 找出哪个参数是 RunContext
    self._ctx_param_name = None
    for pname, _param in self._sig.parameters.items():
      hint = self._type_hints.get(pname)
      if (hint and _is_run_context_type(hint)) or pname == "ctx":
        self._ctx_param_name = pname
        break

    # 2. 生成 JSON Schema 时剔除 ctx
    self._parameters_schema, self._params_model = self._build_parameters()

  def _build_parameters(self):
    fields = {}
    for param_name, param in self._sig.parameters.items():
      if param_name == self._ctx_param_name:
        continue  # LLM 看不到 ctx
      annotation = self._type_hints.get(param_name, Any)
      default = ... if param.default == inspect.Parameter.empty else param.default
      fields[param_name] = (annotation, default)
    model = create_model(f"{self.name}_params", **fields)
    return model.model_json_schema(), model

  def to_json_schema(self):
    return {
      "type": "function",
      "function": {
        "name": self.name,
        "description": self.description,
        "parameters": self.parameters,  # 已剔除 ctx
      },
    }

运行时调用时再把 ctx 塞回去:

async def invoke(self, ctx: RunContext, arguments: dict):
  validated = self._params_model.model_validate(arguments)
  kwargs = validated.model_dump()
  if self._ctx_param_name:
    kwargs[self._ctx_param_name] = ctx  # 运行时注入
  return await self._func(**kwargs)

这样,同一个函数签名被切成两个视图:LLM 看到干净的 JSON Schema,运行时看到带泛型 RunContext[AppDeps] 的完整签名。

b.b. 装饰器边界的类型识别

RunContext[AppDeps] 是工具函数和运行时之间的”暗号”:框架看到它,就知道这个参数要由运行时注入,不能暴露给 LLM。但用户在写工具时,可能不会只写最标准的形态。同一个上下文参数,可能有四种写法:

async def t1(ctx: RunContext): ...                       # 裸类型
async def t2(ctx: RunContext[AppDeps]): ...              # 泛型别名
async def t3(ctx: Annotated[RunContext, "inject"]): ...  # PEP 593 包装
async def t4(ctx: RunContext | None): ...                # Optional / Union

如果框架只识别第一种,t2/t3/t4 里的 ctx 就会被当成普通参数,最终出现在 JSON Schema 里。LLM 看到 ctx 后要么困惑,要么编造一个它根本不该拥有的对象。

nonoka 的做法是在 decorator 边界做一次递归类型识别:

# nonoka/core/tool.py
def _is_run_context_type(hint: Any) -> bool:
  """Handles:
  * ``RunContext`` (bare)
  * ``RunContext[AppDeps]`` (generic alias)
  * ``Annotated[RunContext[AppDeps], ...]`` (PEP 593 wrapper)
  * ``Optional[RunContext]`` (Union)
  """
  # 1. 裸类型:直接就是 RunContext
  if hint is RunContext:
    return True

  # 2. 泛型别名:typing.get_origin(RunContext[AppDeps]) -> RunContext
  origin = typing.get_origin(hint)
  if origin is RunContext:
    return True

  # 3. Annotated:剥掉外层元数据,检查被包装的实际类型
  if hasattr(hint, "__metadata__") and hasattr(hint, "__args__"):
    return _is_run_context_type(hint.__args__[0])

  # 4. Union / Optional:任一成员是 RunContext 即命中
  if origin is typing.Union and hasattr(hint, "__args__"):
    return any(_is_run_context_type(arg) for arg in hint.__args__)

  return False

类型识别是 tool decorator 的第一道关卡。它做不好,后面所有注入逻辑都会站错位置:LLM 可能拿到不该看的参数,运行时可能漏掉该注入的上下文。这类 bug 在单元测试里很难发现,因为普通调用时你通常会手动传 ctx;只有当工具被 LLM 通过 JSON Schema 调用、框架层将 schema 注入 ctx 时才会暴露。

c.c. RunContext:Session 的只读视图

nonoka 的 RunContext 被刻意设计成 Session 的只读视图:

# nonoka/core/context.py
class RunContext(Generic[DepsT]):
  """Runtime context, passed to every tool invocation.

  ``RunContext`` is a read-only view of the current ``Session``.
  It gives tools access to *deps*, *memory*, and *session_id* without
  exposing mutable execution state (``completed_steps``, ``current_plan``).
  """

  def __init__(self, session: "Session"):
    self._session = session

  @property
  def deps(self) -> DepsT:
    return self._session.deps

  @property
  def session_id(self) -> str:
    return self._session.session_id

  @property
  def memory(self) -> "WorkingMemory | None":
    return self._session.memory

  @property
  def session(self) -> "Session":
    """Read-only access to the underlying ``Session``."""
    return self._session

注意 session 属性虽然返回了整个 Session 对象,但注释明确说”mutating execution state directly is discouraged”。为什么保留它?因为后面发现有些工具确实需要读取 completed_steps 或 turn_count 来做决策,但框架不鼓励也不保证这些可变状态的安全性。这是一种”留门但不鼓励”的设计。

RunContext 还提供 call_tool 重入能力,让工具内部能调用同 Session 下的其他工具,同时保持同样的上下文注入规则:

async def call_tool(self, name: str, **args: Any) -> Any:
  target_tool = next((t for t in self.agent.tools if t.name == name), None)
  if not target_tool:
    raise ValueError(f"Tool '{name}' not found in the current Agent.")
  return await target_tool.invoke(self, args)

该设计是基于 Pydantic AI 进行了些魔改实现的。

4. 意图路由与工具授权分层

当 Agent 可调用的工具变多时,“让模型自己决定调什么”会很快失控。真正可靠的做法是把”理解用户想做什么”和”决定能调用什么”分成两层:意图路由只负责语义分类,工具授权由运行时根据分类结果强制执行。

a.a. 从用户输入到工具执行的五层判断

可以把一次请求从输入到执行抽象成下面这条链路:

用户输入
  ↓
[意图路由]  这段话属于哪类语义?
  ↓
[参数提取]  用户提到了哪些可用于执行的关键信息?
  ↓
[实体验证]  这些信息能不能在原文或外部资源中找到证据?
  ↓
[工具授权]  该类意图允许使用哪些工具?
  ↓
[预算守卫]  调用次数、轮次、token、时间是否超限?
  ↓
执行

核心设计原则是:模型负责语义分类和候选参数生成,服务端负责参数校验、实体提取和权限决策。这可以将模型的不确定性限制在它擅长的领域,不让它进入事实校验和权限决策的领域。

b.b. 三层意图路由

意图路由的目标是”用尽量少的模型调用、尽量高的确定性,把输入映射到一个语义类别”。可以拆成下面的三层:

第一层:强规则识别结构信号:有些信号本身足够明确,不需要 LLM。例如任务 ID、明确的错误码、Hello 之类的问候语,都可以用正则或轻量规则直接判定:

function classifyStrongIntentWithConfig(input, hasHistory, config) {
  const trimmed = input.trim();
  if (!trimmed) return unknownIntent(...);

  const taskId = extractTaskId(trimmed);
  if (taskId) {
    return { type: 'task_failure', taskId, source: 'rule_strong', confidence: 1 };
  }

  const errorCode = extractErrorCode(trimmed);
  if (errorCode) {
    return { type: 'system_error', errorCode, source: 'rule_strong', confidence: 0.99 };
  }

  if (config.greetingPatterns.test(trimmed) || config.nonDiagnosticPatterns.test(trimmed)) {
    return { type: 'chitchat', source: 'rule_strong', confidence: 0.98 };
  }

  return undefined;  // 进入 LLM fallback
}

第二层:LLM fallback 做语义分类与候选提取:强规则未命中时,才调用 LLM。但 LLM 不是自由发挥,而是被限制成一个固定输出格式的分类器:

// src/core/intent-router.ts
const result = await this.llm.chatCompletion({
  messages: [
    {
      role: "system",
      content: [
        "你是 Agent 的意图分类器,不是工具选择器。",
        '仅输出 JSON:{"type":"...","confidence":0到1,"reason":"简短理由","candidate?":{可选候选实体}}。',
        "不确定时必须返回 unknown。follow_up 只有在 hasHistory=true 时允许。",
        "可以给出候选实体,但服务端会按类型校验;无法回查原文的候选会被丢弃或请求澄清。",
      ].join("\n"),
    },
    {
      role: "user",
      content: JSON.stringify({
        input: input.slice(0, 4000),
        hasHistory,
        ruleSignals: weak.signals,
      }),
    },
  ],
  temperature: 0,
  maxTokens: 220,
  enableThinking: false,
  tracePhase: "intent_classification",
});

这里也可以用小模型,不过我没有做过,就没写了。

注意几个关键约束:

  • 输出格式固定:返回 { type, confidence, reason, candidate? },candidate 是可选的候选实体;
  • 不返回工具名或权限:LLM 只能建议”这是什么意图”,不能决定”能调用什么工具”;
  • 温度设为 0,关闭 thinking:减少创造性输出和额外 token。

第三层:服务端校验与参数落地:LLM 输出后服务端会进行校验。不同实体的校验严格程度可以不同,比如对于下面的输出:

// 改进后的 LLM 输出格式示例
{
  "type": "task_failure",
  "confidence": 0.92,
  "reason": "用户提到任务失败",
  "candidate": { "taskId": "120000082" }
}

服务端拿到后分情况处理。校验函数的核心语义是:先接受候选,再做可验证提取;最终返回的值必须能在用户原文或可信外部资源中找到证据。

if (type === "task_failure") {
  // validateTaskId 会先检查 candidate.taskId 是否合法且在 input 中出现;
  // 如果 candidate 为空,则尝试从 input 中按规则提取;最终仍要求能在原文中验证。
  const taskId = validateTaskId(candidate?.taskId, input);
  if (!taskId) {
    return askForClarification(
      "您提到任务失败,但我没找到具体的任务 ID,请提供括号里的数字。",
    );
  }
  return { type, taskId, source: "llm_fallback" };
}

if (type === "task_inquiry") {
  const keyword = candidate?.keyword ?? extractMeaningfulKeyword(input);
  if (!keyword || !fuzzyMatchInCorpus(keyword)) {
    return askForClarification(
      "您似乎在询问某个任务,但我没定位到具体名称,请补充任务名或关键词。",
    );
  }
  return { type, keyword, source: "llm_fallback" };
}

这里的关键是不直接相信 LLM 给的实体,但也不一律降级为 unknown:

  • 对于 taskId、错误码这类强格式实体,必须能在原文中验证;
  • 对于业务关键词,可以做模糊匹配或检索式验证;
  • 校验失败时,优先给出澄清提示,而不是默默降级成 unknown。

服务端可以用不同严格程度的策略把候选参数”落地”成可执行参数。常用的策略有下面几种:

策略适用场景例子
原文回查强格式、必须在用户输入中出现的实体任务 ID、错误码
规则格式化有明确格式要求,但 LLM 输出可能不标准统一大小写、补齐前缀、标准化日期
外部资源 link自然语言别名需要映射到标准实体业务术语 → 标准任务名/物料编码/客户 ID

为什么不干脆让 LLM 输出完整 JSON,里面写好 taskId、errorCode、keyword,然后直接用?原型阶段确实可以这样,速度也更快。但生产环境里模型会”合理编造”——它可能给你一个看起来像任务 ID 的数字,或者一个听起来像业务词的关键词,而这些值并不来自用户原文。

不确定性的降级路径:以下几种情况仍然需要安全降级:

场景降级原因signals
LLM fallback 关闭功能未启用weak_rule
返回 type 不在枚举中非法类型llm_invalid_or_low_confidence
confidence < 0.75置信度不足llm_invalid_or_low_confidence
confidence > 1置信度越界llm_invalid_or_low_confidence
follow_up 但无历史上下文缺失llm_follow_up_without_history
JSON 非法 / 解析失败格式异常llm_error
LLM 调用抛异常模型异常llm_error

实体校验失败则不一定降级为 unknown,而是根据业务场景选择”请求澄清”或”保守检索”。这样比一律降级更自然,也减少了用户的挫败感。

c.c. 工具授权

意图确定之后,运行时把它映射到允许的工具集合。关键设计是:无法确认意图时,默认不暴露任何工具。

private selectToolNames(intent) {
  const scoped = this.config.agent.routingMode === 'bounded'
              || this.config.agent.toolPolicy === 'intent_scoped';
  if (!scoped) return undefined;  // all_tools / legacy 模式:暴露全部

  if (intent === 'task_failure') {
    return ['task_finder', 'code_finder', 'log_analyzer', ...];
  }
  if (intent === 'task_inquiry') {
    return ['code_finder', 'log_analyzer'];
  }
  ......

  // bounded / intent_scoped 下,无法确认意图时不暴露工具
  return [];
}

注意最后 return [] 和 return undefined 的区别:

  • undefined 会让 getToolSchemas() 暴露全部工具;
  • [] 表示”这个意图没有可用工具”,LLM 看不到任何工具,只能基于已有上下文直接回答或要求补充信息。

这是 fail-closed 设计:当系统不确定用户想做什么时,它选择缩小权限而不是放大权限。比如 task_inquiry 这种最常见的查询被限制只能用 code_finder 和 log_analyzer,禁止调用 task_finder,从而消除”查任务说明却去查任务运行实例”的越权行为。

d.d. 预算守卫与确定性路径

工具授权之后,还需要运行时守卫防止无限循环和资源耗尽。配置层面给单次请求加了多重上限:

agent: z.object({
  toolPolicy: z.enum(["all_tools", "intent_scoped"]).default("all_tools"),
  routingMode: z.enum(["legacy", "bounded"]).default("legacy"),
  maxRounds: z.coerce.number().int().min(1).max(10).default(3),
  maxToolCalls: z.coerce.number().int().min(1).max(20).default(4),
  maxRunMs: z.coerce.number().int().min(1000).default(45000),
  maxToolMs: z.coerce.number().int().min(1000).default(15000),
  maxTotalTokens: z.coerce.number().int().min(1000).default(12000),
  maxOutputTokens: z.coerce.number().int().min(64).default(800),
  maxConsecutiveNoProgress: z.coerce.number().int().min(1).max(10).default(2),
});

bounded 模式下,主循环会检查:

  • 运行时间上限:超 maxRunMs 终止,记录 run_timeout;
  • token 预算:当前 prompt tokens + 预估输出超过 maxTotalTokens 终止,记录 token_budget;
  • 工具调用次数上限:超过 maxToolCalls 终止,记录 max_tools;
  • 重复调用指纹:同一工具相同参数再次调用时终止,记录 duplicate_tool_call;
  • 连续无进展:连续多轮工具返回空结果时终止,记录 no_progress。

这些守卫让”模型不会无限绕”从一个希望变成一个可被检测、可被评测的性质。

对于 task_inquiry 这类高频意图,bounded 模式还可以走一条确定性路径:先调用 code_finder 检索一次,预加载证据后再交给解释 Agent,且解释阶段不再暴露 code_finder。这样把”检索”和”解释”拆成两个确定性步骤,而不是让模型在解释阶段还能自由调用检索工具。

e.e. 路由结果持久化,防止 drift

同一请求只做一次正式路由。路由结果应作为可恢复、可审计的决策记录,后续工具选择、诊断策略、提示词和恢复流程都复用这份结果。

const durableIntent = await executeCurrentDurableStep(
  {
    stepKey: "route:intent:v3",
    stepType: "intent_route",
    checkpointType: "ROUTE_DECISION",
    input: { userInput, hasHistory },
    projection: {
      observationName: "route.intent",
      traceOutput: (value) => {
        const routed = value as UserIntent;
        return {
          intent: routed.type,
          confidence: routed.confidence,
          signals: routed.signals,
          routerVersion: routed.routerVersion,
          source: routed.source,
        };
      },
    },
  },
  async ({ userInput: input, hasHistory: history }) =>
    this.intentRouter.classify(input, history),
);

route:intent:v3 是 durable execution 里的 step key,不是 trace 本身。可以把它理解为给”意图路由”这个操作起的持久化名字。它会被写进 checkpoint,也会通过 trace 输出 intent/confidence/signals/routerVersion/source 等字段用于观测。

将路由结果做成 durable step 持久化后有两个好处:

  1. 崩溃恢复不重复分类:checkpoint 里已经保存了路由结果,恢复后直接使用,不需要再调用 LLM;
  2. 消除 route/tool-policy drift:Agent 最初调用时把请求判成 follow_up,后续阶段又以不同历史上下文重新分类,导致同一个请求在不同阶段拥有不同的工具集。新版强制后续阶段使用入口路由结果。

f.f. 工具调用评测

工具约束有没有生效还需要看”工具行为本身规不规范”。下面是一些比较常见的评测维度:

评测维度测什么典型例子和工具约束的关系
工具策略合规性是否只调用当前意图允许的工具”查物流”请求没有调用”取消订单”直接衡量约束是否生效
意图-工具匹配度调用的工具是否与识别出的意图匹配意图是”查库存”,却调了”查物流”发现工具选择层面的问题
调用必要性是否存在冗余、重复、无进展调用同一个检索参数连续调用 3 次衡量模型是否”聪明地”使用工具
参数合规性参数是否经过验证、是否在允许范围内订单 ID 是否真的来自用户原文防止模型编造或越界参数
预算守卫触发率异常情况下预算守卫是否按预期生效构造循环请求,观察是否在 max rounds 前终止衡量兜底机制是否有效
效率指标调用次数、轮次、token、延迟平均工具调用次数、端到端耗时优化成本,但不等于合规
任务正确性最终答案是否解决了用户问题用户问 A,系统答 B独立维度,不自动由合规保证

一个健康的评测报告应该同时给出工具调用是否合规以及是否通过调用工具得到了正确的结果。