Dark Dwarf Blog background

Agent 会话管理的原理

Agent 会话管理

在前面的文章中我们提到:nonoka cli 每轮对话 spawn 一个全新进程,进程跑完就自己关掉了。初步接触这个流程的人可能会有下面的疑惑:“进程死了,内存里的执行状态全没了。新进程是从头开始跑入口代码的,它凭什么知道接着上次聊?怎么知道该拿哪个 checkpoint?”

弄清这个问题需要明确下面三个东西:

  1. 代码在磁盘上(Python 源文件,随时可以重新启动);
  2. 执行状态在 SQLite 里(checkpoint 快照,进程死了它还在);
  3. 唯一需要跨进程传递的是一个短字符串——session_id。前两个都不怕进程死,这个才是需要拿到的。

下面就针对这些问题详细讲讲 Session 存储与被调用的一些具体细节:session_id 怎么在 CLI 和 Provider 之间传递、真正的状态对象 SessionState 长什么样、数据库怎么存、恢复时怎么用、以及审查历史时怎么消费这份状态。

1. Session ID 的存储

session_id 本身就是一个 uuid 字符串,它会被 nonoka opencode provider 落盘到下面的地方:

// packages/nonoka-opencode-provider/src/nonoka-language-model.ts
export function loadChatSessionId(cwd: string): string | undefined {
  const file = getChatSessionIdFile(cwd); // <cwd>/.nonoka/provider-session.id
  if (!existsSync(file)) return undefined;
  const id = readFileSync(file, "utf-8").trim();
  return id || undefined;
}

export function saveChatSessionId(
  cwd: string,
  sessionId: string | undefined,
): void {
  // ...
  fs.mkdirSync(path.dirname(file), { recursive: true });
  writeFileSync(file, sessionId, "utf-8");
}

模型实例构造时加载:

this.chatSessionId = settings.sessionId ?? loadChatSessionId(config.cwd);

前面提到过进程间通过 NDJSON bridge 协议来传递信息,ChatRequest 携带 session_id 字段;响应方向,CLI 在第一次成功响应时发出 session_init 事件把新会话的 id 交给 provider:

# src/nonoka_cli/bridge/handler.py
# Provider explicitly requested a brand-new nonoka session (e.g. /new).
if msg.new_session:
  new_id = await self._orchestrator.new_session()
  self._session_id = new_id
  self._session_init_sent = False

# Use provided session_id or the orchestrator's current session.
await self._apply_session(msg.session_id)

# Let the provider know the session id on the first successful response.
if not self._session_init_sent and self._session_id:
  await self._send(SessionInitEvent(session_id=self._session_id))
  self._session_init_sent = True

provider 收到 session_init 后写文件:

onSessionInit: (sessionId) => {
  if (isTitle) { this.titleSessionId = sessionId; return; }  // 标题生成不污染 chat session
  this.chatSessionId = sessionId;
  saveChatSessionId(this.config.cwd, sessionId);
},

2. 新会话还是续聊:一次请求的决策链

“要不要新会话”并不是某一处的一次判断,而是一条贯穿 provider 和 CLI 的决策链。每一层的判据不同,职责也不同:provider 决定“请求长什么样”,CLI 决定“最终用哪个会话”。

先给出完整的决策流程:

┌─ Provider 侧(构造请求时)──────────────────────────┐
│ 遍历 prompt:有 assistant / tool 消息吗?             │
│   ├─ 没有 → 新对话:清掉 chatSessionId,              │
│   │         请求不带 session_id                       │
│   └─ 有   → 续聊:附上 chatSessionId,               │
│             并计算 conversation_key 一并带上           │
└─────────────────────────┬────────────────────────────┘
                          ▼
┌─ CLI 侧(收到请求后,最终裁决)────────────────────────┐
│ new_session == true?                  → 强制新建      │
│ session_id 非空且库里存在?             → 续聊          │
│ session_id 非空但库里不存在?           → 降级新建 + warning │
│ session_id 为空:                                     │
│   ├─ 按 conversation_key 找回成功 → 切回旧会话,标记 recovered │
│   ├─ 找回失败但历史显示是续聊     → 新建 + 警告列出最近会话 │
│   └─ 找回失败且是新对话特征       → 正常新建           │
└───────────────────────────────────────────────────────┘

下面分层展开。

2.1 Provider 侧:看消息历史

provider 用 nonoka-language-model.ts 的 isNewConversation 做判断:OpenCode 的 /new 会把消息历史重置成 system + user 两条,所以只要 prompt 里没有 assistant / tool 消息,就判为新会话;若有,就判为继续:

private isNewConversation(options: LanguageModelV3CallOptions): boolean {
  // OpenCode's /new resets the message history to system + user only.
  for (const message of options.prompt) {
    if (message.role === 'assistant' || message.role === 'tool') {
      return false;
    }
  }
  return true;
}

判为新会话时 chatSessionId = undefined(清掉手里的 session_id),请求里 session_id 就为空;判为继续时,session_id 取自文件里读到的 id。

注意:.nonoka/provider-session.id 文件在不在,不参与这个判断,它只是 session_id 的持久化拷贝,用来跨进程恢复 SessionState 的。

2.2 CLI 侧:看请求字段

CLI 不猜意图,只看请求里的字段:new_session: true 就强制新建;session_id 非空且库里存在就续聊;带了但数据库查不到,就降级新建并记录 session_not_found_starting_new warning;session_id 为空则进入 2.4 的兜底分支。由于每个请求都是 spawn 的新 nonoka 进程、默认的新 uuid,CLI 侧必须完全依赖请求里传来的信息。

2.3 异常场景:判定为续聊,但 session_id 丢了

把两层的判据放在一起,会产生一个危险的组合:场景是续聊(历史里有 assistant)但 id 文件丢了。provider 不认为这是新对话,可它手里的 session_id 是从文件读的——文件丢了,请求里自然没有 id。CLI 收到一个”没有 id、但消息历史明显是续聊”的请求,无法区分”这是新对话”还是”这是丢了 id 的续聊”,只好按新对话处理。用户以为在接着聊,实际旧会话已经躺在库里成了孤儿。

2.4 兜底:用 conversation_key 指纹找回

解决上面这个异常场景的方案是:让每个请求自带一个“能重新算出”的指纹,id 丢了就按指纹找回。注意它只作用于 2.3 的异常分支,不参与正常的新/续判断。

  1. provider 算指纹:每个 chat 请求附带 conversation_key:首条 user 消息文本的 sha256。这个选择很关键:OpenCode 每轮传的是全量历史,所以首条 user 消息在整段对话中不变,指纹天然稳定;而且它是即时计算的、不依赖任何持久化的完整性,id 文件丢了也不影响。
// packages/nonoka-opencode-provider/src/nonoka-language-model.ts
private computeConversationKey(prompt): string | undefined {
  // Fingerprint the conversation by its first user message.
  for (const message of prompt) {
    if (message.role !== 'user') continue;
    const text = extractText(message.content);
    if (!text) return undefined;
    return createHash('sha256').update(text, 'utf8').digest('hex');
  }
  return undefined;
}
  1. CLI 存指纹:会话创建时把 conversation_key 写进 cli_sessions.metadata,查找用 json_extract(metadata, '$.conversation_key') = ? ORDER BY last_active DESC LIMIT 1,同一个指纹命中最近活跃的会话。命中就 switch_session 切回去,并把 session_init 事件标成 recovered: true + recovery_reason(session_id_missing / session_not_found),provider 收到后打日志、重新持久化 id:
# handler.py 的核心逻辑(简化)
if not session_id:
  match = await self._find_recovery_candidate(conversation_key)
  if match is not None:
    await self._try_switch(match.session_id)
    self._mark_recovered("session_id_missing", match.session_id)
  elif self._has_continuation_messages(msg):
    logger.warning("session_id_missing_starting_fresh",
                   recent_sessions=await self._recent_sessions_summary())
  1. 兜底保留:指纹也找不到、但历史证明这是续聊时,返回 session_id_missing_starting_fresh 警告并列出最近活跃会话。

到这一步,我们解决了“新进程怎么拿到同一个 session_id”。但 session_id 只是一把钥匙,真正要恢复的状态是 SessionState——下一节讲它里面到底存了什么。

3. SessionState 里到底存了什么

nonoka core 把会话的一切状态塞在一个 Pydantic 模型里:SessionState,定义在 nonoka/core/session.py:62-92。

class SessionState(BaseModel):
  """
  Immutable snapshot of a Session.

  This is a pure data object used to save to a database or deserialise
  from a database to restore a session.
  """
  schema_version: int = SESSION_STATE_SCHEMA_VERSION
  session_id: str
  status: SessionStatus

  current_plan: Plan | None = None

  completed_steps: dict[str, StepResult] = Field(default_factory=dict)
  failed_steps: dict[str, StepFailure] = Field(default_factory=dict)
  step_statuses: dict[str, StepStatus] = Field(default_factory=dict)

  # Memory snapshot for conversational checkpoint/resume
  memory_entries: list[dict[str, Any]] = Field(default_factory=list)

  start_time: datetime | None = None
  end_time: datetime | None = None
  turn_count: int = 0
  step_count: int = 0
  trace: dict[str, Any] | None = None
  runtime_state: SessionRuntimeState | None = None
  completion_contract: CompletionContract | None = None
  # Namespaced, JSON-serialisable state owned by loop extensions. Keeping it
  # in the checkpoint prevents external-tool pause/resume boundaries from
  # resetting progress detectors and bounded repair counters.
  extension_state: dict[str, dict[str, Any]] = Field(default_factory=dict)

注意源码注释里的关键词:immutable snapshot。SessionState 是不可变的纯数据对象,负责进出数据库;运行时有一个对应的可变对象 Session,负责预算检查、工具调用预留、上下文压缩、取消等逻辑。Session.to_state() 把运行时状态序列化成 SessionState,Session.from_state() 把 SessionState 还原成 Session。

按职责可以把 SessionState 分成四块:

SessionState
├── 元数据
│   ├── schema_version      # 结构版本号,迁移用
│   ├── session_id          # uuid
│   └── status              # created / running / paused / completed / failed / cancelled
│
├── ① 计划与步骤
│   ├── current_plan        # PlanExecutor 使用的 DAG
│   ├── step_statuses       # 每个 step 的当前状态
│   ├── completed_steps     # step 成功产出(dict[str, StepResult])
│   └── failed_steps        # step 失败信息(dict[str, StepFailure])
│
├── ② 对话历史
│   └── memory_entries      # WorkingMemory.entries 的快照
│
├── ③ 运行时账本
│   └── runtime_state       # SessionRuntimeState:limits / usage / deadline / termination
│
└── ④ 扩展层
    ├── completion_contract # 完成契约
    └── extension_state     # 插件状态(HITL、检测器、修复计数器等)

3.1 计划与步骤

current_plan 是 Plan 类型,定义在 nonoka/core/plan.py:76-77,包含 steps 与预计算的 layers(按依赖分层)。PlanExecutor 按层推进,同层 step 可以并行。

step_statuses、completed_steps、failed_steps 三个字典一起记录 DAG 的执行进度。completed_steps 不是列表,而是 dict[str, StepResult]——这个细节在之前的文章里写错了,这里纠正过来。键是 step id,值是 step 完成时的产出。

3.2 对话历史:memory_entries

memory_entries 是 WorkingMemory.entries 的快照。每条 entry 是一个 MemoryEntry,定义在 nonoka/core/memory.py:22-27:

class MemoryEntry(BaseModel):
  role: MemoryRole
  content: str
  metadata: dict[str, Any] = Field(default_factory=dict)
  tokens: int = 0  # Token count

role 可以是 user / assistant / tool / system 等。assistant 消息在 metadata.tool_calls 里带工具调用请求;tool 消息在 metadata.tool_call_id 里带对应的调用 id。判断一次工具调用是否完成,就是看 tool_calls[].id 有没有配对的 tool_call_id——恢复悬空调用时用的就是这个机制。

WorkingMemory 本身在 nonoka/core/memory.py:442-480,负责上下文窗口、token 预算、压缩、可选的长期 memory backend。SessionState 里只存 memory_entries 这个序列化后的列表,不直接持有 WorkingMemory。

3.3 运行时账本:runtime_state

runtime_state 不是简单的 dict,而是 SessionRuntimeState 模型,定义在 nonoka/core/runtime.py:280-298:

class SessionRuntimeState(BaseModel):
  """Checkpoint-owned runtime state. Resumes must restore, never recreate, it."""

  limits: RuntimeLimits
  usage: RuntimeUsage = Field(default_factory=RuntimeUsage)
  started_at: datetime = Field(default_factory=lambda: datetime.now(timezone.utc))
  deadline_at: datetime | None = None
  cancelled_at: datetime | None = None
  termination: Termination | None = None

注释说得很清楚:Resumes must restore, never recreate it。恢复时必须复用旧的 runtime state,而不是新建一个。否则已经消耗的 token 数、已经经过的 wall time、已经调用的工具数都会被重置,导致预算被绕过。

3.4 扩展层

completion_contract 定义完成契约:满足什么条件才算任务完成。extension_state 是插件的命名空间状态,比如 HITL 审批状态、外部工具的进度检测器、bounded repair 计数器。把它放在 SessionState 里,是为了防止进程重启后这些状态被重置。

4. 持久化:整份快照 + 增量补丁

SessionState 被序列化成一个 JSON,存在 checkpoints.state_json 这一列。全量快照的写入在 SQLiteCheckpointStore.save_session(),nonoka/backends/checkpoint/sqlite.py:157-175:

async def save_session(self, session_id: str, state: SessionState) -> None:
  """Persist a full session snapshot."""

  def _save() -> None:
    conn = self._ensure_connection()
    conn.execute(
      """
      INSERT INTO checkpoints (session_id, state_json)
      VALUES (?, ?)
      ON CONFLICT(session_id) DO UPDATE SET
        state_json = excluded.state_json,
        updated_at = CURRENT_TIMESTAMP
      """,
      (session_id, self._state_to_json(state)),
    )
    conn.commit()

  async with self._lock:
    await asyncio.to_thread(_save)

但每走一步都重写整个 JSON 太浪费,所以步骤级别的变化走另一张表 step_updates。save_step_status / save_step_result / save_step_error 在 sqlite.py:232-309:

async def save_step_result(
  self, session_id: str, step_id: str, result: Any
) -> None:
  """Persist a step result and mark it completed."""

  def _save() -> None:
    conn = self._ensure_connection()
    payload = json.dumps(
      {"status": StepStatus.COMPLETED.value, "result": result},
      default=str,
    )
    conn.execute(
      """
      INSERT INTO step_updates (session_id, step_id, update_type, payload_json)
      VALUES (?, ?, ?, ?)
      ON CONFLICT(session_id, step_id, update_type) DO UPDATE SET
        payload_json = excluded.payload_json,
        updated_at = CURRENT_TIMESTAMP
      """,
      (session_id, step_id, "result", payload),
    )
    conn.commit()

  async with self._lock:
    await asyncio.to_thread(_save)

step_updates 的复合主键是 (session_id, step_id, update_type),同一 step 同一类型只保留最新一条。update_type 有三种:

update_type含义合并到快照
statusstep 状态变化step_statuses[step_id]
resultstep 完成产出step_statuses[step_id] = COMPLETED + completed_steps[step_id] = StepResult(...)
errorstep 失败step_statuses[step_id] = FAILED + failed_steps[step_id] = StepFailure(...)

加载时在 SQLiteCheckpointStore.load_session(),sqlite.py:177-230:

  1. 读出 checkpoints.state_json;
  2. 反序列化并做 schema 迁移;
  3. 按 updated_at ASC 读取 step_updates,一条条合并到快照里。

合并有三条保护规则,直接看源码注释:

if update_type == "status":
  new_status = StepStatus(payload["status"])
  # Don't overwrite a terminal status (COMPLETED / FAILED) with a
  # non-terminal one.  A later "status" row may race with an earlier
  # "result" / "error" row when timestamps have identical resolution.
  current = state.step_statuses.get(step_id)
  if current in (StepStatus.COMPLETED, StepStatus.FAILED):
    continue
  state.step_statuses[step_id] = new_status
elif update_type == "result":
  state.step_statuses[step_id] = StepStatus.COMPLETED
  state.completed_steps[step_id] = StepResult(data=payload["result"])
  # A successful retry should clear any previous failure record.
  state.failed_steps.pop(step_id, None)
elif update_type == "error":
  state.step_statuses[step_id] = StepStatus.FAILED
  state.failed_steps[step_id] = StepFailure(
    error_type=payload["error"]["error_type"],
    message=payload["error"]["message"],
    traceback=payload["error"].get("traceback"),
  )
  # Ensure a failed step doesn't also carry a stale success record.
  state.completed_steps.pop(step_id, None)

这三条规则分别是:

  1. 终态不被非终态覆盖:已经是 completed/failed 的 step,不会被一个 later 的 status 更新回 running。
  2. 成功时清掉旧失败:某 step 失败后重试成功,旧失败记录不再保留。
  3. 失败时清掉旧成功:某 step 成功后又被判失败(比如幂等重试时出错),旧成功记录不再保留。

这样设计的好处是:数据库层只负责“整份存 + 小增量补”,复杂的业务一致性规则写在应用层,调试和测试都更容易。

5. 恢复流程:从 state_json 回到运行

拿到 session_id 之后,恢复入口是 Runner.resume(),nonoka/core/runner.py:747-794:

async def resume(
  self,
  agent: Agent[DepsT, ResultT],
  session_id: str,
  deps: DepsT,
) -> RunResult[ResultT]:
  """Resume execution from a checkpoint.

  The session's ``current_plan`` determines which paradigm to resume:
  * If a Plan exists → resume via ``PlanExecutor``.
  * Otherwise → resume via ``ReActAgent``.
  """
  state = await self.checkpoint_store.load_session(session_id)
  if not state:
    return RunResult(success=False, error=f"Session {session_id} not found in checkpoint store.")

  # Re-create WorkingMemory so that memory_entries can be restored
  memory = None
  if self.memory_backend is not None:
    from nonoka.core.memory import WorkingMemory
    resumed_limits = getattr(getattr(state, "runtime_state", None), "limits", None)
    summary_llm = (
      self._ensure_llm(agent)
      if getattr(resumed_limits, "summary_enabled", False)
      else None
    )
    memory = WorkingMemory(
      session_id=session_id,
      memory_backend=self.memory_backend,
      summary_llm=summary_llm,
    )

  session = Session.from_state(state, agent, deps=deps, memory=memory)

  if session.status in {SessionStatus.COMPLETED, SessionStatus.FAILED}:
    return RunResult(success=session.status == SessionStatus.COMPLETED, session=session)

  self._ensure_llm(agent)

  # Route to the correct paradigm based on whether a plan was in flight
  if session.current_plan and session.current_plan.steps:
    from nonoka.core.paradigm import PlanExecutor
    executor = PlanExecutor()
    return await executor.resume(session.current_plan, session, self)
  else:
    from nonoka.core.paradigm import ReActAgent
    paradigm = ReActAgent()
    return await paradigm.resume(session, self)

这段代码是整个恢复流程的骨架,可以拆成四步:

  1. load_session(session_id) 读出 SessionState;
  2. 重建 WorkingMemory;
  3. Session.from_state(state, ...) 把不可变快照还原成可变运行时对象;
  4. 根据有没有 current_plan,路由到 PlanExecutor.resume() 或 ReActAgent.resume()。

5.1 从快照还原 WorkingMemory

Session.from_state() 在 nonoka/core/session.py:441-480。关键在最后几行:

# Restore memory entries if memory is provided and state has a snapshot
if memory is not None and state.memory_entries:
  from nonoka.core.memory import MemoryEntry
  memory.entries = [MemoryEntry(**entry) for entry in state.memory_entries]

也就是说,对话历史是随 checkpoint 一起恢复的。之前早期版本有过“带 session_id 的 run_react 会新建 WorkingMemory”的 bug,当前源码已经修复:只要 session_id 存在,_create_session()(runner.py:370-419)就会 load_session 并通过 Session.from_state() 把 memory_entries 还原到 WorkingMemory.entries。

5.2 ReAct 悬空 tool_calls 重放

如果崩溃发生在 assistant 已经发出 tool_calls、但 tool 结果还没落库的瞬间,恢复时就要补跑这些悬空调用。逻辑在 ReActAgent.resume() 与 _replay_dangling_tool_calls(),nonoka/core/paradigm.py:500-584:

async def resume(self, session: Session, runner: Any) -> RunResult:
  """Resume a conversational session from checkpoint.

  A checkpoint saved mid-turn may contain an assistant tool_calls message
  without its tool results (the process crashed between the two).  Repair
  that first by re-executing the dangling calls, then continue the loop.

  Trade-off: replaying a call whose pre-crash execution did finish means
  its side effects may be applied once more; git checkpoint/rollback is
  the mitigation for workspace-mutating tools.
  """
  await self._replay_dangling_tool_calls(session, runner)
  return await self.run(session, runner, prompt="")

_replay_dangling_tool_calls() 的核心扫描逻辑:

# Find the last assistant message carrying tool_calls (same scan as
# resume_approval) and keep only calls with no TOOL result after it.
assistant_index: int | None = None
pending_tool_calls: list[dict[str, Any]] = []
for i in range(len(session.memory.entries) - 1, -1, -1):
  entry = session.memory.entries[i]
  if entry.role == MemoryRole.ASSISTANT and entry.metadata.get("tool_calls"):
    assistant_index = i
    pending_tool_calls = list(entry.metadata["tool_calls"])
    break

然后从该 assistant 消息往后扫,收集已经回应的 tool_call_id:

answered: set[str] = set()
for entry in session.memory.entries[assistant_index + 1:]:
  if entry.role == MemoryRole.TOOL:
    tc_id = entry.metadata.get("tool_call_id")
    if tc_id:
      answered.add(str(tc_id))

差集就是悬空调用:

dangling = [
  tc for tc in pending_tool_calls
  if str(tc.get("id") or tc.get("tool_call_id", "unknown")) not in answered
]

这些调用会被 _execute_tool_calls() 重新执行,结果写回 memory,然后存一次新的 checkpoint。

注意这里的关键词:at-least-once。崩溃点可能在“工具已经执行完、但结果没来得及落库”的位置,所以重放可能导致副作用发生两次。缓解办法是 workspace-mutating 工具配合 git checkpoint/rollback。更详细的讨论已经在《SQLite 在 Agent 中的使用》那篇里写过,这里不再展开。

5.3 PlanExecutor 恢复

PlanExecutor 的恢复更简单:直接调用 execute(),后者会跳过 session.completed_steps 里已经完成的 step。入口在 nonoka/core/paradigm.py:2595-2597:

async def resume(self, plan: Plan, session: Session, runner: Any) -> RunResult:
  """Resume plan execution from checkpoint (skips completed steps)."""
  return await self.execute(plan, session, runner)

跳过逻辑在 execute() 里,paradigm.py:2389-2395:

# Skip steps already completed (resume scenario)
pending_ids = [
  sid for sid in layer_step_ids
  if sid not in session.completed_steps
]
if not pending_ids:
  continue

因为步骤状态已经持久化在 completed_steps / step_statuses 里,PlanExecutor 不需要重算,直接从断点继续按层执行即可。

6. 审查与查看历史

恢复和审查本质上做的是同一件事:把 SessionState 从 SQLite 里读出来。区别只是恢复要重建执行上下文,审查要把状态给人看。

6.1 框架层没有 list/review 命令

nonoka-agent 框架本身没有内置的 list / review / inspect 命令。框架层只提供:

  • Runner.resume() —— 恢复入口;
  • Runner._create_session() —— 按 session_id 加载并运行;
  • Runner.cancel_session() —— 持久化取消状态;
  • MemoryBackend.get_history() —— 查询单 session 的 memory,但不是面向用户的审查入口。

这些能力是给 CLI / Server / 其他上层封装用的。

6.2 CLI 层的 sessions 子命令

nonoka-cli 在 src/nonoka_cli/commands/sessions_cmd.py:31-60 提供了两个命令:

async def execute():
  async with SessionManager(db_path=db_path) as manager:
    if args.sessions_command == "list":
      return await manager.list()
    item = await manager.get(args.session_id)
    if item is None:
      return None, [], None
    store = SQLiteEventStore(_event_db())
    try:
      return item, await store.list(args.session_id, 20), await store.summary(args.session_id)
    finally:
      await store.close()
  • sessions list:列出 cli_sessions 表,按 last_active DESC 排序;
  • sessions show <session_id>:查 cli_sessions 元数据 + events.db 里的时间线与 usage。

SessionManager 的说明写得很清楚(src/nonoka_cli/sessions/manager.py:44-54):

class SessionManager:
  """Manages the CLI session index table in SQLite.

  Uses a dedicated sqlite3 connection to the same database file that
  nonoka's ``SQLiteCheckpointStore`` uses. This lets the CLI maintain its
  own metadata while reusing nonoka's checkpoint persistence for message
  context.
  """

也就是说,CLI 自己维护一张 cli_sessions 索引表,但真正的消息上下文仍然复用 nonoka 的 checkpoint 持久化。

6.3 conversation_key 找回

SessionManager.find_by_conversation_key() 在 manager.py:223-243:

async def find_by_conversation_key(self, key: str) -> SessionInfo | None:
  """Return the most recently active session tagged with *key*, if any."""

  def _find() -> sqlite3.Row | None:
    conn = self._ensure_connection()
    return conn.execute(
      """
      SELECT * FROM cli_sessions
      WHERE json_extract(metadata, '$.conversation_key') = ?
      ORDER BY last_active DESC
      LIMIT 1
      """,
      (key,),
    ).fetchone()

数据库实际路径是项目级的 .nonoka/sessions.db,定义在 manager.py:29-37:

def project_session_db_path(working_dir: Path | str) -> Path:
  """Return the session/checkpoint database owned by one workspace.

  The OpenCode bridge is long-lived and several projects can be open at once.
  Keeping their state below each project avoids cross-project SQLite writer
  contention and prevents session history leaking into another project's TUI.
  """
  return Path(working_dir).expanduser().resolve() / ".nonoka" / "sessions.db"

7. 小结

nonoka 的会话管理可以概括为一句话:进程只负责计算,状态全部交给 SQLite;跨进程只传 session_id。

更具体地说:

  1. session_id 由 CLI 生成,通过 session_init 事件交给 Provider,Provider 把它存在 .nonoka/provider-session.id。Provider 用消息历史判断“新对话 vs 续聊”,CLI 用 new_session / session_id 判断;conversation_key 作为兜底指纹,防止 id 文件丢失后旧会话成为孤儿。
  2. 真正的状态对象是 SessionState,一个完整的 Pydantic 快照,包含计划、步骤、对话历史、运行时账本、扩展状态四大部分。
  3. 持久化采用“整份快照 + 增量补丁”:checkpoints 表存 state_json,step_updates 表存步骤级增量,加载时按时间顺序合并并遵守三条保护规则。
  4. 恢复时,Runner.resume() → load_session() → Session.from_state() → 根据是否有 current_plan 路由到 PlanExecutor 或 ReActAgent。ReAct 恢复会先重放悬空的 tool_calls,PlanExecutor 恢复会直接跳过已完成的 step。
  5. 审查历史是 CLI 层的能力,复用同一份 checkpoint 数据库,但额外维护一张 cli_sessions 索引表,通过 sessions list / sessions show 给人看。

理解了 SessionState 的结构,就能理解为什么 nonoka 敢说“进程死了也没关系”——只要 SQLite 文件还在,session_id 还在,整份执行上下文就能被完整地重建出来。