Repository navigation
fix: retain MCP catalogs per connection; prepare 0.4.5 - #128
Conversation
| // A catalog from startup does not guarantee tools on this turn. Stop | ||
| // before submitting a reduced toolset. Retain the client for disposal | ||
| // and a subsequent retry instead of silently deleting it (#127). | ||
| throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`) |
There was a problem hiding this comment.
🟠 MCP discovery throw hits all tools() callers, not just runs.
| throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`) | |
| Verify every caller of MCP.tools()/MCP.toolEntries() handles the new rejection: either catch at the run-loop call site and rethrow as session_error (scoping the abort to headless runs), or wrap the listing/catalog/interactive call sites (cli/cmd/mcp.ts, cli/cmd/tool-catalog.ts, session/prompt.ts) with error handling that surfaces the message instead of crashing. |
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:623-625):
Problem: MCP discovery throw hits all tools() callers, not just runs
Detail: discoverTools() now throws on any single server's listTools failure, so Promise.all rejects and the whole MCP.tools()/MCP.toolEntries() call fails for ALL servers, where it previously resolved with partial results (failed server skipped). The stated intent (#127) covers headless runs, and the new e2e test proves the `run` path converts this into session_error — but the module's importers also include non-run flows (cli/cmd/mcp.ts, cli/cmd/tool-catalog.ts, session/prompt.ts interactive loop, command/index.ts), and the usual co-change partners (src/cli/cmd/run.ts, src/cli/cmd/tool-catalog.ts, test/tool-catalog.test.ts) are NOT adapted in this PR. Any of those callers without a try/catch now crashes with an unhandled error where it previously degraded gracefully to partial output.
Suggested fix: Verify every caller of MCP.tools()/MCP.toolEntries() handles the new rejection: either catch at the run-loop call site and rethrow as session_error (scoping the abort to headless runs), or wrap the listing/catalog/interactive call sites (cli/cmd/mcp.ts, cli/cmd/tool-catalog.ts, session/prompt.ts) with error handling that surfaces the message instead of crashing.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
discoverTools() now throws on any single server's listTools failure, so Promise.all rejects and the whole MCP.tools()/MCP.toolEntries() call fails for ALL servers, where it previously resolved with partial results (failed server skipped). The stated intent (#127) covers headless runs, and the new e2e test proves the run path converts this into session_error — but the module's importers also include non-run flows (cli/cmd/mcp.ts, cli/cmd/tool-catalog.ts, session/prompt.ts interactive loop, command/index.ts), and the usual co-change partners (src/cli/cmd/run.ts, src/cli/cmd/tool-catalog.ts, test/tool-catalog.test.ts) are NOT adapted in this PR. Any of those callers without a try/catch now crashes with an unhandled error where it previously degraded gracefully to partial output.
error: error instanceof Error ? error.message : String(error),
})
// A catalog from startup does not guarantee tools on this turn. Stop
// before submitting a reduced toolset. Retain the client for disposal
// and a subsequent retry instead of silently deleting it (#127).
throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`)
})
}
export async function tools() {| "packages/util": { | ||
| "name": "@aictrl/util", | ||
| "version": "1.2.16", | ||
| "version": "1.2.17", |
There was a problem hiding this comment.
🟡 Lock bumps @aictrl/util to 1.2.17 with no manifest change.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, bun.lock:213):
Problem: Lock bumps @aictrl/util to 1.2.17 with no manifest change
Detail: bun.lock bumps the @aictrl/util workspace entry 1.2.16→1.2.17, but no packages/util/package.json change appears in this PR. If that manifest still reads 1.2.16 on this branch, the newly-added `bun install --frozen-lockfile` in publish.yml (and RELEASING.md step 1) fails on manifest/lock mismatch and blocks the 0.4.5 release outright. Most likely this is benign catch-up (the cli lock entry jumps 0.3.3→0.4.5, so the lock was stale and util's manifest was presumably already 1.2.17 from an earlier PR) — but confirm the manifest reads 1.2.17 before tagging, since the diff alone cannot prove it.
Suggested fix: Verify packages/util/package.json on this branch reads "version": "1.2.17". If it does not, either bump it in this PR or revert the bun.lock util hunk so lock and manifest agree under --frozen-lockfile.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
bun.lock bumps the @aictrl/util workspace entry 1.2.16→1.2.17, but no packages/util/package.json change appears in this PR. If that manifest still reads 1.2.16 on this branch, the newly-added bun install --frozen-lockfile in publish.yml (and RELEASING.md step 1) fails on manifest/lock mismatch and blocks the 0.4.5 release outright. Most likely this is benign catch-up (the cli lock entry jumps 0.3.3→0.4.5, so the lock was stale and util's manifest was presumably already 1.2.17 from an earlier PR) — but confirm the manifest reads 1.2.17 before tagging, since the diff alone cannot prove it.
},
"packages/util": {
"name": "@aictrl/util",
"version": "1.2.17",
"dependencies": {
"zod": "catalog:",
},
| }) | ||
| const mcpConfig = config[clientName] | ||
| const timeout = | ||
| (isMcpConfigured(mcpConfig) ? mcpConfig.timeout : undefined) ?? defaultTimeout ?? DEFAULT_TIMEOUT |
There was a problem hiding this comment.
🟡 Discovery and tool-call timeout fallbacks diverge.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:643-651):
Problem: Discovery and tool-call timeout fallbacks diverge
Detail: The same conceptual value is now computed three divergent ways: tools() pre-loop uses (per-server timeout) ?? defaultTimeout ?? DEFAULT_TIMEOUT; the tools() result loop re-derives entry?.timeout ?? defaultTimeout (no DEFAULT_TIMEOUT fallback) for the converted tools; toolEntries() inlines a third variant with cfg.experimental?.mcp_timeout. When neither a per-server timeout nor defaultTimeout is set, discovery is bounded by DEFAULT_TIMEOUT but each tool call gets timeout=undefined (SDK default / unbounded), so a hung server still stalls tools/call. The duplication also invites future drift between the three sites.
Suggested fix: Extract one helper, e.g. mcpTimeout(cfg, clientName) { const c = cfg.mcp?.[clientName]; return (isMcpConfigured(c) ? c.timeout : undefined) ?? cfg.experimental?.mcp_timeout ?? DEFAULT_TIMEOUT }, and use it at all three sites — including the per-tool timeout passed to convertMcpTool — so discovery and tool calls share the same bound.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The same conceptual value is now computed three divergent ways: tools() pre-loop uses (per-server timeout) ?? defaultTimeout ?? DEFAULT_TIMEOUT; the tools() result loop re-derives entry?.timeout ?? defaultTimeout (no DEFAULT_TIMEOUT fallback) for the converted tools; toolEntries() inlines a third variant with cfg.experimental?.mcp_timeout. When neither a per-server timeout nor defaultTimeout is set, discovery is bounded by DEFAULT_TIMEOUT but each tool call gets timeout=undefined (SDK default / unbounded), so a hung server still stalls tools/call. The duplication also invites future drift between the three sites.
const mcpConfig = config[clientName]
const timeout =
(isMcpConfigured(mcpConfig) ? mcpConfig.timeout : undefined) ?? defaultTimeout ?? DEFAULT_TIMEOUT
const toolsResult = await discoverTools(clientName, client, timeout)
return { clientName, client, toolsResult }
}),
)
for (const { clientName, client, toolsResult } of toolsResults) {
const mcpConfig = config[clientName]
const entry = isMcpConfigured(mcpConfig) ? mcpConfig : undefined
const timeout = entry?.timeout ?? defaultTimeout| const mcpConfig = config[clientName] | ||
| const timeout = | ||
| (isMcpConfigured(mcpConfig) ? mcpConfig.timeout : undefined) ?? defaultTimeout ?? DEFAULT_TIMEOUT | ||
| const toolsResult = await discoverTools(clientName, client, timeout) |
There was a problem hiding this comment.
🟡 Failed discovery leaves server status "connected" in state.
| const toolsResult = await discoverTools(clientName, client, timeout) | |
| Before throwing in discoverTools (or at the tools()/toolEntries() call sites), still write s.status[clientName] = { status: "failed", error: message } while keeping the client in s.clients for disposal and retry. |
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:644-645):
Problem: Failed discovery leaves server status "connected" in state
Detail: The removed catch block set s.status[clientName] = { status: "failed", error } and deleted the client from s.clients; the replacement only logs and throws. Retaining the client for disposal/retry is deliberate (#127 comment), but the persisted status is never updated, so after a discovery failure the state still reports the broken server as connected/healthy and records no failure at all — any consumer of s.status (e.g. `aictrl mcp` status display) is misled.
Suggested fix: Before throwing in discoverTools (or at the tools()/toolEntries() call sites), still write s.status[clientName] = { status: "failed", error: message } while keeping the client in s.clients for disposal and retry.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The removed catch block set s.status[clientName] = { status: "failed", error } and deleted the client from s.clients; the replacement only logs and throws. Retaining the client for disposal/retry is deliberate (#127 comment), but the persisted status is never updated, so after a discovery failure the state still reports the broken server as connected/healthy and records no failure at all — any consumer of s.status (e.g. aictrl mcp status display) is misled.
const toolsResults = await Promise.all(
connectedClients.map(async ([clientName, client]) => {
const mcpConfig = config[clientName]
const timeout =
(isMcpConfigured(mcpConfig) ? mcpConfig.timeout : undefined) ?? defaultTimeout ?? DEFAULT_TIMEOUT
const toolsResult = await discoverTools(clientName, client, timeout)
return { clientName, client, toolsResult }
}),
)
for (const { clientName, client, toolsResult } of toolsResults) {| for attempt in {1..10}; do | ||
| if npm install "@aictrl/cli@${AICTRL_VERSION}" 2>&1 | tee install.log; then | ||
| echo "::notice::npm install @aictrl/cli@${AICTRL_VERSION} succeeded on attempt ${attempt}" | ||
| INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version) |
There was a problem hiding this comment.
⚪ Missing CLI binary aborts smoke step without diagnostics.
| INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version) | |
| INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version || echo "binary-missing") so the existing != comparison emits the proper ::error message instead of a bare exit code. |
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, .github/workflows/publish.yml:193-194):
Problem: Missing CLI binary aborts smoke step without diagnostics
Detail: INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version) runs under `set -euo pipefail` inside the success branch. If npm install exits 0 but the bin shim is missing or non-executable, the command substitution's non-zero status terminates the step with a bare exit code — no ::error annotation and no retry, unlike every other failure path in the loop.
Suggested fix: INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version || echo "binary-missing") so the existing != comparison emits the proper ::error message instead of a bare exit code.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version) runs under set -euo pipefail inside the success branch. If npm install exits 0 but the bin shim is missing or non-executable, the command substitution's non-zero status terminates the step with a bare exit code — no ::error annotation and no retry, unlike every other failure path in the loop.
# v0.4.4 took longer than 2.5 minutes to become installable (#126).
# Ten attempts allow 8.5 minutes of propagation backoff, bounded by
# the step timeout even if npm itself stalls.
for attempt in {1..10}; do
if npm install "@aictrl/cli@${AICTRL_VERSION}" 2>&1 | tee install.log; then
echo "::notice::npm install @aictrl/cli@${AICTRL_VERSION} succeeded on attempt ${attempt}"
INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version)
if [ "$INSTALLED_VERSION" != "$AICTRL_VERSION" ]; then
echo "::error::Installed CLI reports $INSTALLED_VERSION, expected $AICTRL_VERSION"
exit 1
fi| for attempt in {1..10}; do | ||
| if npm install "@aictrl/cli@${AICTRL_VERSION}" 2>&1 | tee install.log; then | ||
| echo "::notice::npm install @aictrl/cli@${AICTRL_VERSION} succeeded on attempt ${attempt}" | ||
| INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version) |
There was a problem hiding this comment.
⚪ Smoke step executes registry-fetched CLI in publish job.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, .github/workflows/publish.yml:193-194):
Problem: Smoke step executes registry-fetched CLI in publish job
Detail: The new version check executes code just fetched from the public npm registry inside the release job, which holds repo and npm publish credentials. The incremental risk over the pre-existing install is small (npm install already runs dependency lifecycle scripts, and the artifact was built from this same tagged commit), but the explicit execution sink is new — worth hardening given the job's privileges.
Suggested fix: Use `npm install --ignore-scripts` for the install attempts so only the explicit `--version` invocation runs foreign code, and/or run the smoke check in a separate job with a minimal permissions block.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The new version check executes code just fetched from the public npm registry inside the release job, which holds repo and npm publish credentials. The incremental risk over the pre-existing install is small (npm install already runs dependency lifecycle scripts, and the artifact was built from this same tagged commit), but the explicit execution sink is new — worth hardening given the job's privileges.
for attempt in {1..10}; do
if npm install "@aictrl/cli@${AICTRL_VERSION}" 2>&1 | tee install.log; then
echo "::notice::npm install @aictrl/cli@${AICTRL_VERSION} succeeded on attempt ${attempt}"
INSTALLED_VERSION=$(./node_modules/.bin/aictrl --version)
if [ "$INSTALLED_VERSION" != "$AICTRL_VERSION" ]; then
echo "::error::Installed CLI reports $INSTALLED_VERSION, expected $AICTRL_VERSION"
exit 1
fi
exit 0
fi| // A catalog from startup does not guarantee tools on this turn. Stop | ||
| // before submitting a reduced toolset. Retain the client for disposal | ||
| // and a subsequent retry instead of silently deleting it (#127). | ||
| throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`) |
There was a problem hiding this comment.
⚪ discoverTools drops the original error cause.
@@ -623 +623,3 @@
- throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`)
+ throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`, {
+ cause: error,
+ })
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:623):
Problem: discoverTools drops the original error cause
Detail: The rethrown Error carries only the generic message; the original listTools rejection (stack, error code) is lost to callers — only the log.error keeps it. Attaching { cause: error } preserves root-cause diagnostics for the session_error/event payloads shown to users.
Suggested fix: Pass the original error as cause: throw new Error(msg, { cause: error }).
Suggested patch:
@@ -623 +623,3 @@
- throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`)
+ throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`, {
+ cause: error,
+ })
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The rethrown Error carries only the generic message; the original listTools rejection (stack, error code) is lost to callers — only the log.error keeps it. Attaching { cause: error } preserves root-cause diagnostics for the session_error/event payloads shown to users.
error: error instanceof Error ? error.message : String(error),
})
// A catalog from startup does not guarantee tools on this turn. Stop
// before submitting a reduced toolset. Retain the client for disposal
// and a subsequent retry instead of silently deleting it (#127).
throw new Error(`MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`)
})
}
Code reviewVerdict: Address the major findings before merging. · 🔴 0 · 🟠 1 · 🟡 3 · ⚪ 3 · 0/7 resolved
🤖 Fix all 7 open findings with your agent📋 Out-of-diff findings (7)
Reviewed 8 files · 0 inline · view all 7 findings ↗ aictrl · AI code review for fast-moving teams · aictrl.dev |
Review response — PR #128Verified the initial review against its reviewed revision, applied the confirmed fixes, and reran the relevant regressions. Issues addressed (pushed to this PR)
Review claims verified false (no change needed)
Not addressed hereNone. GitHub CI and the next automated review must still pass on the updated head before promotion. Verification
The recorded findings currently have null GitHub comment IDs. Exact records were resolved from the returned review ID, full claim title, file, starting line and matching claim text; every match was unique. |
| // A catalog from startup does not guarantee tools on this turn. Stop | ||
| // before submitting a reduced toolset. Retain the client for disposal | ||
| // and a subsequent retry instead of silently deleting it (#127). | ||
| throw new Error(message, { cause: error }) |
There was a problem hiding this comment.
🟠 Discovery throw aborts unvetted tools() callers.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:646):
Problem: Discovery throw aborts unvetted tools() callers
Detail: discoverTools now throws on every listTools failure, so MCP.tools()/toolEntries() reject instead of degrading to a reduced toolset. This PR updates and tests only the headless run path; the MCP module's other importers — session/prompt.ts (interactive turn loop), cli/cmd/mcp.ts, cli/cmd/tool-catalog.ts, command/index.ts, cli/error.ts — are untouched, so one flaky MCP server can abort an interactive session turn or crash the `aictrl mcp` / `tool-catalog` commands with an unhandled rejection where they previously continued. Repro: Given an interactive session with one healthy and one failing (500 on tools/list) MCP server, When session/prompt.ts builds the per-turn toolset via MCP.tools(), Then the failing server's rejection propagates out of Promise.all and aborts the turn with "MCP tool discovery failed for server ..." instead of proceeding with the healthy server's tools.
Suggested fix: Scope the hard-throw to the headless run path (e.g. an abortOnError option on tools()/toolEntries() set by command/index.ts), or catch the discovery error in session/prompt.ts and the mcp/tool-catalog commands so interactive callers surface the failure as a turn error / partial listing instead of an unhandled rejection.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
discoverTools now throws on every listTools failure, so MCP.tools()/toolEntries() reject instead of degrading to a reduced toolset. This PR updates and tests only the headless run path; the MCP module's other importers — session/prompt.ts (interactive turn loop), cli/cmd/mcp.ts, cli/cmd/tool-catalog.ts, command/index.ts, cli/error.ts — are untouched, so one flaky MCP server can abort an interactive session turn or crash the aictrl mcp / tool-catalog commands with an unhandled rejection where they previously continued. Repro: Given an interactive session with one healthy and one failing (500 on tools/list) MCP server, When session/prompt.ts builds the per-turn toolset via MCP.tools(), Then the failing server's rejection propagates out of Promise.all and aborts the turn with "MCP tool discovery failed for server ..." instead of proceeding with the healthy server's tools.
log.error("MCP tool discovery failed", {
clientName,
error: error instanceof Error ? error.message : String(error),
})
// A catalog from startup does not guarantee tools on this turn. Stop
// before submitting a reduced toolset. Retain the client for disposal
// and a subsequent retry instead of silently deleting it (#127).
throw new Error(message, { cause: error })
})
}
export async function tools() {| # Registry install and binary execution receive no publication permissions. | ||
| permissions: {} | ||
| steps: | ||
| - uses: actions/setup-node@v6 |
There was a problem hiding this comment.
🟡 setup-node pinned to mutable v6 tag, not SHA.
| - uses: actions/setup-node@v6 | |
| Pin to an immutable ref: `- uses: actions/setup-node@<full-40-char-SHA> # v6.x.y`. |
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, .github/workflows/publish.yml:178-179):
Problem: setup-node pinned to mutable v6 tag, not SHA
Detail: The new smoke job references actions/setup-node by mutable major tag instead of a full commit SHA. If the tag is repointed (tag hijack or maintainer compromise), arbitrary code runs inside the release workflow. Blast radius is reduced by the job's `permissions: {}` and absence of checkout/publish tokens, but this workflow performs npm publications, so supply-chain hardening here is cheap and high-value.
Suggested fix: Pin to an immutable ref: `- uses: actions/setup-node@<full-40-char-SHA> # v6.x.y`.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The new smoke job references actions/setup-node by mutable major tag instead of a full commit SHA. If the tag is repointed (tag hijack or maintainer compromise), arbitrary code runs inside the release workflow. Blast radius is reduced by the job's permissions: {} and absence of checkout/publish tokens, but this workflow performs npm publications, so supply-chain hardening here is cheap and high-value.
permissions: {}
steps:
- uses: actions/setup-node@v6
with:
node-version: 22
- name: Smoke test - verify @aictrl/cli installs cleanly via npm
# Catches the v0.3.3-class bug where the published manifest carries| // A catalog from startup does not guarantee tools on this turn. Stop | ||
| // before submitting a reduced toolset. Retain the client for disposal | ||
| // and a subsequent retry instead of silently deleting it (#127). | ||
| throw new Error(message, { cause: error }) |
There was a problem hiding this comment.
🟡 Stale client's discovery failure still aborts run.
| throw new Error(message, { cause: error }) | |
| Only rethrow when `s.clients[clientName] === client`; for a stale client, log and resolve with an empty result (mirroring the stale success path, which silently returns) so a superseded client's in-flight failure cannot fail the catalog build. |
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:646):
Problem: Stale client's discovery failure still aborts run
Detail: The catch block gates all state writes (status, discoveryFailures) behind the staleness check `s.clients[clientName] === client`, but the rethrow is unconditional. If add()/connect()/finishAuth() replaces the client (or disconnect() removes it) while its listTools is in flight, the superseded client's transport-closed rejection still rejects tools()/toolEntries() and aborts the run — precisely the outcome the staleness guard exists to prevent. Repro: Given a per-turn tools() discovery in flight for server X, When finishAuth(X) replaces X's client in the same process, Then the old client's listTools rejects, the guard skips the state writes, but the unconditional throw still aborts the turn for a server that now has a healthy replacement.
Suggested fix: Only rethrow when `s.clients[clientName] === client`; for a stale client, log and resolve with an empty result (mirroring the stale success path, which silently returns) so a superseded client's in-flight failure cannot fail the catalog build.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The catch block gates all state writes (status, discoveryFailures) behind the staleness check s.clients[clientName] === client, but the rethrow is unconditional. If add()/connect()/finishAuth() replaces the client (or disconnect() removes it) while its listTools is in flight, the superseded client's transport-closed rejection still rejects tools()/toolEntries() and aborts the run — precisely the outcome the staleness guard exists to prevent. Repro: Given a per-turn tools() discovery in flight for server X, When finishAuth(X) replaces X's client in the same process, Then the old client's listTools rejects, the guard skips the state writes, but the unconditional throw still aborts the turn for a server that now has a healthy replacement.
.catch((error) => {
const message = `MCP tool discovery failed for server "${clientName}". Check the MCP server and retry.`
if (s.clients[clientName] === client) {
s.status[clientName] = { status: "failed", error: message }
s.discoveryFailures.add(clientName)
}
log.error("MCP tool discovery failed", {
clientName,
error: error instanceof Error ? error.message : String(error),
})
// A catalog from startup does not guarantee tools on this turn. Stop
// before submitting a reduced toolset. Retain the client for disposal
// and a subsequent retry instead of silently deleting it (#127).
throw new Error(message, { cause: error })| const timeout = configuredTimeout(cfg, clientName) | ||
| // Listing uses the connection/discovery default. Calls retain the | ||
| // SDK's existing default when no explicit timeout is configured. | ||
| const toolsResult = await discoverTools(clientName, client, timeout ?? DEFAULT_TIMEOUT) |
There was a problem hiding this comment.
🟡 30s discovery floor converts slow servers to aborts.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:665):
Problem: 30s discovery floor converts slow servers to aborts
Detail: tools() and toolEntries() now bound every per-turn listTools with `timeout ?? DEFAULT_TIMEOUT` (module constant, 30s). Previously the per-turn listTools() call passed no timeout and used the MCP SDK request default (60s), so servers enumerating large toolsets in 30-60s succeeded; now they exceed the 30s cap, and because a discovery failure throws, the entire invocation aborts rather than merely being slow. A slow-but-healthy server becomes a hard run failure. Repro: Given a remote MCP server whose tools/list takes ~40s with no explicit timeout configured, When a headless run builds its per-turn catalog, Then listTools times out at 30s and the run aborts, where before this PR the same server enumerated successfully.
Suggested fix: Use a discovery-specific ceiling for listTools (e.g. keep the SDK's 60s request default, or max(DEFAULT_TIMEOUT, SDK default)) so slow-but-healthy servers are not converted into run-aborting failures; only genuine errors should trigger the #127 abort.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
tools() and toolEntries() now bound every per-turn listTools with timeout ?? DEFAULT_TIMEOUT (module constant, 30s). Previously the per-turn listTools() call passed no timeout and used the MCP SDK request default (60s), so servers enumerating large toolsets in 30-60s succeeded; now they exceed the 30s cap, and because a discovery failure throws, the entire invocation aborts rather than merely being slow. A slow-but-healthy server becomes a hard run failure. Repro: Given a remote MCP server whose tools/list takes ~40s with no explicit timeout configured, When a headless run builds its per-turn catalog, Then listTools times out at 30s and the run aborts, where before this PR the same server enumerated successfully.
const toolsResults = await Promise.all(
connectedClients.map(async ([clientName, client]) => {
const timeout = configuredTimeout(cfg, clientName)
// Listing uses the connection/discovery default. Calls retain the
// SDK's existing default when no explicit timeout is configured.
const toolsResult = await discoverTools(clientName, client, timeout ?? DEFAULT_TIMEOUT)
return { clientName, client, toolsResult, timeout }
}),
)| 2. Run the headless release regressions from `packages/cli`: | ||
|
|
||
| ```bash | ||
| bun test test/cli/run-mcp-discovery.test.ts test/cli/run-provider-finish.test.ts test/cli/run-signal-cancellation.test.ts test/cli/classify-session-error.test.ts test/session/idle.test.ts test/session/processor-idle.test.ts |
There was a problem hiding this comment.
🟡 Release regressions omit discovery-recovery test.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, RELEASING.md:9-10):
Problem: Release regressions omit discovery-recovery test
Detail: Step 2's headless regression list includes test/cli/run-mcp-discovery.test.ts but not packages/cli/test/mcp/discovery-recovery.test.ts — the direct regression this same PR adds for the exact #127 retry/retention semantics being shipped in 0.4.5 (failed status, client retained, retry succeeds). Following the doc verbatim skips the test that most tightly covers the fix.
Suggested fix: Add test/mcp/discovery-recovery.test.ts to the `bun test` command in RELEASING.md step 2.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
Step 2's headless regression list includes test/cli/run-mcp-discovery.test.ts but not packages/cli/test/mcp/discovery-recovery.test.ts — the direct regression this same PR adds for the exact #127 retry/retention semantics being shipped in 0.4.5 (failed status, client retained, retry succeeds). Following the doc verbatim skips the test that most tightly covers the fix.
`packages/cli/package.json`, and the matching `packages/cli` workspace
version in `bun.lock`. Run `bun install --frozen-lockfile`.
2. Run the headless release regressions from `packages/cli`:
```bash
bun test test/cli/run-mcp-discovery.test.ts test/cli/run-provider-finish.test.ts test/cli/run-signal-cancellation.test.ts test/cli/classify-session-error.test.ts test/session/idle.test.ts test/session/processor-idle.test.ts
```
Wait for CI build, workspace typecheck and tests to pass and address reviews.
| return undefined | ||
| }) | ||
| return { clientName, client, toolsResult } | ||
| const timeout = configuredTimeout(cfg, clientName) |
There was a problem hiding this comment.
⚪ timeout differs semantically in tools() vs toolEntries().
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:662-666):
Problem: timeout differs semantically in tools() vs toolEntries()
Detail: In tools(), `timeout` is the configured per-server/experimental value (may be undefined) and is threaded through to convertMcpTool for tool calls, with DEFAULT_TIMEOUT applied only at the discoverTools call site. In the adjacent toolEntries(), the same identifier already has DEFAULT_TIMEOUT folded in and is then discarded. Two subtly different meanings for one name in sibling functions sharing the new helper invites passing the wrong one to a future caller.
Suggested fix: Name them distinctly, e.g. `callTimeout` in tools() (threaded to convertMcpTool) and `discoveryTimeout` in toolEntries().
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
In tools(), timeout is the configured per-server/experimental value (may be undefined) and is threaded through to convertMcpTool for tool calls, with DEFAULT_TIMEOUT applied only at the discoverTools call site. In the adjacent toolEntries(), the same identifier already has DEFAULT_TIMEOUT folded in and is then discarded. Two subtly different meanings for one name in sibling functions sharing the new helper invites passing the wrong one to a future caller.
connectedClients.map(async ([clientName, client]) => {
const timeout = configuredTimeout(cfg, clientName)
// Listing uses the connection/discovery default. Calls retain the
// SDK's existing default when no explicit timeout is configured.
const toolsResult = await discoverTools(clientName, client, timeout ?? DEFAULT_TIMEOUT)
return { clientName, client, toolsResult, timeout }
}),
Code reviewVerdict: Address the major findings before merging. · 🔴 0 · 🟠 1 · 🟡 4 · ⚪ 1 · 0/6 resolved
🤖 Fix all 6 open findings with your agent📋 Out-of-diff findings (6)
Reviewed 9 files · 0 inline · view all 6 findings ↗ aictrl · AI code review for fast-moving teams · aictrl.dev |
Review response — PR #128Verified the follow-up findings recorded by the automated review against their reviewed revision. These exact recorded findings are available in the data layer before the pending GitHub review delivery; IDs are taken directly from those structured records, with no location-based guesses. Issues addressed (pushed to this PR)
Review claims verified false (no change needed)
Not addressed hereInformational or intentional behavior retained:
VerificationA real SDK response delayed 31 seconds failed against the reviewed head at 30.23 seconds, then passed concurrently through both runtime tools and catalog with the corrected SDK default. Recovery: 2 passed; focused publication/headless cases: 15 passed; compiled 0.4.5 cases: 5 passed. Native build, typecheck and normal push hooks passed. GitHub must validate the new head. |
| const catalog = toolCatalogs.get(client) | ||
| if (!catalog) throw new Error(`Missing MCP tool catalog for server "${serverName}"`) | ||
| await catalog.refresh | ||
| if (toolCatalogs.get(client) !== catalog) |
There was a problem hiding this comment.
🟡 Server swap mid-turn kills the in-flight turn.
| if (toolCatalogs.get(client) !== catalog) | |
| When the post-await identity check fails, re-resolve the client by name from state (s.clients[serverName]) once and use the fresh client's catalog instead of throwing; or have tools()/toolEntries() tolerate replaced catalogs for the current call, since the next turn picks up the new client. |
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:166-167):
Problem: Server swap mid-turn kills the in-flight turn
Detail: catalogTools throws "MCP connection changed ... Retry the invocation" when add()/disconnect() replaces a client while tools() is awaiting that server's catalog.refresh. Because tools()/toolEntries() Promise.all every server, replacing or reconnecting ANY server during an in-flight turn (e.g. finishAuth() re-adds the server after an OAuth callback, or a concurrent MCP.add()/disconnect() in-process) rejects the entire turn with a session error, where the previous implementation still returned a usable toolset for the duration of the call. The error is actionable and covered by tests, but a routine reconnection now aborts the active model turn.
Suggested fix: When the post-await identity check fails, re-resolve the client by name from state (s.clients[serverName]) once and use the fresh client's catalog instead of throwing; or have tools()/toolEntries() tolerate replaced catalogs for the current call, since the next turn picks up the new client.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
catalogTools throws "MCP connection changed ... Retry the invocation" when add()/disconnect() replaces a client while tools() is awaiting that server's catalog.refresh. Because tools()/toolEntries() Promise.all every server, replacing or reconnecting ANY server during an in-flight turn (e.g. finishAuth() re-adds the server after an OAuth callback, or a concurrent MCP.add()/disconnect() in-process) rejects the entire turn with a session error, where the previous implementation still returned a usable toolset for the duration of the call. The error is actionable and covered by tests, but a routine reconnection now aborts the active model turn.
async function catalogTools(client: MCPClient, serverName: string) {
const catalog = toolCatalogs.get(client)
if (!catalog) throw new Error(`Missing MCP tool catalog for server "${serverName}"`)
await catalog.refresh
if (toolCatalogs.get(client) !== catalog)
throw new Error(`MCP connection changed for server "${serverName}". Retry the invocation.`)
if (catalog.error) throw catalog.error
return catalog.tools
}|
|
||
| log.info("create() successfully created client", { key, toolCount: result.tools.length }) | ||
| toolCatalogs.set(mcpClient, { tools: result.tools, refresh: Promise.resolve() }) | ||
| registerNotificationHandlers(mcpClient, key, mcp.timeout ?? cfg.experimental?.mcp_timeout) |
There was a problem hiding this comment.
🟡 Discovery and refresh resolve timeouts differently.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:576):
Problem: Discovery and refresh resolve timeouts differently
Detail: The same tools/list operation now runs under three different timeout resolutions. create()'s initial discovery uses mcp.timeout ?? DEFAULT_TIMEOUT (30s, line 551); notification-driven refreshes use mcp.timeout ?? cfg.experimental?.mcp_timeout captured once at create time (line 576) — undefined there falls back to the MCP SDK's own 60s request default; and per-tool execution re-derives configuredTimeout() on every tools() call. With only experimental.mcp_timeout configured, startup discovery times out at 30s while refreshes use the experimental value; with neither configured, a hung refresh can wait 60s where startup waited 30s. The refresh timeout is also frozen at create time, so later config changes never affect it, while tool-call timeouts track current config — the same fact (a server's timeout) derived in multiple places with different lifetimes.
Suggested fix: Compute one resolver, e.g. `const discoveryTimeout = mcp.timeout ?? cfg.experimental?.mcp_timeout ?? DEFAULT_TIMEOUT`, and use it for both the initial withTimeout(mcpClient.listTools(), ...) call and registerNotificationHandlers (or store it in the ToolCatalog entry) so discovery, refresh and config updates share identical, current timeout semantics.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The same tools/list operation now runs under three different timeout resolutions. create()'s initial discovery uses mcp.timeout ?? DEFAULT_TIMEOUT (30s, line 551); notification-driven refreshes use mcp.timeout ?? cfg.experimental?.mcp_timeout captured once at create time (line 576) — undefined there falls back to the MCP SDK's own 60s request default; and per-tool execution re-derives configuredTimeout() on every tools() call. With only experimental.mcp_timeout configured, startup discovery times out at 30s while refreshes use the experimental value; with neither configured, a hung refresh can wait 60s where startup waited 30s. The refresh timeout is also frozen at create time, so later config changes never affect it, while tool-call timeouts track current config — the same fact (a server's timeout) derived in multiple places with different lifetimes.
log.info("create() successfully created client", { key, toolCount: result.tools.length })
toolCatalogs.set(mcpClient, { tools: result.tools, refresh: Promise.resolve() })
registerNotificationHandlers(mcpClient, key, mcp.timeout ?? cfg.experimental?.mcp_timeout)
return {
mcpClient,| }) | ||
| return { clientName, client, toolsResult } | ||
| }), | ||
| const catalogs = await Promise.all( |
There was a problem hiding this comment.
🟡 One failed MCP server fails tools() for all servers.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:654-658):
Problem: One failed MCP server fails tools() for all servers
Detail: tools() and toolEntries() (lines 681-686) Promise.all over per-server catalogTools, so a single server's cached catalog.error (failed refresh or closed connection) rejects the entire call: every model turn errors and healthy servers' tools are withheld until the failed server happens to send another tools/list_changed notification or the operator reconnects it. With multiple failed servers, Promise.all surfaces an arbitrary first rejection, hiding the other culprits from the error message. Failing the turn is the documented intent, but the blast radius spans unrelated servers and the reported error names only one of possibly several. Repro: with servers A (healthy) and B (dropped connection), the next model turn rejects with only B's error and A's tools are never dispatched, on every subsequent turn.
Suggested fix: Use Promise.allSettled over the per-server catalogTools calls and throw an aggregated error listing every failed server name and its message, so operators see all servers needing attention instead of an arbitrary first rejection.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
tools() and toolEntries() (lines 681-686) Promise.all over per-server catalogTools, so a single server's cached catalog.error (failed refresh or closed connection) rejects the entire call: every model turn errors and healthy servers' tools are withheld until the failed server happens to send another tools/list_changed notification or the operator reconnects it. With multiple failed servers, Promise.all surfaces an arbitrary first rejection, hiding the other culprits from the error message. Failing the turn is the documented intent, but the blast radius spans unrelated servers and the reported error names only one of possibly several. Repro: with servers A (healthy) and B (dropped connection), the next model turn rejects with only B's error and A's tools are never dispatched, on every subsequent turn.
const s = await state()
const cfg = await Config.get()
const catalogs = await Promise.all(
Object.entries(s.clients).map(async ([clientName, client]) => ({
clientName,
client,
tools: await catalogTools(client, clientName),
})),
)
for (const { clientName, client, tools } of catalogs) {| result[key] = s.status[key] ?? { status: "disabled" } | ||
| const client = s.clients[key] | ||
| const error = client && toolCatalogs.get(client)?.error | ||
| result[key] = error ? { status: "failed", error: error.message } : (s.status[key] ?? { status: "disabled" }) |
There was a problem hiding this comment.
🟡 Server health now has two diverging sources of truth.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:594):
Problem: Server health now has two diverging sources of truth
Detail: This PR makes catalog.error the authoritative failure signal for tools: status() synthesizes {status:"failed"} from it (line 594) while s.status[key] still says "connected", and prompts() (line 700) / resources() (line 721) still gate solely on s.status === "connected". After a transport close (client.onclose at line 130 sets only catalog.error) or a failed notification refresh, status() reports the server "failed" while prompts()/resources() keep invoking listPrompts/listResources on the dead client on every call; their per-call .catch swallows the error, so the server's prompts/resources silently vanish. Two hand-synced views of server health now coexist in one module and disagree in exactly the states this PR adds; they are kept in sync only by hand at the create/add/disconnect sites.
Suggested fix: Derive from one source: either have the onclose handler / refreshCatalog also write s.status[serverName] = { status: "failed", error } (and restore "connected" on successful refresh) so prompts()/resources() observe the failure through the existing gate and status() needs no catalog lookup; or gate prompts()/resources() on the same catalog error via a shared health helper (e.g. serverHealth(client, name) reading toolCatalogs.get(client)?.error).
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
This PR makes catalog.error the authoritative failure signal for tools: status() synthesizes {status:"failed"} from it (line 594) while s.status[key] still says "connected", and prompts() (line 700) / resources() (line 721) still gate solely on s.status === "connected". After a transport close (client.onclose at line 130 sets only catalog.error) or a failed notification refresh, status() reports the server "failed" while prompts()/resources() keep invoking listPrompts/listResources on the dead client on every call; their per-call .catch swallows the error, so the server's prompts/resources silently vanish. Two hand-synced views of server health now coexist in one module and disagree in exactly the states this PR adds; they are kept in sync only by hand at the create/add/disconnect sites.
for (const [key, mcp] of Object.entries(config)) {
if (!isMcpConfigured(mcp)) continue
const client = s.clients[key]
const error = client && toolCatalogs.get(client)?.error
result[key] = error ? { status: "failed", error: error.message } : (s.status[key] ?? { status: "disabled" })
}| RELEASE_TAG: ${{ github.event.release.tag_name }} | ||
| run: | | ||
| set -euo pipefail | ||
| AICTRL_VERSION="${RELEASE_TAG#v}" |
There was a problem hiding this comment.
⚪ Validate semver tag before npm spec and GITHUB_ENV.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, .github/workflows/publish.yml:193-206):
Problem: Validate semver tag before npm spec and GITHUB_ENV
Detail: The smoke job derives AICTRL_VERSION from github.event.release.tag_name with only the "v" prefix stripped and an empty-check, then interpolates it into an npm package spec (`npm install "@aictrl/cli@${AICTRL_VERSION}"`, line 204) and executes the installed `aictrl` bin (line 206); the publish job appends the same derived string to $GITHUB_ENV (line 63). Full quoting blocks shell injection, but npm accepts alias/URL spec syntax (npm:<pkg>, git/https URLs), so a crafted tag could redirect the install to an arbitrary registry package whose binary the step then runs. Today the publish job's strict PKG_VERSION equality gate plus `needs: publish` ordering means a non-matching tag never reaches the smoke job, so this is defense-in-depth rather than an exploitable path.
Suggested fix: After stripping the "v" prefix, validate the format before use in both jobs, e.g. `if ! printf '%s' "$AICTRL_VERSION" | grep -Eq '^[0-9]+\.[0-9]+\.[0-9]+(-[0-9A-Za-z.-]+)?$'; then echo "::error::Non-semver release tag: $RELEASE_TAG"; exit 1; fi`. This closes the npm alias/URL spec sink and the $GITHUB_ENV write independently of the package.json comparison.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The smoke job derives AICTRL_VERSION from github.event.release.tag_name with only the "v" prefix stripped and an empty-check, then interpolates it into an npm package spec (npm install "@aictrl/cli@${AICTRL_VERSION}", line 204) and executes the installed aictrl bin (line 206); the publish job appends the same derived string to $GITHUB_ENV (line 63). Full quoting blocks shell injection, but npm accepts alias/URL spec syntax (npm:, git/https URLs), so a crafted tag could redirect the install to an arbitrary registry package whose binary the step then runs. Today the publish job's strict PKG_VERSION equality gate plus needs: publish ordering means a non-matching tag never reaches the smoke job, so this is defense-in-depth rather than an exploitable path.
RELEASE_TAG: ${{ github.event.release.tag_name }}
run: |
set -euo pipefail
AICTRL_VERSION="${RELEASE_TAG#v}"
smoke="${RUNNER_TEMP}/aictrl-publish-smoke"
rm -rf "$smoke"
mkdir -p "$smoke"
cd "$smoke"
npm init -y > /dev/null| } catch (error) { | ||
| if (toolCatalogs.get(client) !== catalog) return | ||
| catalog.error = new Error( | ||
| `MCP tool discovery failed for server "${serverName}". Check the MCP server and retry.`, |
There was a problem hiding this comment.
⚪ "Check the server and retry" advice never rediscovers.
--- a/packages/cli/src/mcp/index.ts
+++ b/packages/cli/src/mcp/index.ts
@@ -147,7 +147,7 @@
} catch (error) {
if (toolCatalogs.get(client) !== catalog) return
catalog.error = new Error(
- `MCP tool discovery failed for server "${serverName}". Check the MCP server and retry.`,
+ `MCP tool discovery failed for server "${serverName}". Reconnect the MCP server and retry.`,
{ cause: error },
)🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:150):
Problem: "Check the server and retry" advice never rediscovers
Detail: The refresh-failure message tells the user to "Check the MCP server and retry", but retrying does not issue a new tools/list: catalogTools re-throws the cached catalog.error on every call, and the PR's own test asserts the server's tools/list count stays flat after a failed read ("failed catalog reads must not silently retry"). Recovery requires a new tools/list_changed notification from the server or an explicit reconnect — exactly what the sibling onclose message says ("Reconnect the MCP server and retry"). The instruction in this message does not match the implemented recovery path.
Suggested fix: Align with the onclose wording: `MCP tool discovery failed for server "${serverName}". Reconnect the MCP server and retry.` — the existing tests match on the unchanged "MCP tool discovery failed for server" prefix, so they keep passing.
Suggested patch:
--- a/packages/cli/src/mcp/index.ts
+++ b/packages/cli/src/mcp/index.ts
@@ -147,7 +147,7 @@
} catch (error) {
if (toolCatalogs.get(client) !== catalog) return
catalog.error = new Error(
- `MCP tool discovery failed for server "${serverName}". Check the MCP server and retry.`,
+ `MCP tool discovery failed for server "${serverName}". Reconnect the MCP server and retry.`,
{ cause: error },
)
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
The refresh-failure message tells the user to "Check the MCP server and retry", but retrying does not issue a new tools/list: catalogTools re-throws the cached catalog.error on every call, and the PR's own test asserts the server's tools/list count stays flat after a failed read ("failed catalog reads must not silently retry"). Recovery requires a new tools/list_changed notification from the server or an explicit reconnect — exactly what the sibling onclose message says ("Reconnect the MCP server and retry"). The instruction in this message does not match the implemented recovery path.
} catch (error) {
if (toolCatalogs.get(client) !== catalog) return
catalog.error = new Error(
`MCP tool discovery failed for server "${serverName}". Check the MCP server and retry.`,
{ cause: error },
)
log.error("MCP tool discovery failed", {|
|
||
| const connectedClients = Object.entries(clientsSnapshot).filter( | ||
| ([clientName]) => s.status[clientName]?.status === "connected", | ||
| const catalogs = await Promise.all( |
There was a problem hiding this comment.
⚪ tools() and toolEntries() duplicate the catalog fetch.
🤖 Fix with your agent
Fix this code review finding (aictrl-dev/cli PR #128, packages/cli/src/mcp/index.ts:681-686):
Problem: tools() and toolEntries() duplicate the catalog fetch
Detail: tools() (lines 652-658) and toolEntries() (lines 681-686) each hand-roll the identical Promise.all(Object.entries(s.clients).map(...catalogTools...)) fetch over the same client set. The file already establishes the "single source of truth" convention for the shared key format (mcpToolKey, with its own doc comment); the shared catalog-fetch block is copy-pasted instead, so any future change to iteration or health handling must be applied twice and can silently diverge between dispatch and catalog.
Suggested fix: Extract a shared helper, e.g. `async function catalogsForClients(): Promise<{ name: string; client: MCPClient; tools: MCPToolDef[] }[]>`, consumed by both tools() and toolEntries(), matching the existing mcpToolKey single-source-of-truth pattern.
Implement the fix on the PR head branch and add a regression test that fails before the fix and passes after.
Why this matters
tools() (lines 652-658) and toolEntries() (lines 681-686) each hand-roll the identical Promise.all(Object.entries(s.clients).map(...catalogTools...)) fetch over the same client set. The file already establishes the "single source of truth" convention for the shared key format (mcpToolKey, with its own doc comment); the shared catalog-fetch block is copy-pasted instead, so any future change to iteration or health handling must be applied twice and can silently diverge between dispatch and catalog.
export async function toolEntries(): Promise<{ toolKey: string; serverName: string }[]> {
const s = await state()
const catalogs = await Promise.all(
Object.entries(s.clients).map(async ([serverName, client]) => ({
serverName,
tools: await catalogTools(client, serverName),
})),
)
return catalogs.flatMap(({ serverName, tools }) =>
Code reviewVerdict: Looks good — only minor / nit comments below. · 🔴 0 · 🟠 0 · 🟡 4 · ⚪ 3 · 0/7 resolved
🤖 Fix all 7 open findings with your agent📋 Out-of-diff findings (7)
Reviewed 9 files · 0 inline · view all 7 findings ↗ aictrl · AI code review for fast-moving teams · aictrl.dev |
Review response — PR #128Verified all seven findings from the catalog-cache review. Keep the requested minimal scope: fix recovery advice and the inconsistent health checks; retain the intentional fail-fast behavior and existing timeout policy. Issues addressed (pushed to this PR)
Review claims verified false (no change needed)None. The other observations describe actual code; their proposed changes are intentional policy alternatives or optional improvements. Intentional behavior and optional improvements (unchanged)
Not addressed hereNo deferred fixes. Five informational/policy observations are intentionally left unchanged, as explained above. VerificationThe real SDK regression failed at reviewed head: status reported failure while both prompt/resource request counters increased. The corrected regression proves healthy RPCs, suppression after discovery failure or closure, and restoration after notification recovery or reconnect. A programmatically added server absent from config remains usable. Full workspace: 1,575 passed, 7 skipped, 0 failed. Focused suite: 86 passed; isolated MCP suite: 23 passed; compiled 0.4.5 headless fixtures: 6 passed. Native build, workspace typecheck, frozen install and formatting passed. Exact-head CI is monitored separately. The platform records have null GitHub comment IDs. Each exact finding ID was joined uniquely against the current review's file, starting line and complete normalized claim title; no location-only ambiguity was accepted. |
Closes #127, Closes #126
Intent
The CLI discards successful startup MCP discovery and lists tools again for the startup catalog and every model turn. A later discovery failure can silently remove persistence tools while the run continues successfully. Retain one catalog per connection, refresh it on tool-change notifications, and prepare CLI 0.4.5 for publication.
Expected Impact on Users
Healthy review runs keep their discovered tools without repeated network discovery. A failed notified refresh or closed connection produces an actionable error before the next model request. Release operators get version validation before build and a longer npm propagation window.
Expected Outcomes
Implementation
Upstream comparison
Checked pinned current source on 2026-10-06. Latest OpenCode MCP lifecycle retains discovered definitions, refreshes on tool-list changes, and removes definitions with disconnected clients. Pi MCP runtime likewise retains definitions and refreshes on notifications. Codex client catalog stores definitions and replaces them after successful explicit refresh. This implementation follows that catalog lifetime. Our intentional policy difference is to expose failed refreshes and stop subsequent catalog reads, following aictrl's fail-fast requirement.
Scope Caveat
The original discovery exception on execution 06c34a3f was not captured; its initiating trigger remains unproven. Silent tool loss is independently reproduced in unchanged CLI 0.4.4. Initially unavailable optional servers keep their existing behavior. This PR prepares 0.4.5; it does not publish packages or deploy an executor.
Test Plan
Verification
263afabd7687f5d3c3db8820941b5b6dde196f2a: CI verify and both CodeQL analyses passed. New-head automated review was explicitly skipped because the template's per-PR maximum is two executions (confirmed in production logs at 2026-10-06 07:15:48 UTC). The supplied catalog-cache review was subsequently delivered on this head. Follow-upfb096d9355181135cf0137cac511e10c3d88eb65has the two review fixes. Its CI verify passed native build (3 tasks), typecheck (6 tasks) and workspace tests (1,575 passed, 7 skipped, 0 failed); both CodeQL analyses and the aggregate also passed. The supplied seven-finding review was responded to and all seven exact-ID verdicts were written and confirmed by read-back (0 failed, 0 unrecorded). Latest-push automatic re-review was explicitly skipped at the same cap; fresh re-review remains a release gate.Risks and Rollout
A failed notified refresh blocks subsequent catalog reads until notification recovery or reconnection. This avoids silently sending fewer tools, while preserving definitions and connection cleanup. Generic CLI idle detection remains opt-in; application PR #5892 enables the existing five-minute stream guard. After approved merge, require green npm publication and independent installation/version verification before updating the executor Docker pin, then validate a sandbox review before production promotion.