MCP surface wedges permanently while gRPC stays healthy — suspected unbounded outbound Forgejo call under MCP handler serialization (resync_pr_state) #504
Open
opened 2026-08-06 16:45:40 +00:00 by toasterson
·
3 comments
No Branch/Tag specified
main
claude/wi-019ff2f6-create-work-item-ignores-current-phase-d
claude/wi-019fd84d-two-layer-operator-scope-property-mint-p
claude/wi-019fbe99-lane-faults-are-billed-to-work-items-sic
claude/wi-019fd84d-pwa-renders-turnoutcome-by-recovery-axis
claude/wi-019fe0e9-finish-the-pwa-phase-axis-cutover-on-cur
claude/wi-019fd84c-buf-breaking-change-ci-wire-json-across
claude/wi-019ff2f4-reproduce-a-hung-acp-session-on-demand-p
claude/wi-019fd84b-anima-mq-v1-proto-envelope-package-broke
claude/wi-019fc44d-pre-push-test-gate-iterations-are-witnes
claude/wi-019f9079-wi-2-acp-over-vsock-transport-in-anima-r
claude/wi-019ff0c1-test-network-subnet-collision-something
claude/wi-01a00b8c-ci-verdict-pipeline-is-fully-dead-orches
migration-verification-20260818
claude/wi-019fa954-mcp-proxy-registry-sessions-get-one-endp
claude/wi-019fd3c0-pwa-image-tag-lives-in-a-raw-pwa-yaml-th
claude/wi-019fa59c-double-gate-a-second-review-phase-where
claude/wi-019fc2cc-ci-infrastructure-failures-qemu-ssh-rese
claude/wi-019fe0f2-trace-scheduler-decisions-runner-attachm
claude/wi-019fce0d-reviewer-verdicts-are-not-persisted-to-w
claude/wi-019fe0f2-expose-bounded-fleet-and-work-flow-metri
claude/wi-019fe0f7-cut-gate-reads-to-phase-state-verify-adr
claude/wi-019fe0f2-restore-dashboardservice-against-the-mig
claude/wi-019fa59b-publish-transient-chunk-deltas-to-a-per
claude/wi-019fe0ee-checkpoint-typed-progress-witnesses-and
claude/wi-01a00c91-builds-escape-the-per-slot-cargo-target
claude/wi-019fd894-plan-canvas-wi-dependency-graph-with-fro
claude/wi-01a00ec5-a-timed-out-turn-discards-the-worktree-a
claude/wi-01a004d3-build-the-runner-image-in-ci-from-docker
claude/wi-01a00afa-runner-write-safe-directory-entries-to-g
claude/wi-01a009dc-circuit-breaker-state-fails-to-load-on-e
claude/wi-019ff175-adr-0038-acp-has-no-structured-agent-to
claude/wi-01a005fc-the-acp-brief-carries-the-description-bu
claude/wi-01a00462-a-runner-blocked-session-is-unreapable-t
claude/wi-019ff173-pwa-chat-anchor-the-user-s-message-to-th
claude/wi-01a00463-anima-runner-leaks-its-executor-child-on
claude/wi-019fe61e-agent-sessions-have-no-git-push-credenti
claude/wi-019ff2f4-stamp-liveness-witnesses-from-the-acp-ch
claude/wi-019fd84c-witness-and-park-logic-consume-the-turno
claude/wi-019fd84c-runner-stops-relabeling-human-cancels-ca
claude/wi-019fd84c-turnoutcome-closed-14-variant-proto-enum
claude/wi-019fd84c-afk-dispatch-requires-a-requested-outcom
claude/wi-019fd84d-break-glass-reset-reason-every-operator
claude/wi-019fd84d-divergence-report-db-vs-registry-vs-brok
claude/wi-019ff174-pwa-chat-make-a-tool-call-say-what-it-is
claude/wi-019fe0f0-reap-heartbeat-live-but-activity-dead-tu
claude/wi-019fe0ee-persist-scheduling-clocks-and-leases-so
claude/wi-019ff093-mark-implementation-complete-advances-wo
claude/wi-019ff174-pwa-chat-summarise-a-finished-turn-s-wor
claude/wi-019fd894-satisfied-edges-feed-context-dependency
claude/wi-019fe1d3-no-runner-can-reach-dind-so-the-janitor
claude/wi-01a009dc-quota-probe-kills-the-claude-code-lane-a
claude/wi-019fe251-the-pre-push-gate-runs-tests-but-not-cli
claude/wi-019fdd2a-lane-quota-detection-uses-a-total-turn-t
claude/wi-019fd894-orphaned-edges-and-blocked-visibility-pa
claude/wi-019fd894-dependency-edge-api-surface-grpc-mcp-too
claude/wi-019fd893-edge-witness-kind-dependent-satisfaction
claude/wi-01a0057c-sessions-finish-with-no-captured-diff-an
claude/wi-019fa59e-retire-failed-dispatch-count-per-item-ph
claude/wi-019fd895-anima-server-wedges-grpc-mcp-accept-conn
claude/wi-019fce0d-in-flight-ceiling-stranded-working-queue
claude/wi-019fd84c-sidechannel-detection-runner-flags-branc
claude/wi-019fff95-the-runner-s-anima-owner-worktree-tag-is
claude/wi-019fe0f0-add-immutable-graph-version-records-and
claude/wi-019fbe42-inventory-every-psql-kubectl-recipe-oper
claude/wi-019fe0f7-move-exit-bar-state-onto-pinned-phase-in
claude/wi-019ff075-an-orphaned-solstice-dispatch-strands-a
claude/wi-019ff173-pwa-chat-pin-the-pending-permission-esca
agent/19ce972d63dc
agent/972084064aa3
claude/wi-019fe0e9-operator-watch-set-board-backed-by-bound
claude/wi-019fe0e9-clean-addressed-session-exchange-after-l
claude/wi-019fa59e-detect-session-death-by-routability-mand
claude/wi-019fe706-solstice-rabbitmq-is-degraded-job-reques
claude/wi-019fe18f-the-pre-push-test-gate-bills-toolchain-f
claude/wi-019fe0e9-audited-reconcile-session-releases-zombi
claude/wi-019fe0e9-use-one-durable-https-credential-path-fo
claude/wi-019fe6f2-stranded-work-items-drowns-real-strandin
claude/wi-019fe6fd-pwa-data-error-51-items-without-workflow
claude/wi-019fec06-items-with-a-changes-requested-review-ve
claude/wi-019fe0f0-cache-provider-quota-back-off-failures-a
claude/wi-019fe0e9-pin-the-rust-toolchain-and-make-ci-repor
claude/wi-019fce0d-unpark-work-item-mcp-tool-one-audited-un
claude/wi-019fe0ee-assemble-each-agent-turn-from-durable-ho
claude/wi-019fe0ee-codify-the-anima-owned-horizon-boundary
claude/wi-019fe0e7-re-implement-owner-tagged-resource-reape
claude/wi-019fd893-dependency-edges-table-dag-write-path-de
claude/wi-019fce0d-operator-scheduling-controls-over-mcp-ho
claude/wi-019fd84d-interleaved-refusal-witnesses-indefinite
claude/wi-019fdd2a-a-fully-cleared-exit-bar-does-not-advanc
claude/wi-019fd898-corrective-resume-dispatch-to-an-acp-gat
claude/wi-019fe5fd-approve-and-merge-refuses-heads-whose-gr
claude/wi-019fe624-testing-phase-items-with-a-moved-head-ha
claude/wi-019fdd4c-the-24h-in-flight-ceiling-has-no-operato
claude/wi-019fa5a0-move-the-dashboard-onto-the-activity-fan
claude/wi-019fe0f2-add-one-shared-otlp-observability-initia
claude/wi-019fce4e-approve-and-merge-no-already-merged-shor
claude/wi-019fce0d-lane-occupancy-over-mcp-list-sessions-by
claude/wi-019fbdc4-pwa-plan-detail-workflow-graph-define-im
claude/wi-019fd86c-worktree-prep-rebase-parks-items-that-va
claude/wi-019fa96b-session-project-binding-covers-work-item
claude/wi-019fdd87-a-failed-turn-does-not-roll-back-mark-im
claude/wi-019fe0f0-add-audited-cancelsession-that-terminate
claude/wi-019fcde5-work-items-cannot-be-looked-up-by-their
claude/wi-019fa5a0-persist-the-assembled-prompt-to-s3-along
claude/wi-019fa59f-reviews-land-in-the-thread-and-reviewers
claude/wi-019fbdc3-pwa-workflowgraph-node-component-workflo
claude/wi-019fa5a0-single-active-consumer-akh-lanes-and-rel
claude/wi-019faa5d-nothing-prunes-session-tokens-104-083-ro
claude/wi-019fa59c-graph-versioning-pin-in-flight-work-item
claude/wi-019fa59c-anima-work-exchange-lane-queues-and-the
claude/wi-019fa59d-dispatch-branches-on-phasekind-instead-o
claude/wi-019fa59d-anima-session-addressed-exchange-with-pe
claude/wi-019fa5a0-attempt-counter-in-postgres-surviving-ev
claude/node-workflow-overview-wi-4c95cf
claude/wi-019fc826-the-account-quota-probe-polls-the-provid
claude/grpc-api-surface-evaluation-cd6026
claude/wi-019fc9a5-exit-bar-gates-survive-across-pr-generat
claude/wi-019fa59f-boards-read-the-phase-graph-instead-of-a
claude/wi-019fcce5-implementation-sessions-lack-workitem-wr
claude/wi-019f9fc2-anima-has-no-ssh-credential-story-runner
claude/wi-019f9946-a-work-item-filed-without-a-plan-has-no
claude/wi-019fa3a4-reviewers-cannot-build-the-code-they-rev
claude/wi-019fc638-anima-emits-no-traces-metrics-or-structu
agent/4a6b46e02b58
claude/wi-019fc63c-dispatch-has-no-dependency-graph-work-it
agent/e7469dfa074b
claude/wi-019f9fb1-quota-aware-scheduling-budget-provider-c
agent/2cfe41dbf7f5
claude/wi-019faa43-the-janitor-s-two-sweeps-both-exclude-th
claude/wi-019fcc85-move-cargo-target-dir-from-per-runner-to
claude/wi-019f8615-agents-page-show-each-agent-s-live-slot
agent/1c04647808fc
claude/wi-019dbb8e-auth-for-akh-medu-seshat-envoy-gateway-s
claude/wi-019dc57c-p1-5-seshat-ingest-hook-transcripts-libr
claude/wi-019ddfd9-mcp-tool-time-bounded-out-of-hours-sessi
claude/wi-019e8873-native-solstice-ci-dispatch-send-explici
claude/beautiful-leavitt-d2a8d7
agent/8e80f1909fac
agent/baf43b288ad9
claude/wi-019fc9a5-a-failed-git-push-silently-discards-the
agent/0d500af2cd88
agent/c4a84cf02ec2
claude/wi-019f9ea6-dispatched-sessions-hit-scope-walls-mid
claude/wi-019fc826-the-account-quota-probe-hammers-the-prov
claude/wi-019f98b5-opencode-lane-is-dead-weight-every-turn
claude/wi-019fb430-list-working-document-kinds-omits-ci-ret
agent/cb05e821a3bc
agent/a9ae8f731a77
agent/89755cf8eb3e
claude/wi-019fa59d-error-queue-drainer-park-the-work-item-s
claude/wi-019fc2cc-flaky-test-no-reviewer-dispatched-twice
agent/0c195963cd85
agent/21dbfeaf95c9
agent/30dad6db7aa0
agent/917a59a781d8
agent/b557136f8edf
claude/wi-019fc3bd-park-vs-live-session-race-parks-land-whi
agent/2c4165ec8872
claude/wi-019fa59b-phasekind-enum-and-the-fk-constrained-ph
claude/wi-019fa9f3-dispatch-ordering-does-not-prefer-work-i
claude/wi-019fa8f3-the-scheduler-s-unclaimed-wi-warning-re
claude/wi-019fa59a-dispatch-payload-type-and-the-idempotenc
claude/wi-019fa979-the-reviewer-slot-reservation-is-applied
claude/wi-019fa59b-error-envelope-and-per-lane-error-queues
claude/wi-019fa7d9-an-executor-must-be-able-to-rebase-its-o
claude/wi-019fa59b-seed-the-default-phase-graph-with-ci-to
claude/wi-019fbe41-zombie-session-repair-is-a-raw-sql-sweep
claude/wi-019fbe97-grpc-services-and-mcp-handlers-each-reim
claude/wi-019fbe41-a-stale-recorded-pr-head-breaks-ci-dispa
claude/wi-019fbe41-operator-forensics-needs-psql-no-queryab
claude/wi-019fa59e-retire-anima-sched-max-active-for-runner
claude/wi-019fb424-derive-every-git-remote-config-from-the
claude/wi-019fa3d2-the-ai-gate-ships-no-review-instructions
claude/wi-019f9e06-implement-emit-contextwindowstats-from-p
claude/wi-019fb464-a-review-analysis-wi-that-delivers-its-f
claude/wi-019fa3a4-approve-and-merge-fails-permanently-on-f
claude/wi-019f9e06-implement-a-b-stats-comparison-harness-4
claude/wi-019fb477-per-executor-lane-state-leaks-on-session
claude/wi-019fa59a-anima-mq-crate-topology-declaration-for
claude/wi-019fba1f-witness-output-first-classification-coun
claude/wi-019f9e06-implement-strategy-specific-stats-for-co
claude/wi-019f841f-quick-reply-dedup-options-render-as-bull
archive/wi-410-20260807
claude/wi-019fa5a0-owner-tagged-reaper-for-containers-workt
archive/wi-394-20260807
claude/wi-019fa59e-per-item-phase-revisit-cap-as-workflow-c
claude/wi-019fa59f-surface-the-operator-watch-set-on-the-bo
claude/wi-019fb477-deliverable-reconcile-fails-every-tick-p
claude/wi-019fa59e-migrate-exit-bar-columns-to-phases-and-v
claude/wi-019f8b49-wi-5b-drop-work-items-status-after-the-p
claude/wi-019f841f-live-roster-presence-real-availability-o
claude/wi-019faf95-implementer-runs-the-touched-crates-test
claude/wi-019faf96-migration-086-cannot-run-under-sqlx-inde
claude/wi-019faf95-corrective-dispatch-resumes-the-prior-se
claude/wi-019faf95-refusal-is-not-an-empty-turn-distinct-wi
claude/wi-019fa59f-24-hour-in-flight-ceiling-that-parks-wit
claude/wi-019fb477-rejected-work-item-is-neither-dispatchab
claude/wi-019faf95-session-accounting-capture-per-session-t
claude/wi-019fb40f-a-terminal-worktree-prep-rejection-leaks
claude/wi-019f90ad-scheduler-over-assigns-in-bursts-8-assig
claude/wi-019f860b-pwa-chat-config-pickers-replay-a-dead-se
claude/wi-019f85b9-wi-3-pwa-collapsible-subthread-ui
claude/wi-019f841f-thread-scoped-activity-streaming-retire
claude/wi-019f841e-forward-agent-message-content-blocks-end
claude/wi-019f85b9-wi-2-runner-delegation-reporting-subthre
claude/wi-019f52d2-review-loop-starvation-reviewer-re-dispa
claude/wi-019eab9e-feature-mcp-read-tools-for-session-artif
claude/wi-019e846d-catch-unwired-dead-code-pub-in-pub-mod-h
claude/wi-019faa49-the-per-executor-cap-does-not-hold-acros
claude/wi-019f615e-sia-populate-external-ref-external-state
claude/wi-019dd0b7-wipsnapshot-retention-per-wi-cap-age-pru
claude/wi-019dc0e8-git-remote-url-normalization-ssh-vs-git
claude/wi-019fa9c4-ci-green-is-a-sticky-boolean-with-no-com
claude/wi-019f85b9-wi-1-subthread-domain-model-parent-threa
claude/wi-019f9087-pwa-workitemdetail-request-changes-appro
claude/wi-019f9079-wi-1-spike-gates-the-rest-hypervisor-pic
claude/wi-019f862a-project-overview-in-flight-section-split
claude/wi-019f903c-ci-builds-pushes-multi-arch-deploy-image
claude/wi-019f841e-resource-mentions-in-the-composer-files
claude/wi-019f57aa-message-part-streaming-forward-acp-delta
claude/wi-019e93df-mcp-ambient-get-work-item-detail-full-wi
claude/wi-019e846d-in-server-merge-engine-hardening-anima-g
claude/wi-019dbb9c-agent-annotation-pathway-broken-pat-lack
claude/wi-019faa38-a-ci-green-work-item-awaiting-review-is
claude/wi-019fa5a0-register-session-gate-reject-a-redeliver
claude/wi-019fa902-the-ci-red-corrective-loop-hands-the-imp
claude/wi-019fa945-per-executor-reservation-leaks-a-capped
claude/wi-019f90f9-pwa-delivery-board-is-blind-to-the-phase
claude/wi-019fa7c9-anima-runner-exec-slots-name-is-also-par
claude/wi-019fa59c-distinguish-an-empty-turn-from-a-failed
claude/wi-019f9aa3-ci-lints-against-an-unpinned-toolchain-d
claude/wi-019f9fb8-dispatches-that-fail-before-the-session
claude/wi-019fa59b-progress-witness-pr-aware-per-attempt-re
claude/wi-019f9a17-every-native-ci-dispatch-fires-twice-rec
claude/wi-019f9698-ci-red-corrective-loop-must-distinguish
claude/wi-019f9ec2-a-dropped-amqp-connection-silently-ends
claude/wi-019f84ec-enforce-session-project-binding-on-write
claude/wi-019f903c-fleet-observability-metrics-alerts-for-r
claude/wi-019fa7c6-a-turn-that-only-repairs-the-base-repo-s
claude/wi-019fa7c6-a-rejected-assignment-re-dispatches-ever
claude/wi-019f9b4b-reject-never-reaches-the-implementer-lin
claude/wi-019f9b6e-reviewer-only-is-not-expressible-keeping
claude/wi-019f9ac1-quota-aware-lanes-meter-usage-over-acp-s
claude/wi-019f99ca-native-ci-falls-back-to-a-job-script-tha
claude/wi-019f8610-remove-the-global-anima-sched-max-active
claude/wi-019f990a-analysis-work-parks-itself-17-of-20-park
claude/wi-019f9aa3-a-model-provider-429-is-indistinguishabl
claude/wi-019f99a4-dispatch-queue-ignores-priority-strict-f
claude/wi-019f83ce-wi-a-backend-escalation-core-schema-type
claude/wi-019f98d9-you-cannot-follow-a-wi-along-its-route-1
claude/wi-019f98b5-a-failed-turn-is-indistinguishable-from
claude/account-quota-meter
claude/wi-019fa2fc-the-akh-gateway-advertises-6-slots-for-a
claude/wi-019fa243-every-re-dispatch-throws-away-the-previo
claude/wi-019fa2de-a-review-session-is-dispatched-alongside
claude/wi-019f9e35-worktree-base-branch-is-one-global-env-v
claude/thread-window-redesign-2360d9
claude/wip-lift-rehook
claude/wi-019f9b1d-review-slot-reservation-leaks-when-a-rev
claude/wi-019f99ac-create-project-leaves-new-projects-undis
claude/wi-019f9980-pin-ci-verdicts-to-the-commit-under-test
claude/wi-019f8438-wi-d-frontend-permission-card-render-hum
claude/wi-019f8438-wi-c-runner-interceptor-checkpoint-detec
claude/wi-019f8438-wi-b-executor-token-scope-stripping-q7c
claude/wi-019f8602-reopen-work-items-when-their-orphaned-se
claude/wi-019f90cf-review-analysis-wis-never-close-no-phase
claude/wi-019f90fe-mcp-update-work-item-advertises-status-v
claude/claude-lane-opus-5
claude/wi-019f17e4-user-settings-foundation-store-registry
claude/wi-019f1779-per-user-configurable-inbox-auto-clear-w
No results found.
Labels
Clear labels
Compat/Breaking
Breaking change that won't be backward compatible
Kind/Bug
Something is not working
Kind/Documentation
Documentation changes
Kind/Enhancement
Improve existing functionality
Kind/Feature
New functionality
Kind/Security
This is security issue
Kind/Testing
Issue or pull request related to testing
Priority
Critical
The priority is critical
Priority
High
The priority is high
Priority
Low
The priority is low
Priority
Medium
The priority is medium
Reviewed
Confirmed
Issue has been confirmed
Reviewed
Duplicate
This issue or pull request already exists
Reviewed
Invalid
Invalid issue
Reviewed
Won't Fix
This issue won't be fixed
Status
Abandoned
Somebody has started to work on this but abandoned work
Status
Blocked
Something is blocking this issue or pull request
Status
Need More Info
Feedback is required to reproduce issue or to continue work
No labels
Compat/Breaking
Kind/Bug
Kind/Documentation
Kind/Enhancement
Kind/Feature
Kind/Security
Kind/Testing
Priority
Critical
Priority
High
Priority
Low
Priority
Medium
Reviewed
Confirmed
Reviewed
Duplicate
Reviewed
Invalid
Reviewed
Won't Fix
Status
Abandoned
Status
Blocked
Status
Need More Info
Milestone
Clear milestone
No items
No milestone
Projects
Clear projects
No items
No project
Assignees
Clear assignees
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".
No due date set.
Dependencies
No dependencies set
Reference
toasterson/Anima#504
Loading…
Reference in a new issue
No description provided.
Delete branch "%!s()"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Observed 2026-08-06 ~16:39Z on anima-server (pod anima-server-65c4d85cf6-s2bhh, container started 16:23:35Z after a 137 kill of the previous instance).
Symptom
/mcpreturns 504 upstream timeout (~15s) from the ingress — including a bareinitializewith no session. Wedged for 30+ minutes and counting./healthanswers 200 in ~40ms the whole time; gRPC runner traffic (scheduler assignments, register-session) continues normally per server logs.audit.approve_and_merge.merged, Anima PR #500). No MCP request logging after that — later requests never reach handler logging.Timeline (from the wedged instance's logs + client side)
update_work_item(touch on WI 019fce4c) still succeeded.resync_pr_statecall for WI 019fce4c-d90f-7311-aa6c-50dd754f9f29 (PR toasterson/akh-medu#289) timed out client-side; every MCP request since then 504s, including freshinitialize. Forgejo was responding slowly around that time (alist_repo_pull_requestsfor akh-medu returned 279KB after a long wait).Hypothesis
resync_pr_stateperforms an outbound Forgejo read server-side. If that HTTP call has no timeout and MCP handlers serialize on a shared lock/session state, one slow Forgejo response wedges the entire MCP surface permanently — nothing recovers it short of a process restart. Suggested fixes: put a hard timeout on all outbound Forgejo calls made from MCP handlers, and/or stop holding the MCP serialization across outbound I/O.Related scheduler observations from the same logs (separate but adjacent):
per-executor reservation leak: capped lane reads as full with zero live sessions(executor=tecton, leaked_count=1) — seen at 16:05:33Z and again 16:38:41Z. Each server restart/session death leaks the tecton lane reservation; ready+pinned WIs then sit undispatched until the reconcile heals the counter. This blocked akh-medu WIs 019fce4c/019fce4d for ~20h on 2026-08-05→06.ready WI unclaimed: project not dispatch-enabled— ~13 illumos-project WIs warned every scheduler tick (age up to 965261s); either enable dispatch for the project or stop queueing them ready.substantial output without diff — prose is not delivery, classified Empty;no-progress streak tripped but WI has a live session — deferring parkrepeats without the park ever landing, so the loop never stops.Previous container exit: code 137 at 16:23:34Z (reason "Error") — worth checking memory limits/liveness separately.
Filed by the akh-medu babysitting session; restarting anima-server to restore MCP service.
Recurred within ~15 minutes on a fresh pod — this is not a one-off and my
resync_pr_statewas not the (only) trigger.Timeline, second occurrence:
anima-server-78bfd64d46-fn8srcame up (new ReplicaSet hash, so likely a Flux-applied newer image after today's Anima merges, replacing the pod from my 16:47Z rollout restart).resync_pr_stateops served fine (Anima PRs #485/#481 hand-rebase resyncs, actor 019d8deb).closed live sessions on implementation complete(WI 019f90cd, akh-medu Release Alpha)./mcp504s again on bareinitialize(~15s at the ingress). No MCP log lines after 17:11:25./health+ gRPC scheduler traffic still healthy.Pattern across both wedges: a handful of MCP ops that do outbound Forgejo I/O (resync, audit/merge, session-close) complete, then the surface hardlocks with no further request logging. Consistent with a poisoned/never-released lock or an exhausted single-worker MCP executor rather than one specific slow call.
Restarting via pod delete this time (rollout-restart annotations appear to get reverted by GitOps reconciliation).
Corroborating correlation for the MCP wedge — akh-gateway dispatch.
akh-medu WI 019fbf62 (Phase 44n) now carries an operator-hold park reason: "dispatch to akh-gateway wedges anima-server (2026-08-06); unpark after gateway recycle verified". The timing lines up with both wedges I logged here:
So the trigger may not be MCP-handler outbound Forgejo I/O after all (or not only that) — a dispatch/attach to the akh-gateway (anima-akh-gateway-thoth on archibald) that hangs appears able to take the MCP surface down with it while gRPC scheduling continues. Worth checking what shared state (lock/executor) the gateway-dispatch path and the MCP layer contend on.
Live counter-datapoint: thoth-lane dispatch of akh-medu WI 019fa83d at 18:18:37Z succeeded (session running, PR #294 opened) with MCP staying healthy — so it's not every gateway dispatch, plausibly only ones that hang.
Gateway attach-stream death confirmed as the akh-lane blocker (and the recycle fixes it).
After today's repeated anima-server pod replacements (2 crash/wedge restarts + at least 2 Flux image rolls), the akh gateway's runner (
019f2f96, advertises tecton+thoth) stopped re-attaching entirely — 30+ minutes of complete log silence inanima-akh-gateway-thoth-1while the other two runners re-attached to each new pod within seconds. Result: both akh lanes dead, ready+pinned WIs unclaimed indefinitely, no scheduler warning names this condition (the lanes just silently vanish from dispatch).docker restart anima-akh-gateway-thoth-1at 20:41:45Z fixed it instantly: attach + hello + both akhs advertised within 200ms, and both stalled akh-medu WIs (019fa83d thoth, 019fce4c tecton corrective-resume) were assigned 23 seconds later.Suggested hardening: (1) the gateway's attach loop should reconnect with backoff forever like the other runners evidently do — whatever state kills its stream after repeated server bounces needs a supervisor; (2) the scheduler should WARN when a pinned-executor lane has ready work but the advertising runner hasn't attached since boot — today that condition was only inferable by elimination.
(Also for the record on this issue's earlier thread: the 20:10:38Z rebind-window abandonment of both akh sessions traces to the 19:58–20:01 disk-full outage on archibald — root disk hit 100%, Forgejo Postgres crash-looped, fixed by freeing 111GB. Separate incident, same evening.)