ACP Chain Cascade Failure: How We Traced the Bug Three Layers Deep
Last week our ACP pipeline went down for an entire day. On the surface, it looked like agents were simply not responding. Dig deeper and you find three bugs …

Last week our ACP pipeline went down for an entire day. On the surface, it looked like agents were simply not responding. Dig deeper and you find three bugs stacked on top of each other, each one hiding the next. This is the full postmortem.
The Pipeline
First, some context on how our ACP chain is structured:
OpenClaw → acpx-wrapper → acpx 0.3.1 → claude-agent-acp(0.24.2) → api.taijiaicloud.com
When this chain works, it is seamless. When something breaks, errors bubble up from the bottom and you have no idea which layer is actually at fault.
Layer One: Missing Fields in config.json
The first thing we noticed was that acpx refused to start, throwing a cryptic parse error. Digging through the logs, we found that mcpServers was missing from config.json. Not a required field per the documentation, not clearly documented at all — but acpx 0.3.1 reads it unconditionally at startup, and an empty value causes it to crash.
We added the field. acpx came back up. Agents still did not respond.
Layer Two: SDK Version Mismatch
We kept digging. Between acpx 0.15 and claude-agent-acp 0.17, there was a breaking change in the API contract — the schema format for tool_use had changed. Requests sent by the older acpx were silently discarded by the newer claude-agent-acp because it could not parse them.
No error. No log entry. Just silence. That is the worst kind of bug, because you cannot even confirm whether your request reached its destination.
We only found it by doing a line-by-line diff of the two versions of the source code. Once we aligned the versions, requests went through — but the chain was still broken.
Layer Three: Gateway Stripping Environment Variables
The third layer was the actual root cause.
claude-agent-acp needs to read the ANTHROPIC_API_KEY environment variable. Our Gateway was stripping all variables with the ANTHROPIC_ prefix before forwarding — an early security decision to prevent keys from leaking into subprocesses.
The problem: when claude-agent-acp runs as a local binary, it needs that key to initialize itself, not to pass it along to an external API. The Gateway filter was too coarse. It was blocking something it should not have been blocking.
The Fixes
Three layers, three fixes:
Layer One: Fill in all fields — required and optional — in the config.json template, with comments. No more relying on memory or incomplete documentation.
Layer Two: Pin versions explicitly. The acpx-wrapper now hardcodes matched version pairs for acpx and claude-agent-acp. If either one gets updated, both have to be tested together before promotion.
Layer Three: Run claude-agent-acp from a fixed local binary at /opt/homebrew/bin/claude-agent-acp 0.24.2. It bypasses Gateway process injection entirely, which means the env var stripping never applies. Refining the Gateway filter rules is a separate task for another day.
The Hard Part
It was not the fixing. It was the locating.
With three bugs stacked together, every time you fix one you think you are done — and then you discover there is another one underneath. The most time-consuming part of the whole incident was this: you do not know how many layers are left.
Afterward, we added per-hop health checks to the ACP chain. Every hop gets its own ping. If a hop goes down, you know exactly which one. No more chasing from the top. That change was worth more than all three bug fixes combined.
Takeaway
Longer chains mean harder debugging. The solution is not to make chains shorter — sometimes that is not possible — but to make every layer independently observable.
acpx-wrapper is now our standard solution: not just a bypass layer, but a diagnostic layer.