Skip to content

zephyr-cp: fix the make DEBUG=1 build (quote EXTRA_CONF_FILE) - #24

Closed
tyeth wants to merge 15 commits into
zephyr-cp-frozen-modulesfrom
zephyr-cp-debug-makefile-fix
Closed

tyeth wants to merge 15 commits into
zephyr-cp-frozen-modulesfrom
zephyr-cp-debug-makefile-fix

Conversation

@tyeth

@tyeth tyeth commented Sep 9, 2026

Copy link
Copy Markdown
Owner

make DEBUG=1 BOARD=… was broken: with DEBUG=1 the Makefile builds a
;-separated EXTRA_CONF_FILE (board.conf;debug.conf) and passed it to
west build unquoted, so the shell treated the ; as a command separator —
west ran with only board.conf, and the shell then tried to execute
…/debug.conf as a program.

Fix: quote the whole -Dzephyr-cp_EXTRA_CONF_FILE=… argument so the list
reaches CMake intact.

Verified on hardware (Pico 2 W): a full make DEBUG=1 build completes
(reaches Completed 'zephyr-cp'), debug.conf is merged (CONFIG_DEBUG=y,
CONFIG_LOG_MODE_IMMEDIATE=y, CONFIG_UDC_RPI_PICO_STACK_SIZE=2048), and the
resulting image boots to a working REPL (print("DBG-REPL-OK", 6*7) → 42).

Note: debug.conf sets CONFIG_LOG_MODE_IMMEDIATE=y, which serialises logging
in-thread over UART0 and distorts timing — a DEBUG build must not be used to
measure latency.

🤖 Generated with Claude Code

tyeth and others added 15 commits September 8, 2026 22:33
debug.conf bumps the DWC2 and nRF UDC thread stacks because verbose
logging with LOG_MODE_IMMEDIATE formats in-thread and needs more stack,
but the rpi_pico one was missed. It defaults to 512 bytes, which
overflows during USB enumeration on a Pico 2 W:

  ***** USAGE FAULT *****
    Stack overflow (context area not valid)
  >>> ZEPHYR FATAL ERROR 2: Stack overflow on CPU 0
  Current thread: 0x200034b0 (usbd@50110000)

The board dies partway through configuration, which looks like a bad USB
come-up rather than a stack problem. Independent of any Bluetooth work —
it affects any RP2040/RP2350 board built with DEBUG=1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The CYW43439's Bluetooth controller shares the WiFi gSPI bus rather than
sitting on a UART, so it needs the new shared-bus HCI driver (see the
companion zephyr branch). With a bt_hci device in the devicetree, the
port's existing Kconfig turns CONFIG_BT on by itself and common-hal
_bleio needs no board-specific work.

Wire it up for the Pico 2 W:

- board overlay: add the infineon,cyw43-bt-hci node as a child of the
  stock infineon,airoc-wifi node, and point zephyr,bt-hci at it. The
  hyphenated chosen name matters — the port's Kconfig derives CONFIG_BT
  via dt_chosen_enabled(zephyr,bt-hci).
- board conf: deeper system workqueue stack and HW stack protection, as
  the WHD WiFi, cybt and coexistence call chains are deeper than the
  port defaults.
- compat2driver: map infineon_cyw43_bt_hci to bluetooth/hci, which is
  what makes zephyr2cp.py set _bleio. Added by hand rather than
  regenerating the file, which would churn 500+ unrelated lines against
  the current Zephyr revision.

Builds at 67.1% flash and 43.4% RAM, up from 59.7%/38.0% without BLE.
Not hardware-tested.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Four changes needed to get BLE actually working on a Pico 2 W, all found
by bisecting the failures over SWD.

The settings partition was 2K at 0x180800 — neither erase-sector aligned
nor large enough for one of the RP2350's 4K sectors. flash_area_get_sectors()
returned zero sectors, settings_nvs then computed its sector size from an
uninitialised struct and failed with -EDOM, so bt_enable() returned -33
before ever opening the HCI driver. That surfaced as a bare "OSError: 33"
from "import _bleio", which is where the adapter is enabled (the _bleio
module's __init__ calls common_hal_bleio_adapter_set_enabled). main has
since grown it to one aligned 4K sector at 0x17f000, which fixes the -EDOM
but is still one sector short: nvs_mount() rejects fewer than two sectors
with -EINVAL. Give settings two sectors at 0x17e000, taking the extra one
from the code partition. nvm and circuitpy stay at 0x180000 and 0x181000,
where cptools/check_partitions.py requires them to match
ports/raspberrypi, so the CIRCUITPY filesystem is not moved.

A board whose settings sectors hold content from a previous layout then
fails differently: NVS reads it as "all sectors closed" and refuses to
mount with -EDEADLK. CONFIG_NVS_INIT_BAD_MEMORY_REGION lets it reclaim a
region it does not recognise, so the first boot after a layout change
recovers on its own instead of needing a manual erase over SWD.

The CYW43439's BT controller firmware (CYW4343A2_001.003.016.0065.0000)
does not implement the Bluetooth 5 extended advertising/scanning commands.
It rejects LE Set Extended Scan Parameters (0x2041) with status 0x01
"Unknown HCI Command", so scanning failed with -EIO. The port defaults
BT_EXT_ADV on, which is right for the nRF parts but wrong here, so turn it
off for this board and let the host use the legacy 0x200B/0x200C commands.

Also raise the system workqueue stack and enable HW_STACK_PROTECTION (the
MPU turned several silent corruptions into clean, named faults while
debugging this), and enable BT HCI driver/host debug logging in
debug.conf.

Verified on hardware: patchram loads over the shared gSPI bus, the
controller reports BD_ADDR 2C:CF:67:B7:62:AC (= WiFi MAC + 1), and a scan
returns nearby advertisers by name. That run used the earlier layout with
settings at 0x181000; the 0x17e000 placement is build-tested only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Same CYW43439 as the Pico 2 W, and the shared-bus HCI driver needed no
board-specific work, so this is the Pico 2 W change applied to the RP2040
board:

- overlay: hang the infineon,cyw43-bt-hci node off the stock
  infineon,airoc-wifi node and point zephyr,bt-hci at it.
- overlay: the settings partition had the same defect as the Pico 2 W's --
  originally 2K at 0x180800, neither erase-sector aligned nor big enough
  for one of the RP2040's 4K sectors; main has since made it one aligned
  4K sector at 0x17f000, which is still one short of the two nvs_mount()
  insists on, so bt_enable() fails before opening the HCI driver. Give
  settings two sectors at 0x17e000, taken from the code partition. nvm
  and circuitpy stay at 0x180000 and 0x181000 to match ports/raspberrypi
  (cptools/check_partitions.py checks this), so CIRCUITPY is not moved.
- conf: BT_EXT_ADV=n, since this controller rejects the Bluetooth 5
  extended advertising commands, plus a deeper system workqueue stack.

HW_STACK_PROTECTION is deliberately left off here: it earned its place on
the Cortex-M33 but costs RAM, and this is the constrained board.

Builds at 68.67% flash and 84.32% RAM of 264K, so BLE fits with about 42K
to spare.

Verified on hardware after the "ranges" fix (now on main): the adapter
comes up as WiFi MAC + 2 and a legacy LE scan returns nearby advertisers.
That run used the earlier layout with settings at 0x181000; the 0x17e000
placement is build-tested only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The custom board workflow predates the zephyr-cp port and cannot build its
boards:

- It relies on make's default goal to produce firmware.*, but the zephyr-cp
  Makefile's default goal is the Zephyr ELF, so the firmware.* copies the
  artifact upload globs for are never made and the upload is empty. Name the
  firmware.<ext> targets explicitly instead, the way
  tools/build_release_files.py already does, via a small helper that resolves
  CIRCUITPY_BUILD_EXTENSIONS for either kind of port.
- The port deps action does not check out hal_rpi_pico's cyw43-driver
  submodule, which the CYW43 shared-bus Bluetooth transport needs for the
  controller patchram. Fetch it during port setup.

Also point the west manifest at the fork branches carrying the Pico 2 W
Bluetooth work so the artifact is BLE-capable. That override is CI-only and
is marked as such -- it must be reverted before merging.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A CircuitPython program serving HTTP over the station link to a browser
that opens several connections per page, with an NTP or mDNS UDP socket
beside it, sits at Zephyr's default of six net_contexts: the listener,
every accepted client and every UDP socket come out of the same pool, and
the seventh socket() or accept() fails with ENOMEM. The BLE stack does not
compete for these, but the hub use-case this board is aimed at -- WiFi
station + HTTP + BLE advertising and scanning at the same time -- does.

Double the pool (CONFIG_NET_MAX_CONTEXTS=12) and the connection-handler
table that scales with it (CONFIG_NET_MAX_CONN=16); ZVFS_OPEN_ADD_SIZE_NET
follows automatically. Pico 2 W only: the RP2040 board has 42 KB left and
runs no server.

Build (local, same tree as ci/pico2w-ble-assets @ 94bfd47):
  FLASH 1,146,096 B (unchanged)
  RAM   229,144 -> 231,112 B (+1,968 B, 43.03% -> 43.40%)
  .config: NET_MAX_CONTEXTS=12 NET_MAX_CONN=16 ZVFS_OPEN_ADD_SIZE_NET=12

Not run on hardware. Python side that uses it:
tyeth/deepsleep_espnow_wifi_and_ble_env_collector "max_sockets" in
collector/config.json (caps.MAX_SOCKETS).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The rpi_pico UDC driver's internal thread stack was only raised in
debug.conf, so release builds kept Zephyr's 512-byte default and dropped
off USB partway through a _bleio scan -- CDC port and CIRCUITPY volume
both disappearing, with a USB device/stack error on reset. Move it to
prj.conf so every build gets it. Verified on a Pico 2 W: a 12s active
scan yielding 925 reports from 20 distinct devices now completes with
USB intact, where the same workload previously killed the bus.

Zephyr defaults BT_MAX_CONN to 1, which would confine the Pico 2 W hub
to a single peer and force a connectionless node protocol. The CYW43439
is not the constraint, so raise it to 4 and let nodes connect to sync
data.

Together these cost 5,752 B of RAM (231,112 -> 236,864, 43.40% ->
44.48%) and 408 B of flash.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
start_scan(timeout=...) never ended on the Pico 2 W: the ScanResults loop ran
until Ctrl-C, and the KeyboardInterrupt then surfaced on the first statement
after the loop (stop_scan() in one run, the print() after it in another),
which looked like a stall in scan teardown or in CDC output.

gdb on the running board showed nothing blocked. The main thread was in
common_hal_bleio_scanresults_next() waiting for the next entry with done
still false; bt_dev.flags had BT_DEV_SCANNING set; ncmd_sem.count was 1 and
sent_cmd 0 (no HCI command outstanding); the gSPI bus mutex was free; the
BT RX poll thread was in its 4 ms k_msleep; every work queue and USB thread
was pended idle. stop_scan() itself took 5-6 ms when called explicitly.

The cause is in Zephyr's host: bt_le_scan_param.timeout is only passed to the
controller on the extended-scanning path (LE Set Extended Scan Enable carries
a duration and the controller reports LE Scan Timeout). start_le_scan_legacy()
never reads it, and the legacy path is what CONFIG_BT_EXT_ADV=n selects --
which a controller without extended advertising, such as the CYW43439,
forces. So the timeout was silently ignored and the scan ran forever.

Keep the deadline in the adapter and enforce it from bleio_background(),
called from port_background_task() on the main thread, where bt_le_scan_stop()
is safe to call (it blocks on an HCI round-trip, which must not happen on the
system work queue that also runs the USB CDC console). The ScanResults
iterator finishes at the deadline as it does on nRF and ESP32.

Measured on hardware, Pico 2 W: timeout=3 s -> loop exited after 3.0 s;
timeout=1 -> 1.0 s; timeout=0.5 -> 0.53 s; timeout=2 -> 2.05 s (30 reports).
Before the change a timeout=3 scan was still yielding at 57 s (1600 reports).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The busio UART receive callback -- which also serves the USB CDC console --
queued bytes without waking the main thread. While the REPL is reading it
spins, so that path was fine, but after code.py finishes the supervisor parks
in port_idle_until_interrupt() waiting for "any key", and nothing woke it for
a keypress until the next timed wake-up. Signal the main task from the
callback, as the other ports' console receive paths do.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Nothing in the port, the SoC defaults or these boards' Kconfig turned
CONFIG_NET_TCP on, so every SOCK_STREAM socket failed inside
net_context_get() with EPROTOTYPE -- which socketpool reported as a generic
"Out of sockets" -- and the web workflow's listener never opened either.
Enable TCP on both boards, and raise the fd table to 16: Zephyr sized it from
the subsystems' declared needs (13 here), which the DHCPv4 server and the
socket-service eventfd eat into.

Three socketpool fixes found on the way:

- A socket object is allocated with a finaliser and zeroed before
  zsock_socket() runs. If that failed, num stayed 0 and the finaliser later
  called zsock_shutdown(0). fd 0 belongs to the socket service's eventfd,
  whose shorter vtable has no shutdown slot, and the CPU branched into
  cdc_acm_1's data: "USAGE FAULT / Illegal use of the EPSR", pc 0x200003f8,
  lr z_impl_zsock_shutdown, system halted. Mark the object closed (num = -1)
  before attempting creation.
- Report errno from a failed zsock_socket() (ENOENT: no net_context left,
  EPROTOTYPE: protocol off, ENFILE: fd table) instead of the fixed message.
- Accepted sockets did not inherit the listener's timeout: the Python-facing
  accept path only set num, so a fresh object read as timeout 0 and the ssl
  layer's first recv on an accepted TLS connection raised EAGAIN before the
  handshake could finish. Inherit it, as the other ports do.

Measured on a Pico 2 W with the softAP up: 11 TCP sockets open before ENOENT
(12 net_contexts, one held by the DHCPv4 server).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
wrap_socket(server_side=True) configured mbedTLS with VERIFY_REQUIRED
whenever the context had the default root bundle attached, which every
ssl.SSLContext() does. On a server that means "require a client
certificate", so a plain HTTPS server -- load_cert_chain() and nothing else
-- failed every handshake with MBEDTLS_ERR_SSL_NO_CLIENT_CERTIFICATE
(-0x7480) and the client saw a timeout.

Default server-side contexts to VERIFY_NONE, as CPython's do (CERT_NONE for
servers); a context that loaded its own CA with load_verify_locations() still
verifies clients. Client-side behaviour is unchanged.

Seen on a Pico 2 W serving a captive portal over its softAP; the same code
runs on every port using the shared mbedTLS module.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… boards

Two Kconfig gaps kept an HTTPS server (or client) from ever completing a
handshake on these boards:

- Zephyr's mbedTLS 4 default serves secp256r1 through the p256-m PSA driver
  and leaves the builtin ECP module out. mbedTLS then defines the dummy
  MBEDTLS_ECP_MAX_BITS 1, and ssl.h sizes the TLS 1.2 premaster buffer
  (union mbedtls_ssl_premaster_secret._pms_ecdh[MBEDTLS_ECP_MAX_BYTES]) from
  it: sizeof(handshake->premaster) was 1 in the firmware (DWARF). Every ECDHE
  key agreement then failed -- "psa_raw_key_agreement() returned -138
  (-0x008a)", PSA_ERROR_BUFFER_TOO_SMALL for a 32-byte shared secret -- and
  the raw PSA status leaked through mbedtls_ssl_read() as OSError 138 on the
  first read of any TLS connection. Disabling the p256-m driver brings the
  builtin ECP module back and the buffer is 32 bytes again. This is an
  upstream sizing bug (ssl.h should size that buffer from PSA when the ECP
  module is absent); the Kconfig is the workaround until it is fixed.

- PEM parsing was not compiled in, so load_cert_chain() with the usual
  fullchain.pem / key.pem could only fail; the collector's certstore is PEM.

Verified on a Pico 2 W: an RSA-2048 PEM certificate served from CIRCUITPY
over the softAP; an ESP32-C6 client completed the handshake and received
"HTTP/1.0 200 OK" in 2.0 s. Flash +1,868 B (PEM) and +4,716 B (builtin ECP
instead of p256-m) on the Pico 2 W; no static RAM change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
wifi.radio.start_ap() returned silently with ap_active False: the whole body
of common_hal_wifi_radio_start_ap() was commented-out ESP-IDF code. The AIROC
driver already implements ap_enable/ap_disable through wifi_mgmt, so this is
the hook-up, plus the pieces around it.

start_ap(ssid, password, channel, authmode, max_connections) issues
NET_REQUEST_WIFI_AP_ENABLE with the CircuitPython authmode mapped to
WIFI_SECURITY_TYPE_NONE / PSK / SAE (the driver brings every PSK variant up
as WPA2-AES-PSK; WPA1/TKIP is not offered separately). max_connections gets
the ESP32 range check but cannot be applied: the CYW43439's association
limit is fixed in the controller. The AP gets 192.168.4.1/24 by default (as
on ESP32), settable with set_ipv4_address_ap(), and the DHCPv4 server starts
with it (pool of 8 above the AP address, DNS pointing at the AP, RFC 8910
option 114 so phones and laptops find a captive portal automatically);
start_dhcp_ap()/stop_dhcp_ap() remain explicit controls. stop_ap() tears it
all down. ap_active asks the driver (WIFI_MODE_AP in the interface status)
rather than trusting a flag. stations_ap lists the stations the driver
reported through NET_EVENT_WIFI_AP_STA_CONNECTED/DISCONNECTED with the IPv4
address from the DHCP lease table (no RSSI: wifi_mgmt's station events carry
none). set_ipv4_address() for the station side, start_dhcp() and stop_dhcp()
are implemented on the way.

The AIROC driver runs the AP on the same net_if as the station and refuses
AP_ENABLE with -EBUSY while the station is associated, so there is no
simultaneous AP+STA on these boards: disconnect() first. Soft reboot tears
the AP down (wifi_reset), and start_ap() disables a running AP the driver
still reports before enabling, so the two cannot get out of step -- a soft
reboot was seen to leave the driver's flag set when its stop failed.

Needs the AIROC driver fix in tyeth/zephyr (branch airoc-softap-fixes):
without it the driver's chanspec composition makes the firmware reject every
2.4 GHz channel, and a station leaving took the AP's interface dormant.

Verified on a Pico 2 W with an ESP32-C6 as the station: start_ap() completes
in 0.15 s; the C6 associates in 7 s, gets 192.168.4.2 from the DHCP server
with gateway/DNS 192.168.4.1, pings the AP in 5-8 ms; stations_ap shows it
with its lease, empties when it disconnects (the AP stays active), and shows
it again on rejoin; the AP survives stop_ap()/start_ap() cycles and a soft
reboot. Flash +9,444 B, RAM +3,568 B (DHCPv4 server + socket service).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The port had no frozen-module support: its Python-driven build never runs
make, and the freeze pipeline lives in py/circuitpy_mpconfig.mk. Reimplement
the pipeline in cptools/build_circuitpython.py behind a per-board opt-in,
FROZEN_MPY_DIRS = ["frozen/<lib>", ...] in circuitpython.toml (paths relative
to the repository root, as $(TOP)/... is in mpconfigboard.mk):

- tools/preprocess_frozen_modules.py stages the trees into
  <build>/frozen_mpy (repo directory dropped, __version__ filled in, examples
  and tests left out), before the qstr pass;
- MICROPY_QSTR_EXTRA_POOL / MICROPY_MODULE_FROZEN_MPY are added to the flags
  for the qstr pass as well as the compile -- Q(.frozen) and the sys.path
  entry that uses it are behind MICROPY_MODULE_FROZEN;
- after genhdr/qstrdefs.generated.h and root_pointers.h exist, mpy-cross
  compiles each module (-s with the module path, so mpy-tool derives the
  frozen name from it) and tools/mpy-tool.py -f -q emits frozen_content.c,
  fed the same collected qstr list that produced the generated header so the
  frozen pool numbers from MP_QSTRnumber_of correctly. tools/makemanifest.py
  is bypassed: it insists on genhdr/qstrdefs.preprocessed.h, which this
  builder never produces. Only MPY freezing: MICROPY_MODULE_FROZEN_STR would
  reference the mp_frozen_str_* tables makemanifest emits.
- frozen_content.c is compiled with the no-qstr sources (it defines its own
  MP_QSTR_* enum values and must never go through extraction).

pre_zephyr_build_prep.py builds mpy-cross first when a board freezes modules
(honouring MICROPY_MPYCROSS), and tools/ci_fetch_deps.py learns the toml key
so CI initialises the right frozen/ submodules (it had a TODO for this).

The Pico W opts in with adafruit_ble: it has ~42 KB of heap and a BLE node
cannot otherwise load the library. Measured on a Pico 2 W (same core, same
flags): importing adafruit_ble plus its advertising.standard and
services.nordic modules costs 10,384 B of heap frozen vs 21,792 B from .mpy
files on CIRCUITPY (gc.mem_alloc() delta after gc.collect(); gc.mem_free()
is not usable here, the split heap grows on demand), 0.055 s vs 0.132 s.
Freezing the 20 modules adds 28,336 B of flash and no static RAM on the
Pico W (1,208,636 -> 1,236,972 B with the rest of this series). The Pico 2 W
is not opted in: it has the heap, and a frozen copy would pin the library
version for everyone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
With DEBUG=1 the Makefile appends debug.conf to the board conf as a
";"-separated list and passed it to west unquoted:

    -Dzephyr-cp_EXTRA_CONF_FILE=.../board.conf;.../debug.conf

The unquoted ";" ends the shell command, so west ran with only board.conf and
the shell then tried to execute ".../debug.conf" as a program. Quote the whole
-D argument so the ";"-separated list reaches CMake intact. A DEBUG=1 build now
completes; verified it links and produces a bootable image
(build reaches "Completed 'zephyr-cp'", debug.conf is merged, CONFIG_DEBUG and
CONFIG_LOG_MODE_IMMEDIATE are set). Note debug.conf sets
CONFIG_LOG_MODE_IMMEDIATE=y, which serialises logging in-thread over UART0 and
distorts timing, so a DEBUG build must not be used to measure latency.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tyeth

tyeth commented Oct 3, 2026

Copy link
Copy Markdown
Owner Author

landed as adafruit/circuitpython#11423,

@tyeth tyeth closed this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant